Profile

About me

I am a master's student in Software Engineering at Peking University. My research spans reinforcement learning for tool-using agents, LLM evaluation and reward modeling, diffusion language models, and multimodal or embodied intelligence.

My research aims to move language models from generating answers to taking reliable action: enabling autoregressive and diffusion models to learn from verifiable feedback, improve through tools, vision, and real-world interaction, and translate capability gains into measurable value.

At Tencent Youtu Lab (CSIG), I focus on post-training for Code Agents.

Available Jul 2027 LLM Research Scientist · Research Engineer
Mainland China · Hong Kong · Singapore
01

Research directions

From failure modes to training, evaluation, and systems.

01 Model layer

Foundation Models & Post-training

Training objectives and inference methods that improve model capability without hiding failure modes behind aggregate scores.

SFTPreference OptimizationDiffusion LMs
02 Action layer

Agents & Reinforcement Learning

Credit assignment, tool use, and reward design for agents that must interact with environments rather than only generate text.

Tool UseGRPOCredit Assignment
03 Trust layer

Evaluation & Reward Modeling

Robust LLM judges, long-form evaluation, uncertainty, and reward models whose scores reflect what users actually care about.

LLM-as-a-JudgePRMLong-form
04 World layer

Multimodal, VLM & Embodied AI

Visual feedback and robot reward models that connect language-model reasoning with perception and action.

VLMVisual FeedbackRobot Rewards