Profile
About me
I am a master's student in Software Engineering at Peking University. My research spans reinforcement learning for tool-using agents, LLM evaluation and reward modeling, diffusion language models, and multimodal or embodied intelligence.
My research aims to move language models from generating answers to taking reliable action: enabling autoregressive and diffusion models to learn from verifiable feedback, improve through tools, vision, and real-world interaction, and translate capability gains into measurable value.
At Tencent Youtu Lab (CSIG), I focus on post-training for Code Agents.
Research directions
From failure modes to training, evaluation, and systems.
Foundation Models & Post-training
Training objectives and inference methods that improve model capability without hiding failure modes behind aggregate scores.
Agents & Reinforcement Learning
Credit assignment, tool use, and reward design for agents that must interact with environments rather than only generate text.
Evaluation & Reward Modeling
Robust LLM judges, long-form evaluation, uncertainty, and reward models whose scores reflect what users actually care about.
Multimodal, VLM & Embodied AI
Visual feedback and robot reward models that connect language-model reasoning with perception and action.

