Foundation Models & Post-training
Training objectives and inference methods that improve model capability without hiding failure modes behind aggregate scores.
Profile
I am a master's student in Software Engineering at Peking University. My research spans reinforcement learning for tool-using agents, LLM evaluation and reward modeling, diffusion language models, and multimodal or embodied intelligence.
My research aims to move language models from generating answers to taking reliable action: enabling autoregressive and diffusion models to learn from verifiable feedback, improve through tools, vision, and real-world interaction, and translate capability gains into measurable value.
At Tencent Youtu Lab (CSIG), I focus on post-training for Code Agents.
From failure modes to training, evaluation, and systems.
Training objectives and inference methods that improve model capability without hiding failure modes behind aggregate scores.
Credit assignment, tool use, and reward design for agents that must interact with environments rather than only generate text.
Robust LLM judges, long-form evaluation, uncertainty, and reward models whose scores reflect what users actually care about.
Visual feedback and robot reward models that connect language-model reasoning with perception and action.
Research, publication, and career updates.
Our paper Length-Adaptive Decoding for Masked Diffusion Machine Translation was accepted to the EMNLP 2026 Main Conference. The paper, code, data, and model weights are now publicly available.
I am currently working on Code Agent post-training at Tencent Youtu Lab, CSIG.
Our paper RefineSVG was accepted to ACM Multimedia 2026, a CCF-A conference.
We released new work on tool-calling reinforcement learning, diffusion-language-model reasoning, and robot reward modeling.
I joined Tencent to work on multilingual large-language-model research and agent systems.
I started my master’s study at Peking University.
Filter by topic, then open a paper for details.
Showing all 9 papers.
A closed-loop visual-feedback framework that lets multimodal language models render, compare, and iteratively correct SVG output, supported by an SVG-oriented vocabulary and agentic reinforcement learning.
ACM Multimedia (ACM MM) 2026
Entropy-Valley is a training-free method for selecting a masked diffusion model's target canvas before denoising, improving translation adequacy without a learned length predictor.
EMNLP 2026 Main Conference
Segment-Locked Credit Assignment routes execution rewards to tool tokens and preference rewards to summary tokens, reducing cross-segment credit contamination in tool-calling reinforcement learning.
Submitted to NeurIPS 2026
A matched-compute diagnostic showing that deterministic PRM guidance can trail simple outcome-reward reranking because of intermediate-state signal decay, diversity collapse, and readout errors.
Submitted to NeurIPS 2026
A selective ranking framework that starts from efficient pointwise scores and invokes pairwise comparison only inside ambiguous near-tie spans.
Submitted to AAAI 2027
A lightweight predictor of cross-rubric support for a frozen LLM judge's verdict, improving confidence alignment and selective review without changing the verdict.
Submitted to AAAI 2027
A systematic study of 23 aggregation rules showing that error localization and impact scoring are nearly independent capabilities that should be evaluated and modeled separately.
Submitted to AAAI 2027
Counterfactual masked views and crossed supervision train separate process and answer scorers so misleading reasoning presentation is less likely to override objective answer correctness.
Submitted to AAAI 2027
A multi-paradigm robot reward-modeling framework whose POISE module calibrates pointwise scores to respect more reliable pairwise preferences.
Submitted to NeurIPS 2026
Under-review manuscripts are shown with their target venues; CCF labels describe venue class only. Classification follows the CCF seventh-edition directory. View directory
Research methods applied to real data, models, and systems.

Large Language Model Algorithm Research Intern
Beijing, China · On-site
Working on post-training for Code Agents at Tencent Youtu Lab, with a focus on reinforcement learning and tool-using agent behavior.

Large Language Model Algorithm Research Intern
Shenzhen, China · On-site
Built a one-million-pair multilingual translation corpus for QQ International spanning 13 Eurasian languages, covering data filtering and cleaning, low-resource language-pair distillation, NER entity repair, 4B-class base-model evaluation, and SFT. The resulting model improved in-domain performance by approximately 10 percentage points while matching a substantially larger MoE translation baseline on FLORES-200 and WMT25.

NLP Algorithm Intern
Beijing, China · On-site
Developed an LLM-based advertising-slogan recognition system for player-community moderation. Combined retrieval-augmented generation with LoRA fine-tuning and deployed efficient inference with vLLM. The system reached an F1 score of 85% and supported hourly offline batch recognition.
Select a school name or mark to visit its official website.

Master's in Software Engineering
GPA: 3.71/4.0. Coursework includes machine learning, natural language processing, large language model development, AI practice, and algorithm design.
Bachelor's in Digital Media Technology
Overall rank: 2/37, including first place in the 2022–2023 academic year.
Competitions, scholarships, and student leadership.
8th China International College Students' 'Internet+' Innovation and Entrepreneurship Competition
National College Student Digital Media Technology Works and Creativity Competition
2021 National Modeling Competition
Beijing Forestry University
Beijing Forestry University
Beijing Forestry University
Beijing Forestry University
Beijing Forestry University