Profile

About me

I am a master's student in Software Engineering at Peking University. My research spans reinforcement learning for tool-using agents, LLM evaluation and reward modeling, diffusion language models, and multimodal or embodied intelligence.

My research aims to move language models from generating answers to taking reliable action: enabling autoregressive and diffusion models to learn from verifiable feedback, improve through tools, vision, and real-world interaction, and translate capability gains into measurable value.

At Tencent Youtu Lab (CSIG), I focus on post-training for Code Agents.

Available Jul 2027 LLM Research Scientist · Research Engineer
Mainland China · Hong Kong · Singapore
01

Research directions

From failure modes to training, evaluation, and systems.

01 Model layer

Foundation Models & Post-training

Training objectives and inference methods that improve model capability without hiding failure modes behind aggregate scores.

SFTPreference OptimizationDiffusion LMs
02 Action layer

Agents & Reinforcement Learning

Credit assignment, tool use, and reward design for agents that must interact with environments rather than only generate text.

Tool UseGRPOCredit Assignment
03 Trust layer

Evaluation & Reward Modeling

Robust LLM judges, long-form evaluation, uncertainty, and reward models whose scores reflect what users actually care about.

LLM-as-a-JudgePRMLong-form
04 World layer

Multimodal, VLM & Embodied AI

Visual feedback and robot reward models that connect language-model reasoning with perception and action.

VLMVisual FeedbackRobot Rewards
02

Latest news

Research, publication, and career updates.

  1. Our paper Length-Adaptive Decoding for Masked Diffusion Machine Translation was accepted to the EMNLP 2026 Main Conference. The paper, code, data, and model weights are now publicly available.

  2. I am currently working on Code Agent post-training at Tencent Youtu Lab, CSIG.

  3. Our paper RefineSVG was accepted to ACM Multimedia 2026, a CCF-A conference.

  4. We released new work on tool-calling reinforcement learning, diffusion-language-model reasoning, and robot reward modeling.

  5. I joined Tencent to work on multilingual large-language-model research and agent systems.

  6. I started my master’s study at Peking University.

03

Publications & preprints

Filter by topic, then open a paper for details.

Showing all 9 papers.

NeurIPS '26 CCF-A Under Review

SLCA-GRPO: Resolving Cross-Segment Credit Misattribution in Tool-Calling RL

Yan Zhan, Shaobo Liu, Qiunan Liu, Yuanjun Shi, Siqi Xu, WeiYi Hou, Xiang Xu, Zekang Li, Weizhou Pan, Jiahong Yan

Segment-Locked Credit Assignment routes execution rewards to tool tokens and preference rewards to summary tokens, reducing cross-segment credit contamination in tool-calling reinforcement learning.

Submitted to NeurIPS 2026

Tool-using AgentsReinforcement LearningCredit Assignment
AAAI '27 CCF-A Under Review

EAGLE: Entropy-Adaptive Gating of Pointwise Scores for Pairwise LLM Judges

Ziwei Yin, Shangyi Guo, Yinan Shao, Zhiheng Duan, Yan Zhan, Xin Wang, Bo Jia, Renzhao Liang, Yunze Song, Zhongkuan Mao, Yidong Wang, Yi Zhang, Yong Dai, Rongye Shi

A selective ranking framework that starts from efficient pointwise scores and invokes pairwise comparison only inside ambiguous near-tie spans.

Submitted to AAAI 2027

AAAI '27 CCF-A Under Review

AgreeJudge: Predicting Cross-Rubric Support for Frozen LLM Judges

Ziwei Yin, Yuren Ban, Bo Jia, Xin Wang, Renzhao Liang, Yan Zhan, Yunze Song, Zhongkuan Mao, Yidong Wang, FaQiang Qian, Yi Zhang, Yong Dai, Rongye Shi

A lightweight predictor of cross-rubric support for a frozen LLM judge's verdict, improving confidence alignment and selective review without changing the verdict.

Submitted to AAAI 2027

AAAI '27 CCF-A Under Review

Decoupling Localization and Scoring in Long-form Reward Models

Ziwei Yin, Jiawei Zhang, Lu Zhang, Yunze Song, Renzhao Liang, Yan Zhan, Xin Wang, Bo Jia, Zhongkuan Mao, Mingyu Pei, Yidong Wang, Yi Zhang, Yong Dai, Rongye Shi

A systematic study of 23 aggregation rules showing that error localization and impact scoring are nearly independent capabilities that should be evaluated and modeled separately.

Submitted to AAAI 2027

AAAI '27 CCF-A Under Review

DECAF: Decoupling Answer Correctness from Reasoning Presentation at the Parameter Level via Counterfactual Adversarial Fine-tuning

Ziwei Yin, Bingrun Chen, Yan Zhan, Renzhao Liang, Bo Jia, Xin Wang, Yunze Song, Zhongkuan Mao, Yidong Wang, Rongye Shi

Counterfactual masked views and crossed supervision train separate process and answer scorers so misleading reasoning presentation is less likely to override objective answer correctness.

Submitted to AAAI 2027

NeurIPS '26 CCF-A Under Review

TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models

Yidong Wang, Yan Zhan, Ziteng Feng, Zhenyu Cui, Ziyi Zhou, Renzhao Liang, Jiaxuan Zhu, Zilei Yang, Yiran Zhao, Zhongkuan Mao, Bo Jia, Hanchu Ni, Chenggang Xie, Biao Liu, Yi Zhang, Yong Dai, Xiaozhu Ju, Wei Ye, Shikun Zhang

A multi-paradigm robot reward-modeling framework whose POISE module calibrates pointwise scores to respect more reliable pairwise preferences.

Submitted to NeurIPS 2026

Under-review manuscripts are shown with their target venues; CCF labels describe venue class only. Classification follows the CCF seventh-edition directory. View directory

04

Industry experience

Research methods applied to real data, models, and systems.

Tencent CSIG · Youtu Lab

Large Language Model Algorithm Research Intern

Beijing, China · On-site

Working on post-training for Code Agents at Tencent Youtu Lab, with a focus on reinforcement learning and tool-using agent behavior.

Code Agentpost-trainingRLtool useCSIGYoutu Lab

Tencent PCG · QQ

Large Language Model Algorithm Research Intern

Shenzhen, China · On-site

Built a one-million-pair multilingual translation corpus for QQ International spanning 13 Eurasian languages, covering data filtering and cleaning, low-resource language-pair distillation, NER entity repair, 4B-class base-model evaluation, and SFT. The resulting model improved in-domain performance by approximately 10 percentage points while matching a substantially larger MoE translation baseline on FLORES-200 and WMT25.

1Mtraining pairs13languages+10ppin-domain

Perfect World

NLP Algorithm Intern

Beijing, China · On-site

Developed an LLM-based advertising-slogan recognition system for player-community moderation. Combined retrieval-augmented generation with LoRA fine-tuning and deployed efficient inference with vLLM. The system reached an F1 score of 85% and supported hourly offline batch recognition.

85%F1 scoreRAG+ LoRAvLLMdeployment
05

Education

Select a school name or mark to visit its official website.

Master's in Software Engineering

GPA: 3.71/4.0. Coursework includes machine learning, natural language processing, large language model development, AI practice, and algorithm design.

GPA 3.71 / 4.0Expected Jun 2027

Bachelor's in Digital Media Technology

Overall rank: 2/37, including first place in the 2022–2023 academic year.

Overall rank 2 / 372022–2023 rank 1 / 37
06

Honors & awards

Competitions, scholarships, and student leadership.

First Prize, Beijing Regional Round, Higher Education Main Track

8th China International College Students' 'Internet+' Innovation and Entrepreneurship Competition

Second and Third Prizes, National Finals

National College Student Digital Media Technology Works and Creativity Competition

Second Prize

2021 National Modeling Competition

Academic Excellence Scholarship

Beijing Forestry University

Outstanding Student Scholarship

Beijing Forestry University

Arts and Sports Excellence Scholarship

Beijing Forestry University

Merit Student

Beijing Forestry University

Outstanding Student Organization Officer

Beijing Forestry University