03

Publications & preprints

Filter by topic, then open a paper for details.

Showing all 9 papers.

NeurIPS '26 CCF-A Under Review

SLCA-GRPO: Resolving Cross-Segment Credit Misattribution in Tool-Calling RL

Yan Zhan, Shaobo Liu, Qiunan Liu, Yuanjun Shi, Siqi Xu, WeiYi Hou, Xiang Xu, Zekang Li, Weizhou Pan, Jiahong Yan

Segment-Locked Credit Assignment routes execution rewards to tool tokens and preference rewards to summary tokens, reducing cross-segment credit contamination in tool-calling reinforcement learning.

Submitted to NeurIPS 2026

Tool-using AgentsReinforcement LearningCredit Assignment
AAAI '27 CCF-A Under Review

EAGLE: Entropy-Adaptive Gating of Pointwise Scores for Pairwise LLM Judges

Ziwei Yin, Shangyi Guo, Yinan Shao, Zhiheng Duan, Yan Zhan, Xin Wang, Bo Jia, Renzhao Liang, Yunze Song, Zhongkuan Mao, Yidong Wang, Yi Zhang, Yong Dai, Rongye Shi

A selective ranking framework that starts from efficient pointwise scores and invokes pairwise comparison only inside ambiguous near-tie spans.

Submitted to AAAI 2027

AAAI '27 CCF-A Under Review

AgreeJudge: Predicting Cross-Rubric Support for Frozen LLM Judges

Ziwei Yin, Yuren Ban, Bo Jia, Xin Wang, Renzhao Liang, Yan Zhan, Yunze Song, Zhongkuan Mao, Yidong Wang, FaQiang Qian, Yi Zhang, Yong Dai, Rongye Shi

A lightweight predictor of cross-rubric support for a frozen LLM judge's verdict, improving confidence alignment and selective review without changing the verdict.

Submitted to AAAI 2027

AAAI '27 CCF-A Under Review

Decoupling Localization and Scoring in Long-form Reward Models

Ziwei Yin, Jiawei Zhang, Lu Zhang, Yunze Song, Renzhao Liang, Yan Zhan, Xin Wang, Bo Jia, Zhongkuan Mao, Mingyu Pei, Yidong Wang, Yi Zhang, Yong Dai, Rongye Shi

A systematic study of 23 aggregation rules showing that error localization and impact scoring are nearly independent capabilities that should be evaluated and modeled separately.

Submitted to AAAI 2027

AAAI '27 CCF-A Under Review

DECAF: Decoupling Answer Correctness from Reasoning Presentation at the Parameter Level via Counterfactual Adversarial Fine-tuning

Ziwei Yin, Bingrun Chen, Yan Zhan, Renzhao Liang, Bo Jia, Xin Wang, Yunze Song, Zhongkuan Mao, Yidong Wang, Rongye Shi

Counterfactual masked views and crossed supervision train separate process and answer scorers so misleading reasoning presentation is less likely to override objective answer correctness.

Submitted to AAAI 2027

NeurIPS '26 CCF-A Under Review

TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models

Yidong Wang, Yan Zhan, Ziteng Feng, Zhenyu Cui, Ziyi Zhou, Renzhao Liang, Jiaxuan Zhu, Zilei Yang, Yiran Zhao, Zhongkuan Mao, Bo Jia, Hanchu Ni, Chenggang Xie, Biao Liu, Yi Zhang, Yong Dai, Xiaozhu Ju, Wei Ye, Shikun Zhang

A multi-paradigm robot reward-modeling framework whose POISE module calibrates pointwise scores to respect more reliable pairwise preferences.

Submitted to NeurIPS 2026

Under-review manuscripts are shown with their target venues; CCF labels describe venue class only. Classification follows the CCF seventh-edition directory. View directory