Publications & preprints
Filter by topic, then open a paper for details.
Showing all 9 papers.
RefineSVG: Visual Feedback-Driven Reinforcement Learning for Image-to-SVG Generation
A closed-loop visual-feedback framework that lets multimodal language models render, compare, and iteratively correct SVG output, supported by an SVG-oriented vocabulary and agentic reinforcement learning.
ACM Multimedia (ACM MM) 2026
Length-Adaptive Decoding for Masked Diffusion Machine Translation
Entropy-Valley is a training-free method for selecting a masked diffusion model's target canvas before denoising, improving translation adequacy without a learned length predictor.
EMNLP 2026 Main Conference
SLCA-GRPO: Resolving Cross-Segment Credit Misattribution in Tool-Calling RL
Segment-Locked Credit Assignment routes execution rewards to tool tokens and preference rewards to summary tokens, reducing cross-segment credit contamination in tool-calling reinforcement learning.
Submitted to NeurIPS 2026
Why Deterministic PRM Guidance Underperforms in Discrete Diffusion Reasoning
A matched-compute diagnostic showing that deterministic PRM guidance can trail simple outcome-reward reranking because of intermediate-state signal decay, diversity collapse, and readout errors.
Submitted to NeurIPS 2026
EAGLE: Entropy-Adaptive Gating of Pointwise Scores for Pairwise LLM Judges
A selective ranking framework that starts from efficient pointwise scores and invokes pairwise comparison only inside ambiguous near-tie spans.
Submitted to AAAI 2027
AgreeJudge: Predicting Cross-Rubric Support for Frozen LLM Judges
A lightweight predictor of cross-rubric support for a frozen LLM judge's verdict, improving confidence alignment and selective review without changing the verdict.
Submitted to AAAI 2027
Decoupling Localization and Scoring in Long-form Reward Models
A systematic study of 23 aggregation rules showing that error localization and impact scoring are nearly independent capabilities that should be evaluated and modeled separately.
Submitted to AAAI 2027
DECAF: Decoupling Answer Correctness from Reasoning Presentation at the Parameter Level via Counterfactual Adversarial Fine-tuning
Counterfactual masked views and crossed supervision train separate process and answer scorers so misleading reasoning presentation is less likely to override objective answer correctness.
Submitted to AAAI 2027
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models
A multi-paradigm robot reward-modeling framework whose POISE module calibrates pointwise scores to respect more reliable pairwise preferences.
Submitted to NeurIPS 2026
Under-review manuscripts are shown with their target venues; CCF labels describe venue class only. Classification follows the CCF seventh-edition directory. View directory

