News
Summer 2026Leading a research program on self-evolving terminal agents with four student researchers.
Feb 2026Joined Tencent Hunyuan LLM Frontier in Seattle as a full-time Research Scientist.
Jan 2026Successfully defended my Ph.D. dissertation at the University of Georgia. Ph.D.
Jan 2026Preprint: MobileGUI-RL — online RL for mobile GUI agents. arXiv
May 2025Papers accepted at ICML 2025 and BIBM 2025. ICML
Jan 2025Three papers accepted at ICLR 2025. ICLR
Current Research
In summer 2026 I am leading a research program on self-evolving terminal agents,
working with Zhongzhi Li, Ruhan Wang, Zongxia Li, and Junyao Yang. We are building an integrated
stack for agents that must keep working over long horizons: dense evaluation, recursive synthesis
of executable environments, readable agent harnesses, and stable asynchronous post-training.
A recursive verified synthesis pipeline that evolves complete executable task bundles into 37,484 increasingly difficult training tasks across 15 rounds.
A staleness-adaptive trust region that tightens PPO-style clipping only for high-mismatch tokens, stabilizing fully decoupled asynchronous RL.
Forty-six long-horizon tasks across nine categories, with dense intermediate rewards that measure meaningful progress through sustained terminal workflows.
A behavior-first map that helps people and coding agents locate where a harness behavior lives before editing it.
Project lead · Agent systems · Ruhan Wang · Jun 2026
Recursive Synthesis for Long-Horizon Terminal Tasks
Zhongzhi Li*, Yucheng Shi*, Zongxia Li*, Ruhan Wang, Anhao Li, Zixun Huang, Junyao Yang, Lei Ke, Ninghao Liu, Haitao Mi, Leowei Liang
Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning
Junyao Yang, Yucheng Shi, Zongxia Li, Zhongzhi Li, Ruhan Wang, Xiangxin Zhou, Kishan Panaganti, Haitao Mi, Leowei Liang
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading
Zongxia Li, Zhongzhi Li, Yucheng Shi, Ruhan Wang, Junyao Yang, Zhichao Liu, Xiyang Wu, Anhao Li, Yue Yu, Ninghao Liu, Lichao Sun, Haotao Mi, LeoweiLiang
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis
Yucheng Shi, Zhenwen Liang, Kishan Panaganti, Dian Yu, Wenhao Yu, Haitao Mi
Preprint 2026
MobileGUI-RL: Advancing Mobile GUI Agent through Reinforcement Learning in Online Environment
Yucheng Shi*, Wenhao Yu*, Zaitang Li, Yonglin Wang, Hongming Zhang, Ninghao Liu, Haitao Mi, Dong Yu
Enhancing Cognition and Explainability of Multimodal Foundation Models with Self-Synthesized Data
Yucheng Shi, Quanzheng Li, Jin Sun, Xiang Li, Ninghao Liu
CORTEX: Concept-Oriented Token Explanation in Vector-Quantized Generative Models
Tianze Yang*, Yucheng Shi*, Mengnan Du, Xuansheng Wu, Qiaoyu Tan, Jin Sun, Ninghao Liu
Black-box Backdoor Defense via Zero-shot Image Purification
Yucheng Shi, Mengnan Du, Xuansheng Wu, Zihan Guan, Jin Sun, Ninghao Liu
MKRAG: Medical Knowledge Retrieval Augmented Generation for Medical Question Answering
Yucheng Shi*, Shaochen Xu*, Tianze Yang*, Zhengliang Liu, Tianming Liu, Quanzheng Li, Xiang Li, Ninghao Liu