Yucheng Shi

史淯城
Tencent Hunyuan LLM Frontier, Seattle — Research Scientist
Ph.D. in Computer Science, University of Georgia
Advised by Prof. Ninghao Liu

I work on agent reinforcement learning and self-evolving AI — building agents that improve through environmental interaction and self-generated experience. I currently lead projects on long-horizon terminal agents: evaluating sustained work, recursively synthesizing executable training environments, and stabilizing asynchronous post-training.

Yucheng Shi

News

Aug 2026Released Recursive Synthesis for Long-Horizon Terminal Tasks — 37,484 verified tasks across 15 rounds at about $0.05 per task. Paper · Blog arXiv
Jul 2026Released Stale but Stable — a staleness-adaptive trust region for stabilizing asynchronous RL. Paper · Blog arXiv
Jul 2026Released Long-Horizon-Terminal-Bench — 46 long-horizon terminal tasks with dense, subtask-level grading. Paper arXiv
Jun 2026New research note: Harness Handbook, with Ruhan Wang. Note
Summer 2026Leading a research program on self-evolving terminal agents with four student researchers.
Feb 2026Joined Tencent Hunyuan LLM Frontier in Seattle as a full-time Research Scientist.
Jan 2026Successfully defended my Ph.D. dissertation at the University of Georgia. Ph.D.
Jan 2026Preprint: MobileGUI-RL — online RL for mobile GUI agents. arXiv
May 2025Papers accepted at ICML 2025 and BIBM 2025. ICML
Apr 2025Received Dissertation Completion Award Assistantship for 2025–2026. Award
Jan 2025Three papers accepted at ICLR 2025. ICLR
Nov 2024MKRAG received Distinguished Paper Award at AMIA 2024. Award

Current Research

In summer 2026 I am leading a research program on self-evolving terminal agents, working with Zhongzhi Li, Ruhan Wang, Zongxia Li, and Junyao Yang. We are building an integrated stack for agents that must keep working over long horizons: dense evaluation, recursive synthesis of executable environments, readable agent harnesses, and stable asynchronous post-training.

A recursive verified synthesis pipeline that evolves complete executable task bundles into 37,484 increasingly difficult training tasks across 15 rounds.
Project lead · Data engine · Aug 2026 · Paper · Blog · Data & Models
A staleness-adaptive trust region that tightens PPO-style clipping only for high-mismatch tokens, stabilizing fully decoupled asynchronous RL.
Project lead · Agentic RL · Jul 2026 · Paper · Blog · Code
Forty-six long-horizon tasks across nine categories, with dense intermediate rewards that measure meaningful progress through sustained terminal workflows.
Project lead · Benchmark · Jul 2026 · Paper · Code · Leaderboard
A behavior-first map that helps people and coding agents locate where a harness behavior lives before editing it.
Project lead · Agent systems · Ruhan Wang · Jun 2026

Experience

Tencent Hunyuan LLM Frontier, Seattle Feb 2026 – Present
Research Scientist
  • Agent RL and self-evolving AI.
Netflix Sep – Dec 2025
ML Research Intern  ·  Dr. Ying Li, Dr. Yu Wang
  • Content recommendation with large language models.
Tencent AI Lab, Seattle May – Aug 2025
Research Scientist Intern  ·  Dr. Wenhao Yu
  • MobileGUI-RL: online RL for mobile GUI agent automation.
  • Multimodal AI for mobile UI understanding.
Harvard Medical School May – Sep 2024
Student Researcher  ·  Dr. Xiang Li
  • Post-trained LLaMA-3 70B on 6.5M+ radiology reports.
  • Built SearchRAG for medical question answering.

Selected Publications full list →

Recursive Synthesis for Long-Horizon Terminal Tasks
Zhongzhi Li*, Yucheng Shi*, Zongxia Li*, Ruhan Wang, Anhao Li, Zixun Huang, Junyao Yang, Lei Ke, Ninghao Liu, Haitao Mi, Leowei Liang
arXiv 2026 · Project lead · Paper · Blog · Data & Models
Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning
Junyao Yang, Yucheng Shi, Zongxia Li, Zhongzhi Li, Ruhan Wang, Xiangxin Zhou, Kishan Panaganti, Haitao Mi, Leowei Liang
arXiv 2026 · Project lead · Paper · Blog · Code
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading
Zongxia Li, Zhongzhi Li, Yucheng Shi, Ruhan Wang, Junyao Yang, Zhichao Liu, Xiyang Wu, Anhao Li, Yue Yu, Ninghao Liu, Lichao Sun, Haotao Mi, LeoweiLiang
arXiv 2026 · Paper · Project · Code · Leaderboard
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis
Yucheng Shi, Zhenwen Liang, Kishan Panaganti, Dian Yu, Wenhao Yu, Haitao Mi
Preprint 2026
MobileGUI-RL: Advancing Mobile GUI Agent through Reinforcement Learning in Online Environment
Yucheng Shi*, Wenhao Yu*, Zaitang Li, Yonglin Wang, Hongming Zhang, Ninghao Liu, Haitao Mi, Dong Yu
arXiv 2025 · Paper
Enhancing Cognition and Explainability of Multimodal Foundation Models with Self-Synthesized Data
Yucheng Shi, Quanzheng Li, Jin Sun, Xiang Li, Ninghao Liu
ICLR 2025 · Paper · Code · Model · 5k+ HF downloads
CORTEX: Concept-Oriented Token Explanation in Vector-Quantized Generative Models
Tianze Yang*, Yucheng Shi*, Mengnan Du, Xuansheng Wu, Qiaoyu Tan, Jin Sun, Ninghao Liu
ICML 2025 · Paper · Code
Black-box Backdoor Defense via Zero-shot Image Purification
Yucheng Shi, Mengnan Du, Xuansheng Wu, Zihan Guan, Jin Sun, Ninghao Liu
NeurIPS 2023 · Paper · Code
MKRAG: Medical Knowledge Retrieval Augmented Generation for Medical Question Answering
Yucheng Shi*, Shaochen Xu*, Tianze Yang*, Zhengliang Liu, Tianming Liu, Quanzheng Li, Xiang Li, Ninghao Liu
AMIA 2024 · Paper · Code · Distinguished Paper Award

Contact

The best way to reach me is email: mailofsyc@gmail.com. I am always glad to talk with people working on self-evolving agents, long-horizon interaction, and agent post-training.