Yunhao Feng

Trustworthy agents, from attacks to defenses

I study how autonomous AI agents fail, and how to make them safer in real execution environments.

My work focuses on computer-use and coding agents, including harmful trajectory benchmarks, skill-level backdoors, runtime monitoring, process-level verification, and training safer guard models from real interaction traces.

2,653 AgentHazard instances for multi-step harmful behavior evaluation
3,000+ SkillTrojan backdoored skills for skill-based agent security
80%+ BraveGuard trajectory-level safety detection accuracy

News

Recent updates and paper milestones.

  1. Google Scholar profile reaches 150 total citations.
  2. AgentHazard is accepted to ACM Multimedia 2026.
  3. JAIL is published in Expert Systems with Applications.
  4. BraveGuard and Safety in Self-Evolving LLM Agent Systems are released.
  5. Joined Ant Group as an algorithm engineer intern, focusing on agent safety.
  6. SkillTrojan is accepted to ICML 2026.
  7. BackdoorAgent is accepted to ACL Findings 2026.

Selected Publications

Curated from Google Scholar and verified public records; starred venues are emphasized.

Loading publications...

Experience

  • Ant Group, Department of Large Security Algorithm Engineer Intern · Agent Safety · 2026.05 - Present
  • Alibaba Future Living Lab Algorithm Engineer Intern · Coding Agent Safety · 2025.12 - 2026.05
  • Fudan University Visiting Student · 2025.05 - 2025.12
  • JD Explore Academy Algorithm Engineer Intern · 2022.02 - 2022.05

Education

  • National University of Defense Technology Ph.D., Management Science and Engineering · 2024.03 - Present
  • National University of Defense Technology M.Eng., Electronic Information · 2022.08 - 2024.03
  • Southwestern University of Finance and Economics B.Sc., Information Management and Information System · 2018.09 - 2022.06

Academic Service

  • Reviewer, NeurIPS 2026
  • Reviewer, ICML 2026
  • Reviewer, PRCV 2026
  • Reviewer, Knowledge-Based Systems
  • Reviewer, ACM Computing Surveys

Awards

  • Champion, Huawei Cup China Graduate AI Innovation Competition, LLM Security Track
  • Third Prize, China Graduate Electronic Design Contest
  • Second Prize, National College Mathematical Modeling Contest

Projects

Representative open-source projects, loaded from GitHub when available.