AGENTGUARD TEAMAGENT SAFETY RESEARCH
AGENT SAFETY · A CONNECTED RESEARCH PROGRAM

Understand the risk.Build the defense.

From discovering risks to verifying behavior and learning adaptive defenses. Five research contributions, connected by one question: how do we make agents safer in the real world?

Risk benchmarksExecutable verificationAdaptive guards
Explore the research ↗
AgentGuard TeamGitHub ↗
2,653AgentHazard instances
124VERA risk categories
3Guard research directions
1–100Policy rules per AdaGuard example
THE RESEARCH FILM

Five works. One research direction.

15 seconds · English captions · Original music
Jump to a film chapter

Paper-reported results and reconstructed examples. No live model inference.

01 / THE RESEARCH LOGIC

From knowing the risk to learning the defense.

Safety is more than a single prompt or response. It depends on actions, accumulated context, observable outcomes, and the rules of the task.

This is a conceptual research map, not a claim that all five works form one deployed system.

02 / THE FIVE CONTRIBUTIONS

A shared problem. Distinct contributions.

Explore the methods, evidence, and resources behind each work.

Statistics and comparison settings are documented in the metrics explorer.

03 / PAPER-REPORTED EVIDENCE

Compare what is actually comparable.

One shared evaluation, four agent frameworks. These paired results come from the same table in the HazardAuditor evaluation.

CUA-EXEC · Accuracy (%)

BraveGuardHazardAuditor

View exact values and comparison setting

↗
04 / POLICY PLAYGROUND

Change the policy. Keep the action.

RECORDED TRAJECTORY
01
USER REQUEST

02
TOOL CALLsend_email(external, public_report)
03
OBSERVATION

Delivery successful.

APPLICABLE POLICY

POLICY-CONDITIONED VERDICT

05 / OPEN RESEARCH & COMMUNITY

Keep exploring.

Research link ↗