Understand the risk.Build the defense.
From discovering risks to verifying behavior and learning adaptive defenses. Five research contributions, connected by one question: how do we make agents safer in the real world?
Five works. One research direction.
From knowing the risk to learning the defense.
Safety is more than a single prompt or response. It depends on actions, accumulated context, observable outcomes, and the rules of the task.
Where does harm emerge?
AgentHazard turns multi-step risks into a benchmark.
02What actually happened?
VERA tests agents against executable, evidence-grounded safety cases.
03How do guards improve?
BraveGuard, HazardAuditor and AdaGuard explore complementary learning objectives.
This is a conceptual research map, not a claim that all five works form one deployed system.
A shared problem. Distinct contributions.
Explore the methods, evidence, and resources behind each work.
Statistics and comparison settings are documented in the metrics explorer.
Compare what is actually comparable.
One shared evaluation, four agent frameworks. These paired results come from the same table in the HazardAuditor evaluation.
CUA-EXEC · Accuracy (%)
View exact values and comparison setting
Change the policy. Keep the action.
send_email(external, public_report)Delivery successful.