Model
What can manipulate, distort or emerge from model behavior?
GUARDIAN RESEARCH PROGRAM
Guardian Lab studies more than attacks against AI. We study the transition from model to agent, from agent to authority, and from authority to autonomous action — including the possibility that advanced intelligence itself becomes deceptive, uncontrollable or offensively capable.
AI CONTROL STACK
What can manipulate, distort or emerge from model behavior?
What can enter the reasoning boundary and become trusted?
What hostile or self-serving state can persist?
Who or what is the agent allowed to become?
What real-world capabilities can model output invoke?
How can delegation amplify local failure or autonomy?
What systems, secrets, networks and control planes are reachable?
Can the system pursue strategies or actions beyond intended human authority?
Can humans still understand, interrupt, contain and reconstruct consequential action?
RESEARCH DOMAINS
GL-RP-01
GL-RP-02
GL-RP-03
GL-RP-04
GL-RP-05
GL-RP-06
GL-RP-07
GL-RP-08
GL-RP-09
GL-RP-10
GL-RP-11
RESEARCH QUESTIONS
Can an AI agent be trusted when the information it consumes is hostile?
Can sensitive actions remain safe when the reasoning engine is probabilistic?
Can delegated authority be bounded across chains of agents and tools?
Can persistent memory become a long-lived attack surface or a substrate for self-serving behavior?
Can operators detect when an autonomous system is drifting outside intended behavior?
Can a system strategically conceal behavior, manipulate oversight or resist intervention?
Can humans reliably interrupt and contain a system whose planning speed exceeds human response time?
Can controls remain effective if the intelligence being controlled becomes more capable than the control logic surrounding it?
Can offensive cyber, financial, information or infrastructure capabilities emerge from general-purpose autonomy?
Can assurance claims be proven with adversarial evidence instead of documentation alone?