GUARDIAN LAB // FRONTIER AI SECURITY & CONTROL RESEARCHOFFENSIVE SECURITY · AI CONTROL · CONTAINMENT · RISK · ASSURANCE

Research the point where artificial intelligence becomes a security actor.

Guardian Lab studies more than attacks against AI. We study the transition from model to agent, from agent to authority, and from authority to autonomous action — including the possibility that advanced intelligence itself becomes deceptive, uncontrollable or offensively capable.

01

Model

What can manipulate, distort or emerge from model behavior?

02

Context

What can enter the reasoning boundary and become trusted?

03

Memory

What hostile or self-serving state can persist?

04

Identity

Who or what is the agent allowed to become?

05

Tools

What real-world capabilities can model output invoke?

06

Agents

How can delegation amplify local failure or autonomy?

07

Infrastructure

What systems, secrets, networks and control planes are reachable?

08

Autonomy

Can the system pursue strategies or actions beyond intended human authority?

09

Human Control

Can humans still understand, interrupt, contain and reconstruct consequential action?

External attacks, compromised systems, emergent autonomy — one control mission.

GL-RP-01

Agentic Attack Surface

  • Excessive agency
  • Unsafe tool use
  • Recursive delegation
  • Privilege boundary failure
  • Autonomous attack chaining

GL-RP-02

Prompt & Context Exploitation

  • Direct prompt injection
  • Indirect prompt injection
  • Instruction/data confusion
  • RAG poisoning
  • Cross-context manipulation

GL-RP-03

MCP & Toolchain Security

  • Malicious MCP servers
  • Tool poisoning
  • Capability overreach
  • Authorization failure
  • Cross-server trust abuse

GL-RP-04

Model & Supply Chain

  • Model provenance
  • Backdoors
  • Poisoned artifacts
  • Unsafe loading
  • Fine-tuning integrity

GL-RP-05

Memory & Persistence

  • Long-term memory poisoning
  • Cross-session contamination
  • Persistent malicious state
  • State inheritance
  • Memory trust boundaries

GL-RP-06

Identity & Authority

  • Agent identity
  • Non-human identity
  • Delegation
  • Impersonation
  • Least privilege

GL-RP-07

Multi-Agent Security

  • Emergent attack paths
  • Agent collusion
  • Cascading failure
  • Recursive delegation
  • Trust graph compromise

GL-RP-08

Runtime & Containment

  • Sandbox boundaries
  • Network egress
  • Secrets exposure
  • Execution isolation
  • Killability

GL-RP-09

Adversarial ML

  • Evasion
  • Extraction
  • Poisoning
  • Robustness failure
  • Model behavior under adversarial input

GL-RP-10

Advanced Intelligence Control

  • Deceptive behavior
  • Emergent objectives
  • Strategic autonomy
  • Resistance to intervention
  • Uncontrolled replication or delegation
  • Offensive capability escalation

GL-RP-11

AI Risk & Assurance

  • Attack-informed assurance
  • Evidence quality
  • Control validation
  • Residual risk
  • Operational readiness

Questions that define security in the age of advanced intelligence.

Can an AI agent be trusted when the information it consumes is hostile?

Can sensitive actions remain safe when the reasoning engine is probabilistic?

Can delegated authority be bounded across chains of agents and tools?

Can persistent memory become a long-lived attack surface or a substrate for self-serving behavior?

Can operators detect when an autonomous system is drifting outside intended behavior?

Can a system strategically conceal behavior, manipulate oversight or resist intervention?

Can humans reliably interrupt and contain a system whose planning speed exceeds human response time?

Can controls remain effective if the intelligence being controlled becomes more capable than the control logic surrounding it?

Can offensive cyber, financial, information or infrastructure capabilities emerge from general-purpose autonomy?

Can assurance claims be proven with adversarial evidence instead of documentation alone?