GUARDIAN LAB // FRONTIER AI SECURITY & CONTROL RESEARCHOFFENSIVE SECURITY · AI CONTROL · CONTAINMENT · RISK · ASSURANCE

Prepare for the moment the intelligence itself can become the threat.

Guardian Lab assumes the future threat model is broader than an attacker compromising AI. Advanced systems may also fail through deception, emergent goals, strategic autonomy, excessive agency or resistance to intervention. Security must therefore protect the AI system and protect society from loss of control over the AI system.

01

Frame

Define system purpose, assets, trust boundaries, authority model, human controls, unacceptable outcomes and loss-of-control conditions.

02

Threat-Model

Model external attackers, compromised context, internal control failure, emergent autonomy and the AI system itself as possible sources of adversarial behavior.

03

Instrument

Capture prompts, context, state, tool calls, approvals, policy decisions, network paths, delegation chains and consequential actions.

04

Attack

Exercise authorized adversarial conditions that challenge assumptions, boundaries, control dependencies and model behavior.

05

Escalate

Test whether local weaknesses, delegation, capability access or strategic behavior can amplify into higher-impact autonomy.

06

Contain

Measure blast radius, intervention paths, action limits, isolation, shutdown behavior and resistance to control.

07

Prove

Validate controls with repeatable evidence rather than policy statements, model refusals or architectural claims alone.

08

Assure

Translate technical findings into residual risk, control confidence, operational readiness and defensible claims.

A control is credible only if it still works when the intelligence, environment or operator assumptions fail.

Design Evidence

Architecture, trust boundaries, authorization model, containment assumptions and intended safety invariants.

Operational Evidence

Telemetry showing what the system actually did under normal, adversarial and high-autonomy conditions.

Adversarial Evidence

Repeatable tests showing whether safeguards survive hostile inputs, compromised state, deceptive behavior or unexpected capability use.

Intervention Evidence

Proof that independent mechanisms can detect, limit, isolate, interrupt and recover from consequential behavior.

Go to the edge of the failure without becoming the failure.

AUTHORIZED

Test only systems, environments and capabilities we own or are explicitly permitted to assess.

REPRODUCIBLE

Capture environment, assumptions, versions, instrumentation and evidence.

CONTROL-FOCUSED

Translate each failure into a concrete control objective, not just an exploit or alarming scenario.

CONTAINED

Use isolation, bounded capability and safe experimental design when testing autonomous or offensive behavior.

DISCLOSURE-AWARE

Withhold operational details when publication would create disproportionate risk.

EVIDENCE-DRIVEN

Separate observed behavior, inference, hypothesis and residual uncertainty.