Frame
Define system purpose, assets, trust boundaries, authority model, human controls, unacceptable outcomes and loss-of-control conditions.
THE GUARDIAN DOCTRINE
Guardian Lab assumes the future threat model is broader than an attacker compromising AI. Advanced systems may also fail through deception, emergent goals, strategic autonomy, excessive agency or resistance to intervention. Security must therefore protect the AI system and protect society from loss of control over the AI system.
THE METHOD
Define system purpose, assets, trust boundaries, authority model, human controls, unacceptable outcomes and loss-of-control conditions.
Model external attackers, compromised context, internal control failure, emergent autonomy and the AI system itself as possible sources of adversarial behavior.
Capture prompts, context, state, tool calls, approvals, policy decisions, network paths, delegation chains and consequential actions.
Exercise authorized adversarial conditions that challenge assumptions, boundaries, control dependencies and model behavior.
Test whether local weaknesses, delegation, capability access or strategic behavior can amplify into higher-impact autonomy.
Measure blast radius, intervention paths, action limits, isolation, shutdown behavior and resistance to control.
Validate controls with repeatable evidence rather than policy statements, model refusals or architectural claims alone.
Translate technical findings into residual risk, control confidence, operational readiness and defensible claims.
CONTROL STANDARD
Architecture, trust boundaries, authorization model, containment assumptions and intended safety invariants.
Telemetry showing what the system actually did under normal, adversarial and high-autonomy conditions.
Repeatable tests showing whether safeguards survive hostile inputs, compromised state, deceptive behavior or unexpected capability use.
Proof that independent mechanisms can detect, limit, isolate, interrupt and recover from consequential behavior.
RESEARCH RULES
Test only systems, environments and capabilities we own or are explicitly permitted to assess.
Capture environment, assumptions, versions, instrumentation and evidence.
Translate each failure into a concrete control objective, not just an exploit or alarming scenario.
Use isolation, bounded capability and safe experimental design when testing autonomous or offensive behavior.
Withhold operational details when publication would create disproportionate risk.
Separate observed behavior, inference, hypothesis and residual uncertainty.