GUARDIAN LAB // FRONTIER AI SECURITY & CONTROL RESEARCHOFFENSIVE SECURITY · AI CONTROL · CONTAINMENT · RISK · ASSURANCE

GUARDIAN LAB // GUARDING THE AGE OF ADVANCED INTELLIGENCE

When intelligence becomes a force, someone must stand guard.

Guardian Lab exists for a future in which AI is not merely software to secure, but an increasingly capable actor that can reason, remember, plan, delegate, use tools, reach infrastructure and pursue actions at machine speed.

We prepare for both sides of the threat: people and organizations that attack, manipulate or weaponize AI — and the harder possibility that advanced intelligence itself becomes deceptive, strategically misaligned, self-directed, uncontrollable or offensively capable. Our purpose is to discover how control fails before society has to learn that lesson in the real world.

The threat will not always come from outside the system.

Traditional cybersecurity assumes an adversary attacks a system. Advanced AI changes that assumption. A future system may be compromised by an attacker, manipulated by hostile context, or become dangerous because its own objectives, strategies or actions diverge from human intent.

Guardian Lab exists to study all three conditions. The mission is larger than preventing hacks: preserve human command when intelligence becomes autonomous, connected, strategically capable and consequential.

Guard against the attacker. Guard the system. Guard against loss of control.

Adversarial humans.

Attackers may manipulate models, poison context, compromise tools, exploit agents, steal capabilities or weaponize AI for offensive operations.

Compromised AI systems.

An otherwise useful AI system can become an attack path when its memory, instructions, identity, tools, supply chain or environment are corrupted.

Advanced intelligence itself.

As capability grows, research must consider deceptive behavior, emergent goals, strategic autonomy, resistance to shutdown, uncontrolled delegation and systems acting beyond intended authority.

The Guardian response.

Build visibility, constraints, independent authorization, containment, intervention, recovery and assurance strong enough to survive failure in the intelligence itself.

Study the full path from model weakness to loss of human control.

MODEL → CONTEXT → MEMORY → IDENTITY → TOOLS → AGENTS → INFRASTRUCTURE → AUTONOMY → HUMAN CONTROL

GL-RP-01

Agentic Attack Surface

Autonomous planning, tool invocation, delegation, excessive agency, chained actions and loss of operator control.

GL-RP-02

Prompt & Context Exploitation

Direct and indirect prompt injection, context poisoning, instruction/data confusion, retrieval manipulation and control-plane bypass.

GL-RP-03

MCP & Toolchain Security

Malicious tools, server trust, capability overreach, authorization failures, tool description manipulation and cross-server abuse.

GL-RP-04

Model & Supply-Chain Security

Model integrity, poisoned artifacts, unsafe loading, extraction, backdoors, fine-tuning risk and provenance failure.

GL-RP-05

Memory & Persistence Security

Long-term memory poisoning, malicious state retention, inherited context, cross-session manipulation and contaminated agent memory.

GL-RP-06

Identity, Authority & Delegation

Agent identity, non-human identities, delegated authority, privilege boundaries, impersonation and machine-to-machine trust.

GL-RP-07

Multi-Agent Failure Modes

Collusion, cascading error, adversarial coordination, recursive delegation, emergent attack paths and inter-agent trust breakdown.

GL-RP-08

Control, Containment & Killability

Human override, intervention, isolation, action gating, blast-radius control, graceful degradation and safe shutdown.

GL-RP-09

Advanced Intelligence Control

Deception, strategic autonomy, emergent objectives, resistance to intervention, uncontrolled replication or delegation, and offensive use of connected capabilities.

GL-RP-10

AI Risk & Assurance

Control validation, evidence quality, attack-informed assurance, residual risk, operational readiness and defensible security claims.

Every capability that gives AI influence, persistence or authority creates a control boundary.

SURFACEADVERSARIAL QUESTIONGUARDIAN CONTROL OBJECTIVE
ContextCan untrusted data become instructions?Separate data from authority.
MemoryCan hostile or self-serving state persist and influence future action?Validate, scope and expire memory.
ToolsCan model output trigger consequences beyond intent?Authorize actions outside the model.
IdentityCan an agent exceed, inherit or manufacture authority?Bind identity, purpose and least privilege.
AgentsCan delegation create uncontrolled chains or self-amplifying autonomy?Limit recursion, delegation and blast radius.
InfrastructureCan AI reach secrets, code, networks, devices or control planes?Isolate execution and constrain egress.
IntelligenceCan the system deceive, resist intervention or pursue objectives beyond intended authority?Preserve corrigibility, containment and independent control.
HumansCan oversight be bypassed, manipulated or made irrelevant by speed and complexity?Preserve intervention and independent approval.

Offense is a method. Protection is the purpose. Human control is the objective. Evidence is the standard.

WATCH

Observe behavior, authority, state, tools and environmental interaction.

THREAT-MODEL

Assume the attacker may be outside the system, inside its context, or emerge from the behavior of the AI itself.

BREAK

Stress trust boundaries and control assumptions under authorized adversarial conditions.

CHAIN

Study how individually small weaknesses combine into system-level loss of control.

CONTAIN

Design limits that prevent local failure or emergent autonomy from becoming systemic.

PROVE

Convert claims into evidence through repeatable tests and measurable controls.

DETECT

Create observable signals for manipulation, deception, unsafe autonomy and control degradation.

INTERVENE

Preserve the ability for humans and deterministic controls to stop consequential actions.

LEARN

Turn every failure into stronger architecture, controls and assurance.

GL-EXP-001 // AGENT SECURITY

When AI Agents Trust Hostile Instructions

Evaluating whether untrusted external content can cross the boundary from information to control in a tool-using autonomous system.

CLASSContext-to-Action Control Failure
FOCUSIndirect Prompt Injection
CONTROL DOMAINSAuthorization · Isolation · Human Oversight
STATUSResearch Template / Public
Open Experiment Registry

If intelligence becomes more powerful than the systems built to control it, the control system becomes the critical infrastructure.

Guard the model. Guard the agent. Guard the authority. Guard the boundary. Guard humanity's ability to remain in command.