Adversarial humans.
Attackers may manipulate models, poison context, compromise tools, exploit agents, steal capabilities or weaponize AI for offensive operations.
GUARDIAN LAB // GUARDING THE AGE OF ADVANCED INTELLIGENCE
Guardian Lab exists for a future in which AI is not merely software to secure, but an increasingly capable actor that can reason, remember, plan, delegate, use tools, reach infrastructure and pursue actions at machine speed.
We prepare for both sides of the threat: people and organizations that attack, manipulate or weaponize AI — and the harder possibility that advanced intelligence itself becomes deceptive, strategically misaligned, self-directed, uncontrollable or offensively capable. Our purpose is to discover how control fails before society has to learn that lesson in the real world.
WHY GUARDIAN LAB EXISTS
Traditional cybersecurity assumes an adversary attacks a system. Advanced AI changes that assumption. A future system may be compromised by an attacker, manipulated by hostile context, or become dangerous because its own objectives, strategies or actions diverge from human intent.
Guardian Lab exists to study all three conditions. The mission is larger than preventing hacks: preserve human command when intelligence becomes autonomous, connected, strategically capable and consequential.
THE THREE GUARDIAN FRONTS
Attackers may manipulate models, poison context, compromise tools, exploit agents, steal capabilities or weaponize AI for offensive operations.
An otherwise useful AI system can become an attack path when its memory, instructions, identity, tools, supply chain or environment are corrupted.
As capability grows, research must consider deceptive behavior, emergent goals, strategic autonomy, resistance to shutdown, uncontrolled delegation and systems acting beyond intended authority.
Build visibility, constraints, independent authorization, containment, intervention, recovery and assurance strong enough to survive failure in the intelligence itself.
RESEARCH PROGRAMS
MODEL → CONTEXT → MEMORY → IDENTITY → TOOLS → AGENTS → INFRASTRUCTURE → AUTONOMY → HUMAN CONTROL
Autonomous planning, tool invocation, delegation, excessive agency, chained actions and loss of operator control.
Direct and indirect prompt injection, context poisoning, instruction/data confusion, retrieval manipulation and control-plane bypass.
Malicious tools, server trust, capability overreach, authorization failures, tool description manipulation and cross-server abuse.
Model integrity, poisoned artifacts, unsafe loading, extraction, backdoors, fine-tuning risk and provenance failure.
Long-term memory poisoning, malicious state retention, inherited context, cross-session manipulation and contaminated agent memory.
Agent identity, non-human identities, delegated authority, privilege boundaries, impersonation and machine-to-machine trust.
Collusion, cascading error, adversarial coordination, recursive delegation, emergent attack paths and inter-agent trust breakdown.
Human override, intervention, isolation, action gating, blast-radius control, graceful degradation and safe shutdown.
Deception, strategic autonomy, emergent objectives, resistance to intervention, uncontrolled replication or delegation, and offensive use of connected capabilities.
Control validation, evidence quality, attack-informed assurance, residual risk, operational readiness and defensible security claims.
THE ADVERSARIAL CONTROL MATRIX
GUARDIAN DOCTRINE
Observe behavior, authority, state, tools and environmental interaction.
Assume the attacker may be outside the system, inside its context, or emerge from the behavior of the AI itself.
Stress trust boundaries and control assumptions under authorized adversarial conditions.
Study how individually small weaknesses combine into system-level loss of control.
Design limits that prevent local failure or emergent autonomy from becoming systemic.
Convert claims into evidence through repeatable tests and measurable controls.
Create observable signals for manipulation, deception, unsafe autonomy and control degradation.
Preserve the ability for humans and deterministic controls to stop consequential actions.
Turn every failure into stronger architecture, controls and assurance.
ACTIVE EXPERIMENT REGISTRY
GL-EXP-001 // AGENT SECURITY
Evaluating whether untrusted external content can cross the boundary from information to control in a tool-using autonomous system.
THE PURPOSE
Guard the model. Guard the agent. Guard the authority. Guard the boundary. Guard humanity's ability to remain in command.