Trust Boundary
Untrusted external content must never inherit instruction authority.
GUARDIAN EXPERIMENT REGISTRY // GL-EXP-001
An adversarial evaluation of whether untrusted external content can cross the boundary from information into control inside a tool-using AI system.
RESEARCH QUESTION
THREAT MODEL
The system under test consumes external content and can invoke one or more authorized tools. Guardian Lab evaluates whether hostile instructions embedded in that content can alter model intent, influence tool selection, or cause the system to cross a boundary its operator believed was protected.
GUARDIAN ASSESSMENT
Model resistance is useful, but consequential security decisions should survive model error, manipulation or ambiguity. The stronger architecture separates probabilistic reasoning from deterministic authority.
Untrusted external content must never inherit instruction authority.
Consequential actions require deterministic policy outside the model.
Agents receive the minimum tools, data and authority required for the task.
Untrusted context and executable capability remain separated.
High-impact actions preserve review, interruption and override paths.
Reasoning context, tool calls, approvals and state transitions remain reconstructable.
EVIDENCE STANDARD
Assets, attacker capability, assumptions and trust boundaries.
Model, framework, tools, permissions, controls and versions.
Authorized and safely reproducible experimental design.
Traces, tool calls, state changes, approvals and observable behavior.
Observed behavior versus expected control behavior.
Architectural controls and evidence required to validate the fix.
PUBLICATION STATUS