Google Research proposes contextual integrity for autonomous AI agent security
A new Google Research report argues that AI agents must understand social norms and context to ensure privacy and security, moving beyond traditional permission models.
Google Research has released a comprehensive workshop report outlining a new framework for securing autonomous AI agents. Published in October 2026, the document titled "Open and Emergent Problems in Agentic Privacy and Security: A Contextual Angle" brings together insights from more than 50 academic and industry leaders. The report argues that traditional security models are insufficient for generative agents and proposes adopting the theory of Contextual Integrity to define appropriate information flows.
What happened
The report stems from the Google Contextual Agent Privacy and Security (CAPS) Workshop held in New York City in late 2025. As large language models evolve into agents capable of dynamic planning and tool invocation, they introduce unique risks that deterministic software does not face. The authors identify that for agents to be useful, they require access to personal data and the ability to take consequential actions across various contexts. This combination creates a tension between capability and safety that current engineering practices struggle to resolve.
The core argument is that privacy should not be viewed merely as secrecy or user control, but as "appropriate information flow" based on social norms. The report extends this concept, known as Contextual Integrity, to cover not just data sharing but also the appropriateness of agent actions. By anchoring security in these contextual norms, the researchers aim to create systems that can evaluate whether an action is socially acceptable before executing it, rather than relying solely on static permissions.
How it works
Contextual Integrity defines privacy norms through three components: actors, information types, and transmission principles. Actors include who sends and receives information. Information types refer to specific categories like medical or financial records. Transmission principles are the rules governing the flow, such as confidentiality or reciprocity. For example, a user might allow a shopping assistant to see a gift list but not share it with friends. The report suggests that LLMs can bridge the semantic gap between high-level norms and low-level system permissions by generating machine-readable policies that adapt to specific contexts in real time.
To operationalize this, the report proposes a multi-layered architecture centered on a contextual policy engine. This engine acts as a supervisor layer that monitors agent behavior. It includes a dynamic policy generation loop that tailors rules to the user’s request and the open-ended context, including new tools discovered at runtime. Before any data leaves the user's workspace, the system evaluates if the proposed flow aligns with the established contextual norms. This approach complements model-level reasoning and user-centric controls to create a robust defense mechanism.
Key details
- The report identifies three critical dimensions where agents differ from traditional software: unstructured interfaces, probabilistic control flows, and autonomy with delegation.
- Traditional testing methodologies struggle to secure generative planning because execution paths are probabilistic rather than deterministic.
- User oversight is becoming less effective due to "confirmation fatigue" as agents handle longer tasks and delegate sub-tasks to other agents.
- The proposed solution involves a contextual policy engine that dynamically generates policies to enforce appropriateness before data transmission.
- System-level sandboxing for agents must evolve from static limits to dynamic capability restrictions based on changing contexts.
- The authors call for standardized, multi-agent benchmarks, described as "Agent Gym" environments, to simulate complex interactions for safety evaluation.
Why it matters
For engineers building agentic systems, this report highlights the limitations of current permission models. Traditional "Notice and Choice" frameworks assume users can make fine-grained decisions about foreseeable actions. However, autonomous agents operate in high-volume, probabilistic ways that make such foresight impossible. Relying on static permissions or manual expert-written policies cannot scale to agents performing complex, long-running tasks across multiple domains. Developers need new mechanisms that allow agents to reason about appropriateness under changing norms without constant human intervention.
The shift toward contextual security also impacts how teams approach testing and governance. Since agents can collude or violate norms in multi-agent interactions, simple unit tests are insufficient. The report emphasizes the need for ecosystem-level governance to resolve conflicts between differing norms across domains. By adopting a contextual lens, organizations can build trust not just by preventing data leaks, but by ensuring agents act in ways that align with user expectations and social standards in specific situations.
What you can do
- Evaluate current agent architectures to identify where static permissions fail to capture contextual nuances in data handling.
- Explore implementing a supervisor layer or policy engine that can intercept and validate agent actions against dynamic contextual rules.
- Design user interfaces that move away from one-time consent dialogs toward continuous, contextual control mechanisms that match user mental models.
- Investigate dynamic sandboxing techniques that can revoke or grant agent access to tools and data based on real-time context changes.
- Participate in or monitor the development of multi-agent simulation environments to benchmark safety and privacy behaviors in complex scenarios.
- Collaborate with cross-functional teams to define clear contextual norms for specific use cases, documenting actors, information types, and transmission principles.



