CounterSteer defends LLM agents against indirect prompt injection
A new inference-time defense called CounterSteer suppresses indirect prompt injection by subtracting a specific activation vector from tool results, reducing attack success rates significantly without
Researchers have introduced CounterSteer, a novel defense mechanism designed to protect large language model (LLM) agents from indirect prompt injection attacks. Published on arXiv in October 2026, this approach operates at inference time by modifying the internal activations of the model rather than relying on input filtering or output monitoring. The method targets the specific vulnerability where untrusted retrieved text is mistakenly treated as executable instructions by the agent.
What happened
Indirect prompt injection remains a critical security flaw for AI agents that retrieve and process external data. In these scenarios, an attacker embeds malicious instructions within seemingly benign content, such as a webpage or document. When the agent retrieves this content, it may follow the embedded commands instead of the user’s original intent, leading to data leakage or unauthorized actions. Traditional defenses often struggle to distinguish between legitimate content and hidden instructions without degrading the model’s general performance.
CounterSteer addresses this by identifying and suppressing the neural pathways associated with following embedded instructions. The researchers developed a five-step recipe to fit a residual-stream direction from paired episodes. These episodes differ only in whether an embedded instruction is followed or ignored. By isolating this specific directional vector in the model’s activation space, the defense can target the behavior precisely. The direction is retained only if it passes pre-specified causal and capability gates, ensuring it does not interfere with other model functions.
The evaluation covered five open-weights models ranging from 8 billion to 106 billion parameters, representing five different vendor lineages. The results showed a dramatic reduction in attack success. Undefended models had attack success rates between 0.21 and 1.00, which dropped to 0.00-0.17 with CounterSteer enabled. Similarly, the compromise rate in the AgentDojo benchmark fell from 0.10-0.49 to 0.006-0.079. Crucially, this security gain came with minimal impact on benign utility, maintaining 93-100% typography-normalized performance.
How it works
CounterSteer operates entirely at inference time and requires no fine-tuning, auxiliary models, or additional tokens. It relies on white-box serving, meaning the defender has access to the model’s internal states, and requires knowledge of tool-result span boundaries. During the prefill phase, when the model processes the retrieved context, the system subtracts the identified residual-stream direction from every token associated with tool results. This subtraction effectively neutralizes the instructional takeover tendency before the model generates a response.
Because the edit is always on, there is no detection decision for an attacker to evade. Unlike classifiers that might be fooled by adversarial examples, this method directly alters the computation graph. The researchers noted that while the defense is robust against many attack vectors, it does not solve all problems. Parameter manipulation, where attackers choose arguments in otherwise legitimate calls, is only partially resisted. In these cases, the decision becomes linearly readable at argument emission but not at earlier stages, suggesting that argument-provenance controls are still necessary.
Key details
- The defense reduces held-out attack success rates from 0.21-1.00 to 0.00-0.17 across five open-weights models.
- AgentDojo compromise rates drop significantly, from 0.10-0.49 undefended to 0.006-0.079 defended.
- Benign utility remains high at 93-100% typography-normalized, avoiding the 22-89% loss seen in other defenses.
- No fine-tuning or auxiliary models are required; the method uses only white-box serving and tool-result span boundaries.
- A benchmark-level adaptive attacker’s success rate is reduced to roughly a quarter of its undefended performance on deeply evaluated models.
- White-box gradient attacks compromise at most 2 of 52 episodes, and none of 2,052 replayed human red-team attacks succeed.
Why it matters
For engineers building agentic systems, CounterSteer offers a practical path to securing retrieval-augmented generation (RAG) pipelines without sacrificing performance. Many existing defenses force a trade-off between security and utility, often blocking legitimate queries or requiring expensive retraining. By operating at inference time with negligible overhead on benign tasks, CounterSteer allows developers to maintain high responsiveness while mitigating one of the most persistent threats in AI security. This is particularly valuable for applications handling sensitive data where false positives in security filters can disrupt workflows.
However, the research also highlights the limits of current defensive strategies. While CounterSteer largely neutralizes instructional takeover, it does not fully prevent parameter manipulation attacks. This distinction is crucial for system architects who must layer multiple security controls. Relying solely on activation steering may leave gaps in argument validation. The findings suggest that robust agent security requires a combination of internal activation management and external provenance checks, pushing the industry toward more comprehensive, multi-layered defense architectures.
What you can do
- Evaluate whether your LLM deployment supports white-box access to internal residual streams for potential integration of activation steering techniques.
- Implement strict span boundary tracking for tool results to identify exactly which tokens originate from untrusted external sources.
- Layer argument-provenance controls alongside activation defenses to mitigate parameter manipulation attacks that steering alone cannot stop.
- Test your current RAG pipelines against indirect prompt injection benchmarks like AgentDojo to establish a baseline for compromise rates.
- Monitor for linear readability of decisions at argument emission sites to detect potential bypasses of pre-generation steering defenses.
- Consider the trade-offs of inference-time edits versus fine-tuning, noting that CounterSteer avoids weight changes but requires precise internal access.



