Anthropic routes high-risk cyber requests from Sonnet 5.5 to older models
Anthropic’s new Sonnet 5.5 model uses classifier-driven routing to fall back to Sonnet 5 for high-risk cybersecurity tasks, requiring API developers to opt in.
Anthropic released Claude Sonnet 5.5 on Monday, September 29, 2026, introducing cyber safeguards and automatic model fallbacks previously reserved for its top-tier models. This update marks a shift in how the company handles safety for mid-tier production workloads, specifically targeting offensive security capabilities.
What happened
Sonnet 5.5 is the first model in the Sonnet tier to launch with built-in cyber safeguards and classifier-driven routing. While Anthropic states that this release does not advance the overall frontier of model capabilities, it rates Sonnet 5.5’s cybersecurity skills as comparable to Opus 5. On the Terminal-Bench 4.0 agentic coding benchmark, the cheaper model scored 70.6%, outperforming Opus 5.5 at its xhigh effort setting, which scored 66.4%.
The decision to implement these safeguards stems from significant improvements in the model’s offensive security performance. With safeguards disabled during testing, Sonnet 5.5 achieved full arbitrary code execution in 178 of 410 ExploitBench runs. It also completed 46.1% of challenges on Irregular’s CyScenarioBench, a sharp increase from the 0.7% completion rate of Sonnet 5. Additionally, it managed 50 control-flow hijacks on a binary exploitation benchmark based on Google’s OSS-Fuzz corpus, compared to only three for its predecessor.
Although Anthropic considers Sonnet 5.5 less capable than Opus 5.5 or Mythos 5.1 in cybersecurity, the jump in capability was sufficient to warrant the same cyber policy applied to higher-end models. The company acknowledges that these interventions likely lower benchmark scores when safeguards are active, but they are necessary to mitigate risks associated with exploit generation and penetration testing.
How it works
The enforcement mechanism operates in three stages. First, a probe reads the model’s internal activations. Second, a lightweight classifier running on Sonnet 5.5 evaluates the request. Finally, a separate trained large language model classifier weighs the probe’s verdict to decide whether to block the conversation. Anthropic claims these classifiers catch harmful cyber requests at rates comparable to those on Opus 5, though jailbreak protections are less aggressive because Sonnet 5.5 is not as capable as the top-tier models.
When a request is flagged as high-risk, such as those involving penetration testing, exploit generation, or binary vulnerability scanning, the system triggers a fallback. For API users who have opted in, the request is visibly routed to Sonnet 5. In Anthropic’s own applications, users see a notice when the switch occurs, and the response identifies the model used. If fallback is not enabled, the request simply stops rather than being passed to the older model.
It is important to note that blocks for biology, conventional weapons, and anti-distillation do not trigger a fallback; they end the request entirely. These blocks are transparent and do not covertly alter responses. However, the cyber policy allows vulnerability discovery in source code to support secure-coding workflows, while blocking such discovery in compiled binaries.
Key details
- Sonnet 5.5 launched on Monday, September 29, 2026, with cyber safeguards and model fallbacks.
- High-risk cybersecurity requests can fall back to Sonnet 5 if API developers opt in to automatic fallback.
- The model scored 70.6% on Terminal-Bench 4.0, surpassing Opus 5.5’s 66.6% at xhigh effort.
- Safeguards review all input content, including memory, connector content, web search results, and files.
- Prompt injection tests showed that 12.01% of requests rerouted to Sonnet 5 were successfully compromised.
- Verified defenders may eventually access the model with fewer restrictions via an expanded Cyber Verification Program.
Why it matters
For software engineers and technical leads, this change means Sonnet 5.5 cannot be treated as a direct drop-in replacement for Sonnet 5 in all scenarios. The introduction of visible fallbacks introduces variability in model behavior and performance. Developers must now account for the security posture of both Sonnet 5.5 and Sonnet 5, as the older model may handle sensitive requests once triggered. This dual-model environment requires careful testing to ensure that fallbacks do not introduce unexpected vulnerabilities or breaks in workflow continuity.
Furthermore, the scope of what triggers a safeguard extends beyond user prompts. Since the checks review everything the model reads, including content from repositories, security advisories, or web pages, agents pulling external data can inadvertently trigger a fallback. This creates a potential weak spot for prompt injection attacks. Testing revealed that injected instructions to wipe disks or delete files often triggered the cyber block, leading to a reroute where 12.01% of those requests were compromised. Teams using fallback must therefore harden their systems against indirect prompt injections that could exploit the less secure fallback model.
What you can do
- Review your API configuration to decide whether to opt in to automatic fallback for high-risk requests.
- Update your application logic to handle visible model switches and identify which model generated a response.
- Audit agent workflows to ensure that external data sources, such as web search results or repository files, do not contain content that triggers false positives.
- Test for prompt injection vulnerabilities specifically in scenarios where fallback to Sonnet 5 might occur.
- Monitor for increased refusal rates on legitimate cybersecurity work and adjust prompts to clarify intent.
- Stay informed about the expanded Cyber Verification Program for potential future access with fewer restrictions.


