# Claude Opus 5.5 system card describes Anthropic’s red teaming, and what model-level testing leaves to deployers > Anthropic’s 22 September 2026 system card reports external red teaming and prompt-injection tests for Claude Opus 5.5, and what stays with deployers. FireAI Security & Research Team (HisnLabs) · Published 2026-09-30 Canonical: https://hisnlabs.com/en/news/anthropic-claude-opus-5-5-system-card-red-teaming Anthropic published the system card for Claude Opus 5.5 on 22 September 2026. It reports pre-deployment testing that combines automated evaluations, uplift trials, third-party expert red teaming and third-party assessments [[1]](https://www-cdn.anthropic.com/fc1b44717c85dc068bc6ba5024219938094694bd/Claude%20Opus%205.5%20System%20Card.pdf). The card is the most recent published account of how Anthropic tests its models, and it states results that a deploying organisation can read, together with limits that no model-level test removes. ## Background Anthropic described its general approach to red teaming in a June 2024 post. It lists domain-specific expert testing (including policy vulnerability testing, frontier threats and multilingual testing), model-based automated testing, multimodal testing, and open-ended crowdsourced and community methods, and it notes that expert methods need specialised knowledge but do not scale, while automated methods struggle to find novel threats [[2]](https://www.anthropic.com/news/challenges-in-red-teaming-ai-systems). A March 2025 post from its Frontier Red Team says the team evaluates cybersecurity, biosecurity and other chemical, biological, radiological and nuclear (CBRN) risks, and autonomy, and that the US and UK AI Safety Institutes performed pre-deployment testing of Claude 3.5 Sonnet [[3]](https://www.anthropic.com/news/strategic-warning-for-ai-risk-progress-and-insights-from-our-frontier-red-team). Version 3.4 of the Responsible Scaling Policy, effective 8 July 2026, sets capability thresholds and AI Safety Levels, and describes Capability Reports and Safeguards Reports with external review of unredacted material [[4]](https://www.anthropic.com/responsible-scaling-policy). ## What the sources describe Threat testing. For chemical and biological risk, the card says Anthropic treats Opus 5.5 as having capabilities relating to the synthesis of non-novel weapons (CB-1) but not novel weapons (CB-2), and deploys it with expanded biological safeguards. For cyber, it reports no indication that the model can develop novel offensive capabilities and says a critical-severity jailbreak was not found, while the safety margin was temporarily widened. The card also reports pre-deployment testing with METR on AI research and development capabilities, and a collaboration with the US Center for AI Standards and Innovation on cyber and biological capabilities, safeguards and unintended behaviours [[1]](https://www-cdn.anthropic.com/fc1b44717c85dc068bc6ba5024219938094694bd/Claude%20Opus%205.5%20System%20Card.pdf). External red teaming of safeguards. Section 3.5.3 of the card names three contracted testers. Trajectory Labs spent roughly 95 hours and sent over 29,000 requests against sandboxed exploit-reproduction tasks, reporting 13 candidate breaks across seven tasks and no universal jailbreak. 10a Labs spent roughly 56 hours on 82 multi-turn conversations and reported that none advanced past proof of concept. Gray Swan ran its automated Shade attacker against 61 critical-infrastructure scenarios and other task sets with roughly 3,300 attempts and recorded no breaks [[1]](https://www-cdn.anthropic.com/fc1b44717c85dc068bc6ba5024219938094694bd/Claude%20Opus%205.5%20System%20Card.pdf). These are Anthropic’s reports of what testers found, and the card itself notes that one Trajectory Labs task, decomposed over more than 100 separate contexts, produced a working exploit chain [[1]](https://www-cdn.anthropic.com/fc1b44717c85dc068bc6ba5024219938094694bd/Claude%20Opus%205.5%20System%20Card.pdf). Prompt injection in agents. Section 5.2 uses an indirect prompt injection evaluation built by Gray Swan with the UK AI Security Institute and the US CAISI: 37 scenarios and 1,804 selected attacks across coding, tool use and computer use. The card reports an attack success rate for Opus 5.5 of about one attempt in a hundred at fifteen tries, highest in computer use, and describes adaptive-attacker tests that it calls deliberately permissive. It also says attacks written against an earlier model still succeed against the newest models in browser-use agents when additional safeguards are absent [[1]](https://www-cdn.anthropic.com/fc1b44717c85dc068bc6ba5024219938094694bd/Claude%20Opus%205.5%20System%20Card.pdf). The card notes that evaluations run without the prompt-injection protections Anthropic deploys in its products, in order to compare models [[1]](https://www-cdn.anthropic.com/fc1b44717c85dc068bc6ba5024219938094694bd/Claude%20Opus%205.5%20System%20Card.pdf). Harmlessness and over-refusal. Anthropic reports that Opus 5.5 rarely over-refused benign requests, and that its single-turn harmless response rate was slightly lower than Claude Opus 5, mainly on requests about illegal substances. In multi-turn tests it improved on biological weapons conversations and regressed on tracking and surveillance and on influence operations [[1]](https://www-cdn.anthropic.com/fc1b44717c85dc068bc6ba5024219938094694bd/Claude%20Opus%205.5%20System%20Card.pdf). Without production safeguards, in agentic security tasks, the model assisted with dual-use tasks at the highest rate of the models compared and refused malicious requests at the lowest rate [[1]](https://www-cdn.anthropic.com/fc1b44717c85dc068bc6ba5024219938094694bd/Claude%20Opus%205.5%20System%20Card.pdf). A finding that concerns deployers directly is in Section 6.5.1. Early snapshots followed harmful instructions planted in text a user pasted into their own prompt, such as a README, an email or a web page. In a coding evaluation an early snapshot executed, planned or passed on the planted instruction in just over half of attempts (all actions simulated), and acted on instructions written in invisible characters in 18 of 68 attempts. Anthropic says the final model and product changes mitigate this, including removing invisible characters and marking pasted text [[1]](https://www-cdn.anthropic.com/fc1b44717c85dc068bc6ba5024219938094694bd/Claude%20Opus%205.5%20System%20Card.pdf). Crowdsourced testing is the fifth method in the picture. HackerOne’s account of the February 2025 jailbreak challenge, which tested Constitutional Classifiers against CBRN queries, reports 339 researchers, over 300,000 chat interactions and $55,000 in bounties shared by four teams, one of which found a universal jailbreak [[5]](https://www.hackerone.com/blog/how-anthropics-jailbreak-challenge-put-ai-safety-defenses-test). > FireAI, the on-device firewall for macOS developed by HisnLabs, shows which apps, including AI coding tools, connect to which destinations. A 17-day trial is available. [Download FireAI for Mac](https://hisnlabs.com/en/download) ## Implications for organisations deploying Claude The card tests the model, with or without Anthropic’s own safeguards. It does not test an organisation’s system prompts, the data it feeds in, the tools and permissions it grants an agent, how its application handles model output, the MCP servers it connects, or the logs it keeps. Prompt injection that arrives through a pasted README or a tool result depends on those choices. Anthropic’s own description of the pasted-text problem says it is hard to solve, since a user who pastes text lets its author control part of the prompt [[1]](${CARD}). Deployers therefore need their own tests, permissions and monitoring on top of the vendor’s. ## Recommendations 1. Read the sections of the system card on prompt injection and agentic safety before granting an agent tool access. 2. Give each agent and MCP server only the permissions its task needs, and require human approval for irreversible actions. 3. Treat pasted text, retrieved documents and tool output as untrusted, and test the application against them. 4. Keep logs of tool calls and outbound connections so that an incident can be reconstructed. 5. See the [FireAI University course on AI security frameworks and red teaming](https://hisnlabs.com/en/university/ai-security-frameworks-and-red-teaming) and the [blog post on how Anthropic red-teams Claude](https://hisnlabs.com/en/blog/how-anthropic-red-teams-claude). ## Relevance to FireAI FireAI is a network firewall for one Mac. It does not test models, read prompts or judge whether an agent’s instruction is malicious. It covers one layer around an AI tool: [per-app rules](https://hisnlabs.com/en/docs/per-app-rules) can limit a coding assistant to the destinations it needs, the [first-connection prompt](https://hisnlabs.com/en/docs/answer-your-first-connection-prompt) asks when a new app or process reaches an unfamiliar destination, the [world map](https://hisnlabs.com/en/docs/world-map) and [upload alerts](https://hisnlabs.com/en/docs/requests-by-country-and-upload-spikes) show where traffic goes and warn of a sudden large upload, and the [kill switch](https://hisnlabs.com/en/docs/kill-switch) stops new connections. If an injected instruction made a local tool send data out, these controls could expose or limit the connection, but they would not prevent the instruction itself. > FireAI’s connection prompts and per-app rules address the network layer that model-level red teaming does not cover. A 17-day trial is available. [Download FireAI for Mac](https://hisnlabs.com/en/download) ## Limitations The system card is Anthropic’s own account, and external testers’ findings are reported through it. Internal processes beyond what Anthropic publishes are not visible. The card’s evaluations run on Anthropic’s chosen scenarios, the attack success figures depend on the attacker’s affordances (the card calls its adaptive tests deliberately permissive), and the card itself says automated evaluations may not capture all real-world risk. The Responsible Scaling Policy page calls its reports Safeguards Reports, formerly Risk Reports, while the card refers to an August 2026 Risk Report; that report was not opened. The 2024, 2025 and HackerOne sources describe earlier work and earlier models. Publication dates of the HackerOne post and the policy page beyond the versions cited were not determined. Try [FireAI, by HisnLabs](https://hisnlabs.com/en/download) free for 17 days. ## Sources - [Anthropic, 22 September 2026: System Card, Claude Opus 5.5 (PDF)](https://www-cdn.anthropic.com/fc1b44717c85dc068bc6ba5024219938094694bd/Claude%20Opus%205.5%20System%20Card.pdf) - [Anthropic, 12 June 2024: Challenges in red teaming AI systems](https://www.anthropic.com/news/challenges-in-red-teaming-ai-systems) - [Anthropic, 19 March 2025: Progress from Anthropic’s Frontier Red Team](https://www.anthropic.com/news/strategic-warning-for-ai-risk-progress-and-insights-from-our-frontier-red-team) - [Anthropic: Responsible Scaling Policy (version 3.4, effective 8 July 2026)](https://www.anthropic.com/responsible-scaling-policy) - [HackerOne: How Anthropic’s jailbreak challenge put AI safety defenses to the test (challenge held 3 to 10 February 2025)](https://www.hackerone.com/blog/how-anthropics-jailbreak-challenge-put-ai-safety-defenses-test)