The FireAI Security Blog

By FireAI Security & Research Team · Published

How Anthropic red-teams Claude, and what it leaves to you

How Anthropic red-teams Claude, and what it leaves to you

Anthropic publishes more about how it tests Claude than most model developers: blog posts on red teaming, a Responsible Scaling Policy and system cards for each model. This note summarises those documents and separates three groups of work: methods that developers and deployers both use, work only a developer can do, and work that remains with the organisation that builds an application on the model. Internal processes beyond what Anthropic has published are not visible to outside readers and are not described here. The note is the second half of a pair with Shared responsibility for LLMs: who secures what.

Background: two different questions

A model developer asks whether a model is dangerous: whether it can provide serious uplift to weapons development, carry out cyber operations or behave deceptively. A deployer asks whether its application is secure: whether a crafted document, tool result or user message can make the application leak data or take an action it should not. The methods overlap, but the questions and the evidence differ, and a passing result on the first question does not answer the second.

What Anthropic describes

Methods that overlap with deployer practice

In a post of 12 June 2024, Anthropic groups its red teaming into domain-specific expert testing, testing that uses language models, work on new modalities and open-ended approaches. Automated red teaming uses a red team and blue team dynamic in which a model generates attacks. Multimodal red teaming covered image and text risks in Claude 3 models before deployment. Multilingual testing included a partnership with Singapore’s Infocomm Media Development Authority across four languages: English, Tamil, Mandarin and Malay. Anthropic also names crowdsourced and community red teaming, including events at DEF CON’s AI Village. [1]

The Claude Opus 5 system card, dated 24 July 2026, shows the same methods applied to one release. Its safeguards chapter uses single-turn harmful and benign requests, ambiguous-context prompts and multi-turn conversations in which a simulated user steers gradually toward harm. It reports harmlessness together with over-refusal, stating that the model kept high harmless response rates on harmful requests while keeping among the lowest over-refusal rates on benign ones. Its agentic safety chapter covers malicious use of coding and computer use agents and prompt injection robustness across coding, computer use and browser use. [7]

The system card defines prompt injection as a malicious instruction hidden in tool results that an agent processes, and notes that the risk is greatest when an agent can both reach private data and act on a user’s behalf. It also reports that safety instructions in the system prompt on claude.ai strengthened the model’s handling of harmful requests compared with the API without a system prompt. [7] The second finding bears directly on deployers, because the system prompt is part of the deployer’s side.

The Responsible Scaling Policy, version 3.4, effective 8 July 2026, sets capability thresholds in advance and requires formal evaluations at six-month intervals, and it provides for external reviewers of risk reports. [4] Fixing thresholds before testing limits the temptation to reinterpret a result after the fact, and the practice transfers to any organisation that defines release criteria.

Work only a model developer can do

Frontier threats red teaming targets chemical, biological, radiological and nuclear (CBRN) risks, cybersecurity and autonomous AI risks. A 2023 post describes domain experts with decades of experience defining threat models, more than 100 hours of expert probing, and a six-month biosecurity study of more than 150 hours. [3] A March 2025 post describes cyber evaluations based on capture-the-flag challenges and simulated network environments, and states that the Frontier Red Team worked with the US AI Safety Institute, the UK AI Security Institute and the US National Nuclear Security Administration (NNSA), the last on classified evaluations of nuclear and radiological knowledge. [2]

Anthropic’s Transparency Hub says the company uses both internal and external red teaming and that the UK AI Security Institute, the US Center for AI Standards and Innovation and Model Evaluation and Threat Research (METR) have carried out additional testing of its models. It also lists bug bounty programmes on HackerOne. [5] The Opus 5 system card includes cyber range testing from the UK AI Security Institute and an alignment assessment based on an automated behavioural audit, as well as a model welfare chapter. [7] The Transparency Hub page does not state which institute tested a given model before release, so the pre-deployment role of each is not established here beyond what the system card reports.

External and crowdsourced testing is a distinct category. HackerOne reports that Anthropic’s jailbreak challenge ran from 3 to 10 February 2025 with 339 participants, more than 300,000 chat interactions and eight difficulty levels, and that four teams shared 55,000 US dollars in bounties. Successful techniques included encoded prompts and ciphers, role-play, substituting harmful keywords with benign ones and prompt injection. [6] For Opus 5, the system card names three contracted external testers: one spent roughly 100 hours and completed one task with task-specific prompting, one spent around 16 hours without a successful jailbreak, and one ran an automated attacker with 150 attempts per task without success. [7]

What stays with deployers

None of the developer testing above evaluates a particular application. The following remain with the organisation that deploys the model, and the system card itself points to several of them, for example by showing that system prompts change model behaviour and that agents are most exposed when they combine private data with the ability to act. [7]

  • The system prompt and its guardrails, including how they behave under multi-turn pressure.
  • RAG data and indirect prompt injection: any document, page or email retrieved into context can carry instructions.
  • Tools, agent permissions and what an injected instruction could do with them.
  • Downstream handling of model output before it reaches a browser, shell, database or person.
  • Supply chain, including third-party plugins and Model Context Protocol (MCP) servers.
  • Cost abuse, such as unbounded loops or requests that consume paid tokens.
  • Sandboxing of code execution and browsing, and control of outbound network access.
  • Logging, monitoring and release gates that decide when a change may ship.

Practices worth borrowing

  1. Commission external red teamers, or run a private bug bounty, before major releases. Anthropic’s own practice includes both: contracted external testers for each release and a public jailbreak challenge with bounties.
  2. Run an alignment-style behaviour check on agents: sample transcripts of tool use and look for attempts to bypass restrictions, as Anthropic’s monitoring did for internal deployment.
  3. Write a short internal system card for each release: what was tested, which tests failed, over-refusal alongside harm, and known gaps.
  4. Set release thresholds before testing, following the Responsible Scaling Policy approach.
  5. Test prompt injection through every channel an agent reads, not only user input.

The FireAI University course on AI security frameworks and red teaming covers these methods in more detail.

Relevance to FireAI

FireAI operates on the Mac, below any model. Rules allow or block an app, or one destination for that app, so a local AI tool can be limited to the hosts it needs. This limits where data can go if an application misbehaves. FireAI does not test models, detect prompt injection, read prompts or filter model output.

Limitations

  • The note relies only on what Anthropic and HackerOne have published. Anthropic’s internal processes beyond that are not visible, and published summaries are selective.
  • Sources differ in date, from July 2023 to July 2026, and methods may have changed since the older posts.
  • Results described in a system card are self-reported by the developer, apart from work attributed to named external testers.
  • The note describes one developer. Other providers publish different material, and the comparison with deployer practice is a synthesis, not a finding from the sources.

How FireAI and HisnLabs fit in

Model testing ends at the API. What an app or agent may reach from your Mac is a separate control, and FireAI provides it.

FireAI is HisnLabs’ own product: an on-device AI firewall for Mac. It shows every connection your apps make, in plain language, and lets you decide what leaves your Mac — its AI runs locally, so your traffic is never sent to us or anyone else. HisnLabs’ security research team is the group that keeps that decision-making accurate: cataloguing which domains are ordinary telemetry versus a real product, tracking the country and network behind a connection, and training the on-device model (its FireAI Pilot feature) on real traffic patterns, all without any of it leaving your Mac.

You can read the technical decisions behind it, or try FireAI for 17 days, at FireAI, by HisnLabs.

Sources