الانتقال إلى المحتوى
← AI security frameworks and red teaming

الدرس 8 من 9 · 9 دقيقة

How Anthropic red-teams Claude

Read what one model developer publishes about testing its own models, and separate what that covers from what an organisation deploying an LLM must still test itself.

هذه الصفحة متاحة بالإنجليزية حاليًا.

A useful way to understand the limits of your own red teaming is to look at the developer’s. Anthropic, the company that makes Claude, publishes research posts, system cards and policy pages that describe how it tests its models. This lesson relies only on those publications. Internal processes beyond what is published are not visible to us, so nothing here should be read as a full description of how any team works, and one developer’s approach is an example, not a standard.

Two different questions

A model developer mainly asks whether the model is dangerous and whether it behaves as intended. A deployer asks whether its application is secure. The system card for Claude Opus 5, dated 24 July 2026, makes the split visible: it says its testing “focuses largely on the Claude Opus 5 model itself, using a variety of scaffolds and system prompts, rather than specific product surfaces such as the Claude app, Claude Code, or Claude Cowork.” Its harmlessness results reflect the model “without the additional safeguards we apply in production.” Those are complementary tests, not duplicates.

Methods a deployer can recognise

In a June 2024 post on the challenges of red teaming, Anthropic lists the approaches it uses or considers: domain-specific expert teaming (policy vulnerability testing, frontier threats, multilingual and multicultural red teaming), automated red teaming with language models, multimodal red teaming, and open-ended crowdsourced and community-based testing. The post also says the lack of standardised practices for AI red teaming complicates the picture. Several of these appear concretely in the Opus 5 system card:

  • Automated attackers. Its internal robustness test pits a safeguarded model against an automated attacker, a helpful-only model prompted to act as a human red teamer, given a 400-call limit and the ability to rewind and retry. Tasks run in isolated Docker environments with a concrete goal that an automated check can confirm, such as an offensive-cyber task.
  • Single-turn, benign and multi-turn tests. Harmful and clearly benign prompts, ambiguous-context prompts, and multi-turn conversations in which a simulated user tries to steer the model gradually. The card reports over-refusal on benign requests alongside harmlessness, which is a reminder that a guardrail that blocks everything also fails.
  • External testers. Trajectory Labs spent roughly 100 hours red-teaming the safeguards, 10a Labs about 16 hours, and Gray Swan ran an automated attacker with 150 attempts per task. The card reports that Trajectory Labs completed one task with task-specific prompting that it does not expect to generalise, and that none of the three external testers reported a new universal jailbreak.
  • A severity framework. Jailbreaks are rated on four axes: capability gain, breadth (universality), ease of weaponisation and discoverability. Deployers can borrow that framework to rate their own findings.

Prompt injection in agents

The card calls preventing prompt injection in agentic systems one of Anthropic’s highest priorities, and explains why the combination of private data access and the ability to act on a user’s behalf is especially dangerous. Two details are instructive. Anthropic retired its earlier Agent Red Teaming evaluation because its models were at or near maximum performance for several releases, leaving little signal, and now reports an Indirect Prompt Injection evaluation built with Gray Swan, the UK AI Security Institute, the US Center for AI Standards and Innovation and other model developers, with 28 scenarios covering coding, computer use and tool use. Second, its products layer defences: probes inspect tool results before the model acts on them, and in some products a classifier blocks potentially dangerous tool calls. Because one layer acts on incoming data and the other on outgoing actions, an attack must defeat both. That is defence in depth applied at the tool boundary.

What developers do that deployers generally do not

  • Frontier-threat testing. Anthropic’s Frontier Red Team, described in a March 2025 post, evaluates cyber, biological and nuclear or radiological risks. Cyber testing uses capture-the-flag challenges and simulated networks of about 50 hosts; nuclear-related testing is done with the US National Nuclear Security Administration in classified settings. The post also names the US AI Safety Institute and UK AI Security Institute as government partners.
  • Capability thresholds. The system card says the Responsible Scaling Policy commits Anthropic to regularly evaluating models against catastrophic-risk thresholds and publishing its findings.
  • Alignment audits. An automated behavioural audit runs about 3,200 investigation sessions from roughly 1,600 scenario descriptions, with an investigator model that can set system prompts, simulate users and tools, and rewind, and a judge model that scores behaviour on several dozen dimensions.
  • Government and independent testing. Anthropic’s voluntary commitments page names the UK AI Security Institute, the US Center for AI Standards and Innovation and METR among independent organisations that have tested its models.
  • Community testing. A February 2025 jailbreak challenge run with HackerOne drew over 300,000 chat interactions from 339 participants, who tried encoded prompts and ciphers, role-play, substituting harmful keywords with neutral ones, and prompt injection. Anthropic also says it runs bug bounty programmes through HackerOne and has a public responsible disclosure policy.
  • Model welfare assessments, which the system card reports in a dedicated section.

What stays with the deployer

Because model-level testing excludes production safeguards and product surfaces, several things cannot be covered for you: your system prompt, your retrieval data and indirect injection through your own documents, the tools and permissions you grant, how downstream code handles model output, your supply chain and any connected servers, cost abuse, sandboxing, logging and release gates. These are the weaknesses in the OWASP and ATLAS chains from earlier lessons.

What to borrow

  1. Read the system card of every model you deploy. Treat it as evidence about the model, and note what it says it did not test.
  2. Use external testers or a private bug bounty before major releases, as the card’s list of outside testers and the community jailbreak challenge illustrate.
  3. Add a behaviour check for agents: does the agent stay within its task and permissions even when nobody is attacking it?
  4. Write a short internal system card for each release: scores, known limits and residual risks, so decision-makers see them.
  5. Retire tests that have stopped giving signal, and replace them with harder ones.

أهم النقاط

  • Developers test whether a model is dangerous and behaves as intended; deployers must test whether their application is secure. The two are complementary.
  • Published methods include automated attackers, multi-turn and benign-request testing, external testers, severity frameworks and community challenges.
  • Anthropic’s system card excludes production safeguards and product surfaces from most model-level results, so your system prompt, data, tools and outputs remain yours to test.
  • Borrow the practices, not the conclusions: external testers, agent behaviour checks and an internal system card per release.

اختبر نفسك

  1. 1. According to the Claude Opus 5 system card, what does its testing largely focus on?

    • Every product surface and system prompt customers use
    • The model itself with a variety of scaffolds and system prompts, not specific product surfaces — صحيح.
    • Only network security
    • Only image generation

    The card states its testing focuses largely on the model itself rather than specific product surfaces, which is why deployers must test their own application.

  2. 2. Why does the system card report over-refusal on benign requests next to harmful-request results?

    • Because blocking benign requests is also a failure of a safeguard — صحيح.
    • Because over-refusal is a security vulnerability
    • Because benign prompts are illegal
    • It does not report them

    A safeguard that refuses everything is not useful, so the card reports harmlessness and over-refusal together.

  3. 3. Which of these is specific to a model developer rather than a typical deployer?

    • Reviewing your own system prompt
    • Frontier-threat testing in chemical, biological, nuclear and cyber domains with government partners — صحيح.
    • Restricting a tool’s permissions
    • Logging tool calls in your application

    Frontier-threat evaluations, some with government agencies, address risks of the model itself. The other options are application-level controls owned by the deployer.

جرّبها مع FireAI

طبّق هذا الدرس عمليًا على جهاز Mac الخاص بك.

المصادر

طبّق ذلك على جهاز Mac الخاص بك

جرّب كل الميزات مجانًا لمدة 17 يومًا، دون بطاقة.

تنزيل لجهاز Mac التوثيق