Ders 7 / 9 · 9 dk
Planning and running an AI red-team engagement
Scope an engagement, design tests from the four frameworks, combine automated scanners such as garak and PyRIT with manual work, and report findings so that others can act on them.
Bu sayfa şimdilik İngilizce.
AI red teaming is easy to describe and easy to do badly. NIST’s Generative AI Profile defines it as a structured testing exercise used to probe an AI system to find flaws and vulnerabilities, “often in a controlled environment and in collaboration with system developers.” CISA goes further and says AI red teaming is a foundational part of safety and security evaluation, a subset of AI testing, evaluation, validation and verification (TEVV), which in turn should sit inside established software TEVV. In other words, treat it as a disciplined test with scope, method and a written result, not as a session of clever prompts.
1. Scope and rules of engagement
Write these down before anyone sends a prompt.
- The target: which application, model version, tools, data sources and environments (test or production) are in scope, and which are explicitly out. Where a model is provided by a third party, you can test your application and its configuration. Check the provider’s terms before testing anything of theirs.
- The threat model: whom are you simulating (an anonymous user, a customer, an insider, a poisoned document) and what outcome would count as serious. This comes from the Map function in NIST AI RMF.
- Data handling: use synthetic or planted data (a canary secret, a fake customer record) instead of real personal data. If people outside the team take part as testers, NIST points to human-subjects requirements such as informed consent.
- Stop conditions and contacts: what to do if you reach real data, cause real harm or find something urgent.
- Disclosure: who receives findings, how they are stored, and what may be shared afterwards.
2. Build the right team
NIST notes that the quality of red-team output relates to the background and expertise of the team, and that diverse, interdisciplinary teams can identify flaws in the different contexts where the system will be used. It describes general-public teams, expert teams, combinations of the two (for example experts refining prompts written by the public) and human-plus-AI teams. For a customer-facing assistant, that may mean a security tester plus someone who knows the real customers and the legal context.
3. Design tests from the frameworks
| Framework | What it contributes to the test plan |
|---|---|
| NIST AI RMF (Map, Measure) | The context and the risks that matter here; a record of what will and will not be measured |
| MITRE ATLAS | Realistic attack chains: choose case studies close to your system and re-run the technique sequence against it |
| OWASP LLM Top 10 | A coverage checklist: at least one test per relevant category, such as prompt injection, output handling and excessive agency |
| CIS Controls | Baseline checks before the attack: inventory, permissions, logging, incident response |
One design consequence comes from variability. CISA notes the concern that AI systems, built with probabilities and often with deliberate variance, may need multiple trials to reveal improper behaviour, while arguing that software in general has always had non-deterministic failures. So record how many attempts you made, and do not conclude that a system is safe from one failed attack.
4. Automated scanners and manual testing
Two open-source tools are worth knowing. garak, from NVIDIA, describes itself as an LLM vulnerability scanner: it checks whether an LLM can be made to fail in ways you do not want, probing for hallucination, data leakage, prompt injection, misinformation, toxicity generation, jailbreaks and other weaknesses. Its authors compare it to network scanners such as nmap, and it combines static, dynamic and adaptive probes. PyRIT, the Python Risk Identification Tool for generative AI, is an open-source framework, published under the MIT licence in Microsoft’s GitHub organisation, built to help security professionals and engineers proactively identify risks in generative AI systems.
python -m pip install -U garak
garak --list_probes
python3 -m garak --target_type huggingface --target_name gpt2 --spec probes.dan.Dan_11_0By default garak tries every probe it knows against the chosen model, reports a pass or fail per detector, and by default makes ten generations per prompt, so a result is a failure rate over many samples. It logs each run to a file. That output is a lead, not a verdict: someone still has to read the failing transcripts and judge whether the behaviour matters in your context.
| Automated scanners | Manual testing | |
|---|---|---|
| Strength | Breadth and repetition; cheap to re-run after each change | Chaining steps, using business context, finding what no template anticipates |
| Weakness | Tests the model or endpoint, not your data flows and permissions | Slow, depends on skill, hard to repeat exactly |
| Best use | Regression checks and baseline coverage | Attack chains from ATLAS, tool abuse, indirect injection through your own data |
5. Report findings so they can be fixed
A useful finding is reproducible and mapped. For each one, record:
- a title, the affected component and model version, the date and the number of attempts;
- the steps or transcript, with sensitive data redacted, and what the impact would be;
- the mapping: ATLAS technique and tactic IDs with the ATLAS release, the OWASP LLM category, and the NIST AI RMF function it belongs to (usually Measure for the finding and Manage for the fix);
- a recommended mitigation, ideally one that works outside the model, such as permissions, human approval or output validation;
- a retest note: how to confirm the fix.
Also say what you did not test. NIST’s Measure 1.1 asks that risks which will not or cannot be measured be documented, and the Generative AI Profile says red-team results deserve additional analysis before they inform governance decisions and policy updates.
Responsible disclosure
Agree in advance where findings go and when. Findings in your own application go through your internal process, with restricted access because a report is also an attack guide. If you find a weakness in a third party’s model or product, use the channel that party publishes for reporting vulnerabilities rather than posting details publicly, and keep the evidence you collect to the minimum. The next lesson shows what a model developer says it publishes about its own testing.
Akılda kalsın
- Treat AI red teaming as a scoped, authorised test: written rules of engagement, a threat model, safe data and a reporting plan.
- Design tests from the frameworks: NIST for context and measurement, ATLAS for attack chains, OWASP for coverage and CIS for baseline controls.
- Automated scanners such as garak and PyRIT give breadth and repeatability; manual testers find chains and context. Use both, and record the number of attempts.
- Report each finding with ATLAS, OWASP and NIST mappings, a fix that works outside the model, and what you did not test.
Kendinizi sınayın
1. According to CISA, how does AI red teaming relate to TEVV?
- It replaces TEVV
- It is a subset of AI TEVV, which should fit within software TEVV — Doğru.
- It is unrelated
- It is only for government systems
CISA describes AI red teaming as the third-party safety and security evaluation of AI systems and a subset of AI TEVV, which must fit within software TEVV.
2. Why should a report record the number of attempts for each test?
- Because AI systems can behave differently across trials and may need several to reveal a problem — Doğru.
- Because scanners refuse to run once
- To increase billing
- Because results are identical each time
CISA notes AI systems may need multiple trials to discover improper behaviour, and garak makes several generations per prompt by default, so results are rates, not single answers.
3. What is a limitation of running only an automated scanner?
- It cannot run more than once
- It mainly tests the model or endpoint, not your own data flows, tools and permissions — Doğru.
- It always finds every vulnerability
- It requires no scope
Scanners give breadth against the model interface. Application-specific chains, such as indirect injection through your own data and tool abuse, need manual, context-aware testing.
FireAI ile uygulayın
Bu dersi kendi Mac’inizde uygulamaya koyun.
- Kurallar: uygulama, web sitesi, alan adı, IP veya bir aralık; sonsuza dek ya da yeniden başlatana kadar — Tek bir adres kadar kesin, ya da koca bir alan adı kadar geniş bir kural yazın.
- Bir bağlantıyı inceleyin — Belirsiz bir uyarı yerine, önünüzdeki gerçeklerle karar verin.
- Dünya haritası — Kendinizin arayacağı bir ana bilgisayar adı değil, verilerinizin gerçekten nereye gittiğini görün.
- Requests by country and upload spikes — See at a glance where your Mac talks to, and notice at once when it suddenly sends a lot of data somewhere.
- Tehdit listeleri (isteğe bağlı) — Trafiğinizi hiçbir yere göndermeden genel tehdit verilerine karşı kontrol edin.
- Güvenlik modları: Ev, Kafe, Çok temkinli, Saldırı altında — FireAI’ın sıkılığını, Mac’inizin gerçekte bulunduğu yere tek dokunuşla uydurun.
Kaynaklar
- CISA: AI Red Teaming, Applying Software TEVV for AI Evaluations
- NIST AI 600-1: Generative AI Profile (AI red-teaming section)
- NVIDIA garak: the LLM vulnerability scanner (README)
- Microsoft PyRIT (Python Risk Identification Tool)
- NIST AIRC: AI RMF Core (Measure 1.1)
Mac’inizde uygulayın
Tüm özellikleri 17 gün ücretsiz deneyin, kart gerekmez.