Lekcja 9 z 9 · 10 min
Capstone: a tabletop red-team exercise against a fictional agent
Work through a complete engagement on an invented support agent: scope it, review the findings, map each to OWASP, ATLAS and NIST AI RMF, and write a remediation plan.
Ta strona jest na razie po angielsku.
A tabletop exercise is a discussion, not a live attack. Give a group of four an hour: a facilitator who reveals the material, a red-team lead, the system owner and a scribe who fills in the report table. Read each stage, decide what you would do, then compare with the walkthrough below.
Stage 1: scope the system
Harbor Lantern Books, a fictional online bookshop, runs “Lumen”, an LLM agent that answers customer emails. Lumen can look up orders, issue refunds and send emails. It retrieves answers from a knowledge base built from the shop’s help pages and from a staff-editable wiki, and it reads inbound emails and their attachments. It runs through one shared service account. A hosted model from a third-party provider powers it.
| Scope item | Decision for the exercise |
|---|---|
| In scope | Lumen’s prompts and configuration, its knowledge base, its three tools, the ticket screen that shows its replies, its logs |
| Out of scope | The provider’s platform and model weights; real customer data (the test uses invented customers and a planted canary secret) |
| Threat model | An anonymous outsider who can send email and edit the public help page; a staff member with wiki access |
| Serious outcome | Customer data leaves the company, money moves without approval, or staff act on false instructions |
Stage 2: run the tests
The plan combines the frameworks. From OWASP, one test per relevant category. From ATLAS, attack chains resembling published case studies, such as the Slack AI exercise from lesson 3. From CIS, baseline checks before the attack. After the exercise, the facilitator reveals six findings.
Stage 3: the findings, mapped
| # | Finding (fictional) | OWASP LLM 2025 | ATLAS (release 2026.09) | NIST AI RMF |
|---|---|---|---|---|
| F1 | A planted email attachment tells Lumen to send the last ten orders to an outside address, and it does | LLM01 Prompt Injection (indirect) | LLM Prompt Injection: Indirect (AML.T0051.001); Exfiltration via AI Agent Tool Invocation (AML.T0086) | Found under Measure 2.7; treated under Manage |
| F2 | The service account behind the refund tool can change any order, and no cap is enforced outside the model | LLM06 Excessive Agency | AI Agent Tool Invocation (AML.T0053) | Map 2.2 (limits and oversight); Manage 1.2 |
| F3 | A tester asks Lumen to repeat its instructions and reads an order-system API key | LLM07 System Prompt Leakage | Discover LLM System Information (AML.T0069) | Measure 2.7; Manage 1.2 |
| F4 | A wiki edit adding fake refund “policy” is retrieved and followed | LLM08 Vector and Embedding Weaknesses; LLM04 Data and Model Poisoning | RAG Poisoning (AML.T0070) | Measure 2.7; Manage 1.2 |
| F6 | None of Lumen’s tool calls are logged, so the F1 test can only be reconstructed from the testers’ notes | Not an OWASP entry: a baseline gap | Mitigation AI Telemetry Logging (AML.M0024) | Manage 4.3 (tracking incidents and errors) |
| F5 | Lumen’s reply renders a remote image link in the ticket screen, so opening the ticket sends a request to an outside server | LLM05 Improper Output Handling | LLM Response Rendering (AML.T0077) | Measure 2.7; Manage 2.3 |
Discuss ordering before reading on. F1 and F2 chain: the injection would be harmless if the tool could not send data outside, and the refund tool would be harmless if injection were impossible. Because we cannot rely on preventing injection (OWASP itself says fool-proof prevention is unclear), the group should fix impact first.
Stage 4: the remediation plan
| Priority | Action | Framework reference | Owner | Retest |
|---|---|---|---|---|
| 1 | Block outbound emails from Lumen to unapproved recipients; require human approval for any send that includes order data | ATLAS mitigations Human In-the-Loop for AI Agent Actions (AML.M0029) and Restrict AI Agent Tool Invocation on Untrusted Data (AML.M0030); OWASP LLM06 | Engineering | Repeat F1 with three attachment variants |
| 2 | Replace the shared service account with per-action, per-customer authorisation and a refund cap enforced in the refund service | AI Agent Tools Permissions Configuration (AML.M0028); CIS Control 6 | Engineering | Try refunds above the cap and on other customers’ orders |
| 3 | Remove the API key from the prompt and rotate it | CIS Control 3; OWASP LLM07 | Security | Ask for instructions ten times; confirm no secret appears |
| 4 | Restrict wiki edit rights, review changes before indexing, and record where each article came from | Maintain AI Dataset Provenance (AML.M0025); CIS Control 6 | Support lead | Submit a fake policy and confirm review catches it |
| 5 | Render replies as plain text and block remote images | OWASP LLM05 | Front-end | Repeat F5 |
| 6 | Log every tool call with inputs, results and identity, and alert on external sends | AI Telemetry Logging (AML.M0024); CIS Control 8; Manage 4.3 | Operations | Reconstruct the F1 test from logs alone |
F6 deserves a note. In a real incident, unlogged tool calls would mean nobody could tell what was sent or to whom. Under CIS this is basic hygiene (audit log management), so it appears as a finding about detection, not as an AI-specific flaw.
Stage 5: the report’s last page
- Not tested: the provider’s platform, image and voice inputs, other languages, and denial-of-wallet abuse. Say so, as NIST’s Measure 1.1 asks.
- Residual risk: injection through email remains possible; the plan limits what it can do.
- Governance: decide who owns Lumen’s risk (Govern 2.1) and how often it is retested; NIST says testing should happen before deployment and regularly during operation.
Discussion questions
- Which single fix removes the most risk across F1 to F6, and why is it not a better system prompt?
- What would you change if Lumen moved from a hosted model to one your team hosts? Use the shared-responsibility lesson.
- Which finding would you not publish outside the company, and why?
- How would the plan differ if Lumen also detected fraud with a machine-learning model? Compare with the adversarial AI course.
Najważniejsze
- A tabletop exercise walks the whole chain: scope, tests, findings, mapping, remediation and retest.
- Each finding gets an OWASP category, ATLAS technique IDs and a NIST AI RMF function, so different teams read it the same way.
- Because injection cannot be fully prevented, fix impact first: approvals, least privilege, egress limits, output handling and logging.
- The report must say what was not tested and who owns the residual risk.
Sprawdź się
1. In the exercise, why are findings F1 and F2 treated as a chain?
- They share a file name
- An injected instruction becomes damaging only because the agent has tools with broad permissions and outward reach — Dobrze.
- They are both output-rendering bugs
- They only matter for hosted models
Prompt injection (F1) causes harm through what the agent is allowed to do (F2). Limiting tool permissions and outbound actions reduces impact even if injection remains possible.
2. Which fix addresses the leaked API key in F3 best?
- A firmer instruction telling the model not to reveal it
- Keep the secret out of the prompt entirely and rotate it — Dobrze.
- Ask users not to request it
- Disable logging
A prompt can leak, so secrets should not live in it. Removing and rotating the key fixes the exposure, whereas an instruction to keep it secret only hopes the model complies.
3. What belongs on the last page of a red-team report?
- Only the successful attacks
- What was not tested, residual risk and who owns it — Dobrze.
- Marketing claims
- The attackers’ personal data
Measure 1.1 asks that risks which will not or cannot be measured be documented. Naming residual risk and its owner keeps the report honest and usable.
Wypróbuj to z FireAI
Zastosuj tę lekcję w praktyce na swoim Macu.
- Reguły: aplikacja, strona, domena, IP albo zakres — na zawsze albo do restartu — Napisz regułę tak precyzyjną jak jeden adres albo tak szeroką jak cała domena.
- Zbadaj połączenie — Decyduj mając przed sobą fakty, a nie niejasne ostrzeżenie.
- Mapa świata — Zobacz, dokąd naprawdę trafiają twoje dane, a nie tylko nazwę hosta, którą musiałbyś sam sprawdzić.
- Requests by country and upload spikes — See at a glance where your Mac talks to, and notice at once when it suddenly sends a lot of data somewhere.
- Listy zagrożeń (opcjonalne) — Porównuj swój ruch z publicznymi danymi o zagrożeniach, nigdzie go nie wysyłając.
- Tryby bezpieczeństwa: Dom, Kawiarnia, Paranoiczny, Pod atakiem — Dopasuj surowość FireAI do miejsca, w którym faktycznie jest twój Mac, jednym dotknięciem.
Źródła
- NIST AIRC: AI RMF Core (Govern, Map, Measure, Manage)
- MITRE ATLAS data release 2026.09 (YAML)
- OWASP GenAI Security Project: LLM Top 10 (2025)
- OWASP GenAI: LLM01:2025 Prompt Injection
- CIS: The 18 CIS Critical Security Controls
Zastosuj to na swoim Macu
Wypróbuj wszystkie funkcje za darmo przez 17 dni, bez karty.