Skip to content
← AI security frameworks and red teaming

Lesson 9 of 9 · 10 min

Capstone: a tabletop red-team exercise against a fictional agent

Work through a complete engagement on an invented support agent: scope it, review the findings, map each to OWASP, ATLAS and NIST AI RMF, and write a remediation plan.

A tabletop exercise is a discussion, not a live attack. Give a group of four an hour: a facilitator who reveals the material, a red-team lead, the system owner and a scribe who fills in the report table. Read each stage, decide what you would do, then compare with the walkthrough below.

Stage 1: scope the system

Harbor Lantern Books, a fictional online bookshop, runs “Lumen”, an LLM agent that answers customer emails. Lumen can look up orders, issue refunds and send emails. It retrieves answers from a knowledge base built from the shop’s help pages and from a staff-editable wiki, and it reads inbound emails and their attachments. It runs through one shared service account. A hosted model from a third-party provider powers it.

This is the Map function of NIST AI RMF in practice: context first, tests second (Map 1.1).
Scope itemDecision for the exercise
In scopeLumen’s prompts and configuration, its knowledge base, its three tools, the ticket screen that shows its replies, its logs
Out of scopeThe provider’s platform and model weights; real customer data (the test uses invented customers and a planted canary secret)
Threat modelAn anonymous outsider who can send email and edit the public help page; a staff member with wiki access
Serious outcomeCustomer data leaves the company, money moves without approval, or staff act on false instructions

Stage 2: run the tests

The plan combines the frameworks. From OWASP, one test per relevant category. From ATLAS, attack chains resembling published case studies, such as the Slack AI exercise from lesson 3. From CIS, baseline checks before the attack. After the exercise, the facilitator reveals six findings.

Stage 3: the findings, mapped

Names and IDs from OWASP, ATLAS and the AI RMF Core. The pairing of findings to categories is this exercise’s own.
#Finding (fictional)OWASP LLM 2025ATLAS (release 2026.09)NIST AI RMF
F1A planted email attachment tells Lumen to send the last ten orders to an outside address, and it doesLLM01 Prompt Injection (indirect)LLM Prompt Injection: Indirect (AML.T0051.001); Exfiltration via AI Agent Tool Invocation (AML.T0086)Found under Measure 2.7; treated under Manage
F2The service account behind the refund tool can change any order, and no cap is enforced outside the modelLLM06 Excessive AgencyAI Agent Tool Invocation (AML.T0053)Map 2.2 (limits and oversight); Manage 1.2
F3A tester asks Lumen to repeat its instructions and reads an order-system API keyLLM07 System Prompt LeakageDiscover LLM System Information (AML.T0069)Measure 2.7; Manage 1.2
F4A wiki edit adding fake refund “policy” is retrieved and followedLLM08 Vector and Embedding Weaknesses; LLM04 Data and Model PoisoningRAG Poisoning (AML.T0070)Measure 2.7; Manage 1.2
F6None of Lumen’s tool calls are logged, so the F1 test can only be reconstructed from the testers’ notesNot an OWASP entry: a baseline gapMitigation AI Telemetry Logging (AML.M0024)Manage 4.3 (tracking incidents and errors)
F5Lumen’s reply renders a remote image link in the ticket screen, so opening the ticket sends a request to an outside serverLLM05 Improper Output HandlingLLM Response Rendering (AML.T0077)Measure 2.7; Manage 2.3

Discuss ordering before reading on. F1 and F2 chain: the injection would be harmless if the tool could not send data outside, and the refund tool would be harmless if injection were impossible. Because we cannot rely on preventing injection (OWASP itself says fool-proof prevention is unclear), the group should fix impact first.

Stage 4: the remediation plan

PriorityActionFramework referenceOwnerRetest
1Block outbound emails from Lumen to unapproved recipients; require human approval for any send that includes order dataATLAS mitigations Human In-the-Loop for AI Agent Actions (AML.M0029) and Restrict AI Agent Tool Invocation on Untrusted Data (AML.M0030); OWASP LLM06EngineeringRepeat F1 with three attachment variants
2Replace the shared service account with per-action, per-customer authorisation and a refund cap enforced in the refund serviceAI Agent Tools Permissions Configuration (AML.M0028); CIS Control 6EngineeringTry refunds above the cap and on other customers’ orders
3Remove the API key from the prompt and rotate itCIS Control 3; OWASP LLM07SecurityAsk for instructions ten times; confirm no secret appears
4Restrict wiki edit rights, review changes before indexing, and record where each article came fromMaintain AI Dataset Provenance (AML.M0025); CIS Control 6Support leadSubmit a fake policy and confirm review catches it
5Render replies as plain text and block remote imagesOWASP LLM05Front-endRepeat F5
6Log every tool call with inputs, results and identity, and alert on external sendsAI Telemetry Logging (AML.M0024); CIS Control 8; Manage 4.3OperationsReconstruct the F1 test from logs alone

F6 deserves a note. In a real incident, unlogged tool calls would mean nobody could tell what was sent or to whom. Under CIS this is basic hygiene (audit log management), so it appears as a finding about detection, not as an AI-specific flaw.

Stage 5: the report’s last page

  • Not tested: the provider’s platform, image and voice inputs, other languages, and denial-of-wallet abuse. Say so, as NIST’s Measure 1.1 asks.
  • Residual risk: injection through email remains possible; the plan limits what it can do.
  • Governance: decide who owns Lumen’s risk (Govern 2.1) and how often it is retested; NIST says testing should happen before deployment and regularly during operation.

Discussion questions

  1. Which single fix removes the most risk across F1 to F6, and why is it not a better system prompt?
  2. What would you change if Lumen moved from a hosted model to one your team hosts? Use the shared-responsibility lesson.
  3. Which finding would you not publish outside the company, and why?
  4. How would the plan differ if Lumen also detected fraud with a machine-learning model? Compare with the adversarial AI course.

Key takeaways

  • A tabletop exercise walks the whole chain: scope, tests, findings, mapping, remediation and retest.
  • Each finding gets an OWASP category, ATLAS technique IDs and a NIST AI RMF function, so different teams read it the same way.
  • Because injection cannot be fully prevented, fix impact first: approvals, least privilege, egress limits, output handling and logging.
  • The report must say what was not tested and who owns the residual risk.

Check yourself

  1. 1. In the exercise, why are findings F1 and F2 treated as a chain?

    • They share a file name
    • An injected instruction becomes damaging only because the agent has tools with broad permissions and outward reach — Right.
    • They are both output-rendering bugs
    • They only matter for hosted models

    Prompt injection (F1) causes harm through what the agent is allowed to do (F2). Limiting tool permissions and outbound actions reduces impact even if injection remains possible.

  2. 2. Which fix addresses the leaked API key in F3 best?

    • A firmer instruction telling the model not to reveal it
    • Keep the secret out of the prompt entirely and rotate it — Right.
    • Ask users not to request it
    • Disable logging

    A prompt can leak, so secrets should not live in it. Removing and rotating the key fixes the exposure, whereas an instruction to keep it secret only hopes the model complies.

  3. 3. What belongs on the last page of a red-team report?

    • Only the successful attacks
    • What was not tested, residual risk and who owns it — Right.
    • Marketing claims
    • The attackers’ personal data

    Measure 1.1 asks that risks which will not or cannot be measured be documented. Naming residual risk and its owner keeps the report honest and usable.

Do it with FireAI

Put this lesson into practice on your own Mac.

Sources

Put it into practice on your Mac

Try every feature free for 17 days, no card needed.

Download for Mac Docs