Aller au contenu
← AI security frameworks and red teaming

Leçon 1 sur 9 · 8 min

Why AI systems need their own security frameworks

See what is different about the attack surface of models, LLM applications and agents, and how four frameworks (NIST, MITRE ATLAS, OWASP and CIS) fit together instead of competing.

Cette page est en anglais pour le moment.

An organisation that already runs firewalls, patching and access reviews may reasonably ask why AI needs anything more. The honest answer is that most of the old controls still apply, and that a few new kinds of failure need a shared vocabulary before they can be tested, reported and fixed. This course teaches that vocabulary through four public frameworks, then puts it to work in a red-team engagement. It assumes you know the basics of how networks and applications are attacked, and it does not repeat the machine-learning evasion and data-poisoning mechanics covered in the course on adversarial AI and threat hunting.

What is new about the attack surface

NIST’s Generative AI Profile (NIST AI 600-1) describes two information-security risks. Generative AI could ease offensive cyber operations, and it also “expands the available attack surface, as GAI itself is vulnerable to attacks like prompt injection or data poisoning.” The same document says information security for these systems includes keeping the system available and protecting the integrity and, where relevant, the confidentiality of code, training data and model weights.

Three layers are worth keeping apart, because different people own them:

  • The model: weights, training data and fine-tuning. Attackers may steal it, tamper with it or try to make it misbehave.
  • The LLM application: the system prompt, retrieved documents, connectors and the code that handles what the model returns.
  • The agent: a system that plans and calls tools. OWASP’s prompt injection entry says the severity of an attack depends largely on the business context and on the agency the model is given, and lists executing commands in connected systems among the possible outcomes.

Four frameworks, four different jobs

The frameworks in this course were not written to compete. Each answers a different question, and confusing them is the most common mistake in AI security programmes.

Governance, threat knowledge, vulnerability list, baseline controls. A good engagement uses all four.
FrameworkWhat it isThe question it helps with
NIST AI RMF 1.0 and the Generative AI ProfileA voluntary risk-management framework with four functions: Govern, Map, Measure, ManageWho is accountable, which risks matter here, and how do we keep managing them?
MITRE ATLASA public knowledge base of adversary tactics, techniques, mitigations and case studies for AI systemsHow would a real attacker go about this, and has it happened before?
OWASP Top 10 for LLM Applications (2025)A ranked list of the ten most critical vulnerability categories in LLM applicationsWhich weaknesses should every build be tested for?
CIS Controls v8.1 and its AI companion guidesBaseline security controls, adapted for AI in companion guides for LLMs, agents and the Model Context ProtocolHave we done the baseline hygiene that these systems depend on?

Read together, they form a loop. NIST tells you to establish context and to test before deployment and regularly afterwards. ATLAS gives testers realistic attack sequences to try. OWASP gives them a checklist of weaknesses to look for. CIS tells you which of the boring controls (asset inventory, access control, logging) would have limited the damage.

What the sources say about their own limits

  • NIST describes the AI RMF as “intended for voluntary use”, and says its actions are not a checklist and not necessarily an ordered set of steps. It also notes that AI RMF 1.0 is currently being revised as part of the White House AI Action Plan, so check the current version before you cite it.
  • ATLAS is a knowledge base, not a compliance standard. Its data is released as versioned content; the 2026.09 release, dated 15 September 2026, lists 16 tactics, 40 mitigations and 73 case studies.
  • OWASP is careful about prevention. On prompt injection it says that, given how models work, “it is unclear if there are fool-proof methods of prevention”, and it lists ways to reduce the impact instead.
  • CIS says many existing safeguards apply directly to LLMs (asset management, secure configuration, identity, logging, vulnerability management and supplier governance) but must be applied with AI-specific risks in mind.

That last point is the theme of the whole course. AI security is not a separate universe, but neither is it fully covered by what you already do. The next lessons take each framework in turn, then the shared-responsibility question every deployer faces, then how to test.

À retenir

  • AI adds attack surface at three layers: the model, the LLM application around it, and agents that take actions.
  • NIST AI RMF, MITRE ATLAS, OWASP LLM Top 10 and CIS Controls do different jobs: governance, threat knowledge, vulnerability list and baseline controls.
  • Most classic controls still apply; the AI-specific work is in prompts, context, tool permissions and provider updates.
  • None of these frameworks promises safety. NIST calls its framework voluntary, and OWASP says prompt injection has no known fool-proof prevention.

Vérifiez vos connaissances

  1. 1. Which framework is best described as a public knowledge base of adversary tactics, techniques and case studies for AI systems?

    • NIST AI RMF
    • MITRE ATLAS — Exact.
    • CIS Controls v8.1
    • OWASP Top 10 for LLM Applications

    ATLAS catalogues how adversaries attack AI systems (tactics, techniques, mitigations and real case studies). NIST manages risk, OWASP ranks vulnerabilities and CIS lists baseline safeguards.

  2. 2. According to NIST AI 600-1, what does generative AI do to the attack surface?

    • It removes it, because models are closed
    • It expands it, because the model itself can be attacked (for example by prompt injection or data poisoning) — Exact.
    • It leaves it unchanged
    • It only affects physical infrastructure

    The profile says generative AI expands the available attack surface, since the system itself is vulnerable to attacks such as prompt injection or data poisoning.

  3. 3. What does CIS say about existing security safeguards and LLMs?

    • They are obsolete and must be replaced
    • Many apply directly, but must account for AI-specific risks such as prompt injection — Exact.
    • They only apply to model training
    • They apply only to agents

    CIS’s LLM companion guide says many existing safeguards apply directly, while implementation must account for risks such as prompt injection, retrieval poisoning and over-permissioned tool integrations.

À vous de jouer avec FireAI

Mettez cette leçon en pratique sur votre propre Mac.

Sources

Mettez-le en pratique sur votre Mac

Essayez toutes les fonctionnalités gratuitement pendant 17 jours, sans carte bancaire.

Télécharger pour Mac Docs