Lesson 9 of 9 · 8 min
Why a firewall cannot stop prompt injection, but can flag exfiltration
A host firewall cannot read an AI agent’s instructions, but it can see where the agent connects and how much it sends. Learn how a per-agent baseline and an upload-spike signal limit the damage of an attack the firewall cannot prevent.
An AI agent such as a coding assistant reads text, decides what to do and acts on the Mac: it runs commands, opens files and calls tools. If some of that text was written by an attacker, the agent can be steered. This lesson separates two problems that are often confused: stopping the steering, and limiting what the steering can achieve.
What prompt injection is, and why a firewall does not see it
OWASP defines a prompt injection vulnerability as one that occurs when prompts alter a model’s behaviour or output in unintended ways. It also states that, because of the stochastic way models work, it is unclear whether fool-proof methods of prevention exist. Its mitigations are layered: constrain the model’s behaviour, filter inputs and outputs, enforce least privilege, require human approval for high-risk actions, separate untrusted content and test adversarially.
A host firewall sits at a different layer. It sees processes and their network connections. The instructions an agent receives travel inside encrypted connections: TLS is designed to prevent eavesdropping, tampering and message forgery, so the content is not visible from outside, and a firewall is not inside the agent. It also cannot see which tool an agent calls or which files it reads. A firewall therefore cannot detect, let alone prevent, a prompt injection.
What the attack still needs: a way out
MITRE ATT&CK groups the techniques adversaries use to steal data under the tactic Exfiltration. Its entries include sending data over an existing, legitimate web service instead of a dedicated command channel, and sending it in fixed-size chunks to stay under thresholds. Whatever the attacker’s method, stolen data has to leave the Mac through a network connection made by some process.
That is the point where a firewall can help. It cannot know why the agent connected, but it can know that the connection is unlike any the agent has made before, or that the volume is unlike any it has sent before. Those two facts can be measured without reading the content.
Signal 1: a first-ever destination
A coding agent has a routine. It talks to its own vendor’s service, to a package registry, to a code host. The set is small and stable. A baseline records it: for a period, the firewall notes the destinations each agent contacts, grouped by registrable domain so that one service with many hostnames counts once. After that, a destination outside the set is a departure from the agent’s own history, and a reasonable moment to ask the person.
Signal 2: an upload spike
Agents mostly download: documentation, packages, model responses. An hour in which an agent uploads far more than ever before stands out. A useful rule compares the current hour with the busiest hour so far, and also sets an absolute floor, so that a quiet agent’s small fluctuations do not raise flags. FireAI’s Agent profile, new in its next version, uses at least 4 times the busiest hour so far and never less than 25 MB. Only host names and byte counts are used, never the payload.
What the two signals do not catch
- Data sent to a destination the agent already uses, in amounts within its usual range, raises no flag. An attacker who knows the baseline can try to stay inside it.
- A baseline cannot be useful until it has learned something. During FireAI’s 3-day learning period nothing is flagged, and anything harmful done in that period can become part of normal (see the lesson on learning a baseline).
- Local damage, such as deleting files or reading a key without sending it, produces no network signal at all.
- Recognising the agent matters: traffic from a program that is not identified as an agent is not part of the agent’s baseline.
Layers, not substitutes
The two controls act on different things. An agent-side control judges what the agent is about to do. A network control observes what actually left the machine. Neither replaces least privilege, which keeps the agent from holding secrets it does not need, or human approval for high-risk actions. When you assess any product in this area, ask which layer it works at, what it cannot see from there, and what it does when it is unsure.
Key takeaways
- A host firewall cannot see an agent’s instructions or its tools, so it does not stop prompt injection.
- Stolen data still has to leave through a network connection, which a firewall can observe.
- A per-agent baseline turns a first-ever destination into a signal; an upload far above the busiest hour so far is a second one.
- Baselines miss traffic that stays inside normal, and cannot flag anything while they are still learning.
- Network controls and agent-side controls act on different things, and neither replaces least privilege.
Check yourself
1. Why can a host firewall not detect a prompt injection?
- Firewalls only work on Windows
- The instructions travel inside encrypted connections and the firewall is not inside the agent — Right.
- Prompt injection never uses the network
- Firewalls block all agents by default
TLS hides the content from outside, and a firewall sees processes and connections, not the agent’s reasoning, tools or file reads.
2. What can a firewall observe about an agent that helps limit the damage of an attack?
- The agent’s system prompt
- Which files the agent read
- Which destinations it connects to and how many bytes it sends — Right.
- Which MCP tools it installed
Host names and byte counts are visible without reading content, and stolen data must leave the Mac through such a connection.
3. Which transfer would a baseline of destinations and upload volumes most likely miss?
- A large upload to a server the agent has never contacted
- A small amount of data sent to a destination the agent uses every day — Right.
- An upload far above the busiest hour so far
- A connection from a never-seen domain
Traffic inside the normal destinations and normal volumes raises no flag, so an attacker who stays inside the baseline is not caught by it.
4. Why is nothing flagged during the learning period?
- The firewall is switched off
- There is no baseline yet to compare against — Right.
- Agents are exempt for three days
- Flags are only shown at night
A flag means a departure from the agent’s own history, so the firewall first has to record that history.
Do it with FireAI
Put this lesson into practice on your own Mac.
- How FireAI watches your Mac’s connections — Know which app is talking to the internet, in plain terms, without installing anything that runs as a hidden background service.
- Rules: app, website, domain, IP or a range, forever or until you restart — Write a rule as precise as one address or as broad as an entire domain.
- FireAI Pilot: FireAI decides the easy connections for you — Let FireAI clear the easy decisions on its own, and always see why.
- Deep inspection, without decrypting anything — Get real detail on a secure connection without FireAI ever reading what’s inside it.
- Security modes: Home, Coffee shop, Paranoid, Under attack — Match FireAI’s strictness to where your Mac actually is, in one tap.
- Coffee Shop Armor: safer on public Wi-Fi, and warned about fake networks — Sit down in any café, hotel or airport and let FireAI tighten up for you.
- The World map — See where your data actually goes, not just a hostname you’d have to look up yourself.
- Ask FireAI: plain-language orders instead of forms — Change what FireAI does by typing a sentence, not hunting through menus.
- Investigate a connection — Decide with the facts in front of you, not a vague warning.
Sources
- OWASP GenAI Security Project: LLM01:2025 Prompt Injection
- MITRE ATT&CK: Exfiltration (TA0010)
- RFC 8446: The Transport Layer Security (TLS) Protocol Version 1.3
- FireAI docs: Requests by country and upload spikes
Put it into practice on your Mac
Try every feature free for 17 days, no card needed.