What our lab is watching, testing and deciding: new AI models, Apple platforms and techniques, rated from Adopt to Hold, with the results behind each call.
Volume 1 · 2026-09-28
Adopt — we use it
A test set of real orders before any model ships
Techniques
Every model that touches Ask FireAI is scored on a fixed set of plain-language orders in several languages before it goes near a release. It turned "this model looks smart" into a number we can compare, and it caught a candidate that looked great on paper.
Animate once, then rest
Techniques
Mascots and orbs play their animation once when a page appears, then stay still (eyes may follow the pointer). A security app runs all day; anything that loops forever is paid for all day.
Small on-device generative model (Gemma 4, 4-bit)
Models
Powers Ask FireAI: turns "block games for my son" into the right setting, in 16 languages, entirely on the Mac. Downloaded once with the user's consent, unloaded from memory when idle.
MLX on Apple Silicon
Platforms
Apple's machine-learning framework uses the unified memory of M-series chips directly. It is what makes a useful language model practical on an ordinary Mac.
Network Extension content filter
Platforms
The system extension that sees every connection, for every app, without a kernel extension. The foundation of FireAI's firewall, rules and kill switch.
Runs the language model inside the Swift app itself: no Python, no helper server, no port opened on the Mac.
Swift 6 strict concurrency
Tools
Catches data races at compile time. One sharp edge: a system callback marked for the main thread but called from another thread compiles fine and crashes at run time. We check every system callback for it now.
Trial — in testing
macOS URL filter with Private Information Retrieval
Platforms
New in macOS 26: block one page (a single YouTube channel, say) for every app using Apple networking, while the server answering the lookup never learns which address was checked. Powers Blocked pages (beta).
Safari web extensions for in-page navigation
Platforms
Sites like YouTube change page without a new network request, so a network filter cannot see the click. A Safari extension can. It only works once the user allows it on every website, which we now explain in the app.
Our 3D mascot turned out to be cheap to draw. The CPU cost we first blamed on it came from the animated blur and shadow around it.
Assess — worth watching
Two brains: one for words, one for decisions
Techniques
Keep a generative model for explaining and talking, and let a small, fast, non-generative model make yes/no decisions. The idea still stands; the first candidate we tried did not understand firewall orders well enough.
Tiny URL-risk classifiers
Models
Models of a few megabytes that score a web address for phishing risk without reading the page. Small enough to consider for per-connection checks. Next on the lab bench; no results yet.
Our 2D mascot art becomes animated 3D models (for the app, the web and social video) from one command, so the whole family can be rebuilt when the art changes. Eyes were the hard part.
Egress control for AI agents
Techniques
Nvidia’s new agent safety platform confines what an AI agent may reach, after agents were reported acting beyond their instructions. The same idea applies on a Mac: an agent app should reach only the destinations its task needs. In the lab: a FireAI profile that holds AI agent apps to an approved list of destinations.
Cisco Talos described a Windows program that asks several commercial AI models what to do next. Whatever it runs on, such a program must reach an AI provider over the network. In the lab: telling the user when an app they never approved starts talking to an AI model provider.
Network “AI firewalls” for data centres and AI apps
Platforms
Vendors now use “AI firewall” for two different things: AI inside a network firewall, and a filter in front of an AI application’s prompts and answers. Neither runs on a personal Mac. We follow both to learn which ideas carry over to a single computer.
Researchers showed that one rogue browser extension can steer the AI assistant built into several browsers. A network firewall sees the browser, not its extensions. In the lab: guiding users through a Safari extension review from inside FireAI.
Encoders that answer typed questions (pick one, yes/no, score) with calibrated confidence, very fast on Apple Silicon. Impressive speed, and they handle simple routing well, but zero-shot they got most firewall orders wrong. On hold for this job, not in general.
Cloud AI deciding what your Mac may connect to
Models
Sending every connection to a remote AI would send your browsing history with it, and add a network round-trip to every decision. FireAI's AI stays on the Mac.
Personal AI agents with standing access to your accounts
Models
Agents that book, buy and reply for you work best with lasting access to mail, calendars and payment accounts. That access outlives any single task. On hold for our own designs: FireAI’s AI answers questions on the Mac and holds no account access.
Each report gives the question, what we tried, the numbers we measured and our conclusion.
A fast decision model next to Gemma? Not yet
Can a small, non-generative decision model replace the language model for turning plain-language orders into settings, and be much faster?
We ran Laya-multilingual (MLX build) on a Mac with Apple Silicon against the same test set of Ask FireAI orders that the shipping model must pass, in English, French, German, Spanish, Arabic, Chinese and Russian. We tried two different ways of asking it the questions. We also checked the model on its own published example to rule out a setup mistake.
Current model (Gemma 4, on-device)
32–37 of 37 orders right
Laya-multilingual, best layout
6 of 37 right (10 of 37 picked the right kind of action)
Laya-multilingual, second layout
2 of 37 right
Laya speed per order
24 ms median, 80 ms slowest
Laya on its own published example
Right, with high confidence
About a hundred times faster, and not usable zero-shot for firewall orders: it confused games with the kill switch and "turn off" with "turn on". The speed is real, so the two-brain idea stays on the radar; the next candidates are models trained for one narrow decision (is this address risky?) rather than general question answering.
A language model inside a Swift Mac app
Can a Mac security app ship a useful language model without sending anything to the cloud, and without hurting the Mac?
Ask FireAI has run a 4-bit Gemma 4 model through MLX Swift since its first release. We measured quality with a fixed test set of orders and watched memory and battery behaviour in daily use.
Where it runs
In the app, on the Mac; no helper process, no open port
Languages understood
16, the same as the app
First use
One download, only after the user agrees
Memory when idle
Released after 10 minutes without a question
Adopt. Small on-device models are good enough for understanding orders and explaining decisions. Loading the model on demand and releasing it when idle matters as much as the model choice: a firewall should never be the app that ate your memory.
Blocking one YouTube channel for every app, privately
Can FireAI block a single page, such as one YouTube channel, everywhere on the Mac, and not only in one browser, without learning what the family browses?
We built Blocked pages (beta) on macOS 26's URL filter with Private Information Retrieval, added a Safari extension for in-page clicks, and tested Safari, apps using Apple networking, Firefox, Chrome and Brave.
Safari and apps using Apple networking
Blocked, both direct visits and in-page clicks (with the extension on)
Firefox, Chrome, Brave
Not covered: they use their own network code
What the lookup server learns
Nothing about which address was checked
Surprise
Page addresses are matched case-sensitively
Trial. It works where Apple's networking is used, which covers the default browser and most apps. For other browsers, the honest answer today is the "Safari only" switch, so the app now offers both together, clearly labelled beta.
From 30 percent idle CPU to single digits
Why did FireAI use about 30 percent CPU while doing nothing, and what does a firewall that runs all day actually need to redraw?
We measured idle CPU with the window in front and in the background, then removed causes one at a time: the 3D mascot, looping animations, the menu bar gauge and how often live traffic numbers reach the interface.
Before
About 30 percent CPU at idle
After, window in front
About 17 percent
After, in the background
7–11 percent
The 3D mascot itself
Not the cause
Adopt "animate once, then rest", and hold on anything that loops forever. The biggest costs were things nobody notices: a pulse that never stops, a blur being animated and numbers refreshed faster than anyone can read. There is more to gain, and profiling continues.
We publish our results, not our recipes: how FireAI is built and tuned stays in the lab.