# The Best Tools of 2026 for AI Model Security > A balanced roundup of open tools that secure AI models and LLM applications in 2026, grouped by job, with maintainers, licences, limits and where network controls fit. FireAI Security & Research Team (HisnLabs) · Published 2026-09-30 Canonical: https://hisnlabs.com/en/blog/best-ai-model-security-tools-2026 AI model security is not one discipline but six jobs, each with its own tools: scanning model files, establishing provenance, probing a model for weaknesses, guarding it at run time, vetting the agent tooling around it, and limiting where the hosting application can connect. This roundup, prepared by HisnLabs, the developer of FireAI, covers open tools whose repositories or documentation were read on 30 September 2026. Licences, maintainers and status are stated as the primary sources give them, and no tool is ranked above another, because the jobs do not compete. ## Scope and method Each tool was included only if its official repository or documentation could be opened, its licence identified and its maintenance state checked (archived or active, and the date of its latest release where the release page shows one). Releases named below are the latest at the time of reading. Nothing here is a performance measurement: the sources do not provide comparable detection results, and none is claimed. Commercial platforms are omitted unless an open component is the subject. The categories follow the order in which a model moves through a project: a file is downloaded, its origin is checked, the model is tested, it is deployed behind guards, it is connected to tools, and the application that runs it reaches the network. The [OWASP Top 10 for Large Language Model Applications](https://owasp.org/www-project-top-10-for-large-language-model-applications/), maintained inside the OWASP GenAI Security Project, is the usual shared vocabulary for the application-level risks in categories three to five. ## 1. Model file scanning Many model files are serialised with Python’s pickle format, which can execute code when loaded. The Hugging Face documentation states that loading a pickle “means that code can be executed” and lists the `GLOBAL`, `STACK_GLOBAL` and `REDUCE` opcodes as the ones that pose the threat. [Hugging Face: Pickle scanning](https://huggingface.co/docs/hub/security-pickle) Scanners read the file’s instructions without loading it and flag dangerous imports. ### ModelScan [ModelScan](https://github.com/protectai/modelscan) is maintained by Protect AI under the Apache-2.0 licence. It scans pickle-based formats (PyTorch, scikit-learn, XGBoost, joblib, cloudpickle), TensorFlow SavedModel and Keras H5 and V3 files, ranks findings by severity and exits with codes suited to CI pipelines. The latest release on its page is v0.8.8 of 18 February 2026, and the repository showed activity on 28 September 2026. Use it as a gate before a downloaded model enters a build. Limitation: the README describes signature-based detection that can produce false positives and false negatives, and findings need human review. ### picklescan [picklescan](https://github.com/mmaitre314/picklescan) is an MIT-licensed scanner maintained by an individual account, mmaitre314. Hugging Face’s documentation names picklescan among the tools it has developed for the Hub. The latest release is v1.0.5 of 1 July 2026. It scans pickle files, PyTorch archives and ZIP containers, and can scan a local path, a URL or a Hugging Face model. Use it for a fast check in a script. Limitation: it matches dangerous operations by pattern, so a novel technique can pass, and it reports findings without removing them. ### fickling [fickling](https://github.com/trailofbits/fickling), from Trail of Bits, is licensed under LGPL-3.0. It is described as a decompiler, static analyser and bytecode rewriter for pickle. It can decompile a pickle into readable Python, check safety through static analysis and raise an error before loading an unsafe file. The latest release is v0.1.12 of 26 June 2026. Use it when a finding must be understood, not only flagged. Limitation: it is an analysis tool for pickle, and it does not cover the other formats that ModelScan reads. Its licence terms differ from the permissive licences of the other scanners. ### Hugging Face Hub scanning and safetensors The Hub scans every pushed file with ClamAV and with a pickle import scan that lists the imports inside each pickled file and highlights suspicious ones. [Hugging Face: file scanning](https://huggingface.co/docs/hub/security-malware) It also shows results from a third-party scanner, Protect AI’s Guardian, on public repositories. [Hugging Face: Protect AI](https://huggingface.co/docs/hub/security-protectai) Limitation, as Hugging Face states it: the pickle scan is not fully foolproof, its lists are maintained on a best-effort basis, and it is the user’s responsibility to check what is safe. The structural alternative is [safetensors](https://github.com/huggingface/safetensors), an Apache-2.0 format maintained by Hugging Face that stores tensors without executable content. It removes the pickle risk for weights but does not say whether the weights themselves are trustworthy. ## 2. Model provenance, signing and fingerprinting Scanning asks whether a file is dangerous to load. Provenance asks who produced it and whether it changed since. The two questions need different tools. ### Sigstore model-transparency [model-transparency](https://github.com/sigstore/model-transparency) is an Apache-2.0 project under the Sigstore organisation that signs and verifies machine-learning models. It can sign through Sigstore, or with keys, certificates or PKCS #11 devices, and produces a bundle containing a DSSE envelope with an in-toto statement. Its latest release is v1.1.1 of 10 October 2025, with commits in September 2026. Use it to let a consumer verify that the files they hold are those a publisher signed. Limitations: the documentation states support for hashing files and file shards, not individual tensors, and a valid signature proves origin and integrity, not that the model is safe. Hugging Face makes the same distinction for signed commits, which guarantee “the origin of the file” and not that it is safe. ### AI and ML bills of materials [CycloneDX](https://github.com/CycloneDX/specification), maintained by the OWASP Foundation and Ecma International’s TC54, lists a Machine Learning Bill of Materials (ML-BOM) among its supported BOM types. Its latest release, 1.7, was published on 21 October 2025, and its schemas are Apache-2.0. The [SPDX 3.0 AI profile](https://spdx.dev/learn/areas-of-interest/ai/) documents AI components such as models, datasets and prompts in one connected record. Use either to keep an inventory of which models, datasets and versions a product contains, which is what an incident response needs first. Limitation: a BOM records what an author declares. Neither format verifies the declaration. ### Fingerprinting and watermarking research tools [Instructional Fingerprinting](https://github.com/cnut1648/Model-Fingerprint) is the code of an academic paper ([arXiv 2401.12255](https://arxiv.org/abs/2401.12255)) that embeds a fingerprint in a language model so an owner can later test whether another model derives from it. It is MIT-licensed research code, and its last repository update was in July 2024. [MarkLLM](https://github.com/THU-BPM/MarkLLM), from the THU-BPM group, is an Apache-2.0 toolkit for watermarking the text an LLM generates, presented as an EMNLP 2024 demo. The two answer different questions: the first marks a model, the second marks its output. Use them to study attribution, not as a production control. Limitation: both are research artefacts, and neither substitutes for signing a distributed file. > FireAI, the on-device firewall for macOS developed by HisnLabs, does not scan model files or sign them. It controls which app on a Mac connects where. A 17-day trial is available. [Download FireAI for Mac](https://hisnlabs.com/en/download) ## 3. LLM red-teaming and vulnerability scanning These tools send adversarial inputs to a model or application and record how it responds. They test behaviour, not files. ### garak [garak](https://github.com/NVIDIA/garak) is an Apache-2.0 LLM vulnerability scanner maintained by NVIDIA. Its README says it checks “if an LLM can be made to fail in a way we don’t want”, with probes for prompt injection, data leakage, hallucination, toxicity and jailbreaks, each paired with a detector. The latest release is v0.17.0 of 9 September 2026. Use it as a broad baseline scan of a model endpoint. Limitations: it needs API access or a local deployment of the target, some probes are resource-intensive, and the README notes limited support for vision models and prototype status for some probes. ### PyRIT [PyRIT](https://github.com/microsoft/PyRIT) is an MIT-licensed framework described as built for security professionals and engineers to identify risks in generative AI systems. It is maintained by Microsoft, and the latest release is v1.1.0 of 4 September 2026. The earlier repository under the Azure organisation was archived on 27 March 2026 with a pointer to the Microsoft one, so links to the old address point to a read-only copy. Use it for scripted, multi-turn red-team campaigns by a team that will write code. Limitation: it is a framework, not a one-command scanner. ### promptfoo [promptfoo](https://github.com/promptfoo/promptfoo) is an MIT-licensed command-line tool and library for evaluating LLM applications and for red teaming and vulnerability scanning, with CI integration. The latest release is 0.123.1 of 18 September 2026. Its README states: “Promptfoo is now part of OpenAI. Promptfoo remains open source and MIT licensed.” Use it to put prompt tests and adversarial checks into a build. Limitation: results depend on the provider and the judge configured, and most providers require API keys. The change of ownership is a governance fact worth noting when choosing a tool for regulated work. ### Giskard [Giskard](https://github.com/Giskard-AI/giskard-oss) is an Apache-2.0 Python library from the Giskard-AI organisation. Its version 3 has two components described as stable: giskard-checks, a testing framework using scenarios and LLM-as-judge assessments, and giskard-scan, a scanner for agent vulnerabilities. Use it to test agent and retrieval-augmented systems. Limitations: version 3 requires Python 3.12 or later, and version 2, which covered tabular and classical machine-learning scanning, is no longer actively maintained. ## 4. Runtime guardrails and prompt-injection detection Guardrails sit in the request path and classify or constrain inputs and outputs while the application runs. ### Llama Guard, Prompt Guard and LlamaFirewall Meta’s [Purple Llama](https://github.com/meta-llama/PurpleLlama) project groups safeguards for open models. Llama Guard is a family of input and output moderation models. Prompt Guard 2 is described as a lightweight classifier for direct prompt-injection attempts, and LlamaFirewall combines it with an alignment-check scanner and a code scanner for agents. Per the project’s licence table, the models are distributed under Llama Community Licences and other components, such as Code Shield and the evaluations, are MIT. The models on Hugging Face are gated. Use them where a small classifier in front of a model is acceptable. Limitations: a classifier reduces attempts that look like known patterns and does not remove the class of attack. ### NeMo Guardrails [NeMo Guardrails](https://github.com/NVIDIA-NeMo/Guardrails) is an Apache-2.0 toolkit from NVIDIA for adding programmable rails to LLM applications, using a modelling language called Colang to control dialogue flow. The latest release is v0.24.1 of 16 September 2026. Use it to constrain topics, formats and dialogue paths. Limitation: rails are policy written by the developer, so their coverage is only as broad as what was anticipated. ### Rebuff and LLM Guard, now archived [Rebuff](https://github.com/protectai/rebuff), a self-hardening prompt-injection detector with heuristics, an LLM check, a vector store of past attacks and canary tokens, was archived on 16 May 2025. Its README states that it is a prototype and cannot provide complete protection against prompt injection. [LLM Guard](https://github.com/protectai/llm-guard), an MIT-licensed toolkit of input and output scanners from the same maintainer, was archived in July 2026. Both remain readable for reference. They illustrate a practical point: detector projects can lose their maintainers, so a guardrail should be replaceable. ## 5. Agent and MCP security scanners Agents add tools, and tool descriptions are text the model reads. Scanners in this group inspect that tooling and its configuration. ### Snyk Agent Scan [Agent Scan](https://github.com/snyk/agent-scan), formerly MCP-Scan, is an Apache-2.0 scanner maintained by Snyk. It discovers agent components on a machine, including MCP servers and skills across clients such as Claude, Cursor, VS Code and GitHub Copilot, and checks them for prompt injection, tool poisoning and related issues. The latest release is v0.6.8 of 29 September 2026. Use it to audit what a developer machine has installed. Limitations: the README states that it does not accept external contributions at this time, and that two release lines with different output formats exist. ### Cisco MCP Scanner [MCP Scanner](https://github.com/cisco-ai-defense/mcp-scanner) is an Apache-2.0 tool from Cisco AI Defense. It analyses MCP tools, prompts, resources and dependencies with three engines that can run alone or together: YARA rules, LLM analysis and the Cisco AI Defense API. The latest release is 4.8.4 of 28 August 2026. Use it for checking a server before adoption. Limitation: the LLM and API engines need external credentials, and a scan of a server’s description says little about what its code does after an update. ## 6. Where network controls fit The tools above examine files, models, prompts and configurations. None of them decides which application may open a connection to which server. That is the role of egress control, and it matters because the risks in the previous sections converge on the network: a tampered model runtime, a poisoned tool or an injected agent has to reach a destination to exfiltrate data or receive instructions. FireAI, developed by HisnLabs, is an outbound firewall for macOS. It runs as an Apple Network Extension content filter, identifies each app by its code signature and can allow, block or ask for each destination, with [rules per app, domain or address](https://hisnlabs.com/en/docs/per-app-rules). Applied to AI work on a Mac, this means an inference server, a coding agent or an MCP client can be held to the destinations its task needs, and a new destination triggers a decision. The [filter documentation](https://hisnlabs.com/en/docs/how-the-network-filter-works) describes what FireAI sees: app identity, remote host or address, port and protocol. FireAI does not scan model files, verify signatures, red-team a model or classify prompts. It does not read the contents of encrypted connections, so it cannot tell a legitimate request from an injected one when both go to an allowed destination. It complements the tools above and does not replace any of them. ## Comparison | Tool | Job | Maintainer | Licence | Status seen on 30 Sep 2026 | | --- | --- | --- | --- | --- | | ModelScan | Model file scanning | Protect AI | Apache-2.0 | Active, v0.8.8 | | picklescan | Pickle scanning | mmaitre314 | MIT | Active, v1.0.5 | | fickling | Pickle analysis | Trail of Bits | LGPL-3.0 | Active, v0.1.12 | | safetensors | Safe weight format | Hugging Face | Apache-2.0 | Active | | model-transparency | Model signing | Sigstore | Apache-2.0 | Active, v1.1.1 | | CycloneDX | ML-BOM specification | OWASP, Ecma TC54 | Apache-2.0 (schemas) | Active, 1.7 | | Instructional Fingerprinting | Model fingerprint (research) | Paper authors | MIT | Last update July 2024 | | MarkLLM | Text watermarking (research) | THU-BPM | Apache-2.0 | Active | | garak | Vulnerability scanning | NVIDIA | Apache-2.0 | Active, v0.17.0 | | PyRIT | Red-team framework | Microsoft | MIT | Active, v1.1.0 | | promptfoo | Evals and red teaming | Promptfoo, part of OpenAI | MIT | Active, 0.123.1 | | Giskard v3 | Agent tests and scan | Giskard-AI | Apache-2.0 | Active; v2 unmaintained | | Llama Guard, Prompt Guard | Guard models | Meta | Llama Community Licences | Models dated April 2025 | | NeMo Guardrails | Programmable rails | NVIDIA | Apache-2.0 | Active, v0.24.1 | | Rebuff, LLM Guard | Injection detection | Protect AI | Apache-2.0, MIT | Archived | | Agent Scan | Agent and MCP scanning | Snyk | Apache-2.0 | Active, v0.6.8 | | MCP Scanner | MCP server scanning | Cisco AI Defense | Apache-2.0 | Active, 4.8.4 | | FireAI | Per-app egress control on macOS | HisnLabs | Commercial | Does not scan models | *Sources: each project’s repository or documentation, read on 30 September 2026. Status reflects the repository page, not a quality judgement.* ## Limitations - No single tool covers the six jobs. A clean file scan says nothing about a model’s behaviour, a signature says nothing about safety and a red-team scan says nothing about a file’s contents. - Scanners and guards are detectors. They can miss novel techniques, as the ModelScan, picklescan and Rebuff documentation each acknowledge, and their coverage changes as attacks change. - Maintenance is uneven. Two detector projects listed here are archived, one tool has changed owner, one has an individual maintainer and one research repository has not changed since 2024. Check the state of a project before building on it. - This roundup was limited to open tools with primary sources that could be read. It excludes commercial platforms and does not compare detection results, because the sources give no comparable figures. - Licences were read from repository pages and licence tables. Verify them before redistribution, in particular the Llama Community Licences and the LGPL terms of fickling. - The network layer sees connections, not meaning. An allowed destination can still receive data from a compromised model runtime. ## How FireAI and HisnLabs fit in Scan the model, sign the model, test the model. Then decide where its app may connect. FireAI is HisnLabs’ own product: an on-device AI firewall for Mac. It shows every connection your apps make, in plain language, and lets you decide what leaves your Mac — its AI runs locally, so your traffic is never sent to us or anyone else. HisnLabs’ security research team is the group that keeps that decision-making accurate: cataloguing which domains are ordinary telemetry versus a real product, tracking the country and network behind a connection, and training the on-device model (its FireAI Pilot feature) on real traffic patterns, all without any of it leaving your Mac. You can read the technical decisions behind it, or try FireAI for 17 days, at [FireAI, by HisnLabs](https://hisnlabs.com/en/download). ## Sources - [GitHub: protectai/modelscan (Apache-2.0)](https://github.com/protectai/modelscan) - [GitHub: mmaitre314/picklescan (MIT)](https://github.com/mmaitre314/picklescan) - [GitHub: trailofbits/fickling (LGPL-3.0)](https://github.com/trailofbits/fickling) - [Hugging Face Hub docs: Pickle scanning](https://huggingface.co/docs/hub/security-pickle) - [Hugging Face Hub docs: Malware scanning](https://huggingface.co/docs/hub/security-malware) - [Hugging Face Hub docs: Third-party scanner, Protect AI](https://huggingface.co/docs/hub/security-protectai) - [GitHub: huggingface/safetensors (Apache-2.0)](https://github.com/huggingface/safetensors) - [GitHub: sigstore/model-transparency (Apache-2.0)](https://github.com/sigstore/model-transparency) - [GitHub: CycloneDX/specification (schemas Apache-2.0, ML-BOM)](https://github.com/CycloneDX/specification) - [SPDX: AI profile](https://spdx.dev/learn/areas-of-interest/ai/) - [GitHub: cnut1648/Model-Fingerprint, Instructional Fingerprinting (MIT)](https://github.com/cnut1648/Model-Fingerprint) - [arXiv 2401.12255: Instructional Fingerprinting of Large Language Models](https://arxiv.org/abs/2401.12255) - [GitHub: THU-BPM/MarkLLM (Apache-2.0)](https://github.com/THU-BPM/MarkLLM) - [GitHub: NVIDIA/garak (Apache-2.0)](https://github.com/NVIDIA/garak) - [GitHub: microsoft/PyRIT (MIT)](https://github.com/microsoft/PyRIT) - [GitHub: promptfoo/promptfoo (MIT)](https://github.com/promptfoo/promptfoo) - [GitHub: Giskard-AI/giskard-oss (Apache-2.0)](https://github.com/Giskard-AI/giskard-oss) - [GitHub: meta-llama/PurpleLlama (Llama Guard, Prompt Guard, LlamaFirewall)](https://github.com/meta-llama/PurpleLlama) - [GitHub: NVIDIA-NeMo/Guardrails (Apache-2.0)](https://github.com/NVIDIA-NeMo/Guardrails) - [GitHub: protectai/rebuff (archived)](https://github.com/protectai/rebuff) - [GitHub: protectai/llm-guard (archived)](https://github.com/protectai/llm-guard) - [GitHub: snyk/agent-scan, formerly MCP-Scan (Apache-2.0)](https://github.com/snyk/agent-scan) - [GitHub: cisco-ai-defense/mcp-scanner (Apache-2.0)](https://github.com/cisco-ai-defense/mcp-scanner) - [OWASP Top 10 for Large Language Model Applications](https://owasp.org/www-project-top-10-for-large-language-model-applications/) - [FireAI docs: How FireAI watches your Mac’s connections](https://hisnlabs.com/en/docs/how-the-network-filter-works) - [FireAI docs: Rules](https://hisnlabs.com/en/docs/per-app-rules)