A token is roughly three-quarters of an English word, OpenAI’s own pricing page explained back in 2022, right after noting that a 35-token sentence costs a fraction of a cent. That framing barely survives contact with 2026: the smallest models now cost a fraction of a fraction of a cent per token, the largest cost fifty dollars per million output tokens, and three vendors — OpenAI, Anthropic and DeepSeek — have each cut, raised, split and re-split their prices several times over. Every figure below comes from a page we fetched: either the vendor’s current pricing page, or an Internet Archive snapshot of that same page from the date shown. Where we compute a ratio, both numbers it is built from are cited in the same paragraph.
OpenAI, 2022 to 2026
In November 2022, OpenAI’s pricing page listed four base models, Ada, Babbage, Curie and Davinci, each priced the same for every token regardless of whether it was prompt or output — there was no input/output split yet. Davinci, the most capable of the four, cost $0.02 per 1,000 tokens. OpenAI’s current deprecations page independently confirms that exact figure, $20.00 per million tokens, as the price at which text-davinci-003 — the specific model built on that tier — was billed until its 2024 shutdown. Four months later, GPT-4 launched at $0.03 per 1,000 prompt tokens and $0.06 per 1,000 completion tokens for its 8K-context version, doubling to $0.06 and $0.12 for the 32K version — the first time OpenAI’s public pricing charged output tokens more than input tokens. gpt-3.5-turbo launched the same month at a flat $0.002 per 1,000 tokens, again with no input/output split.
| Model | Input $/1M | Output $/1M | Priced as of | Source |
|---|---|---|---|---|
| text-davinci-003 | $20.00 | $20.00 | Nov 2022 | archived pricing page |
| gpt-3.5-turbo (launch) | $2.00 | $2.00 | Mar 2023 | archived pricing page |
| GPT-4, 8K context (launch) | $30.00 | $60.00 | Mar 2023 | archived pricing page |
| gpt-4o-2024-05-13 (launch) | $5.00 | $15.00 | May 2024 | OpenAI pricing docs |
| o1 (reasoning) | $15.00 | $60.00 | current listing | OpenAI pricing docs |
| gpt-4o (current price) | $2.50 | $10.00 | current | OpenAI pricing docs |
| gpt-6-luna (cheapest current) | $0.10 | $0.50 | current | OpenAI pricing docs |
| gpt-6-astra (current flagship) | $10.00 | $50.00 | current | OpenAI pricing docs |
Two comparisons, both built only from the rows above: text-davinci-003’s $20.00 input price against gpt-6-luna’s current $0.10 is a 200x drop over roughly four years, for OpenAI’s cheapest available tier at each point in time. But GPT-4’s original $30.00 input price against gpt-6-astra’s current $10.00 is only a 3x drop for a flagship-to-flagship comparison — the cheap tier fell far faster than the expensive one. gpt-4o itself was repriced downward after launch, from $5.00/$15.00 in May 2024 to $2.50/$10.00 today, a straightforward case of the same model getting cheaper without changing its capability at all.
Anthropic, 2023 to 2026
Anthropic’s earliest developer pricing sheet we could recover, an archived PDF from July 2023, priced Claude 2 — launched that same month — at $11.02 per million input tokens and $32.68 per million output tokens, and noted that Claude 1 remained available at the identical price. Claude Instant, the cheap tier at the time, was $1.63 input and $5.51 output. By early 2024, before Claude 3 shipped, an archived snapshot shows Claude 2.0 and 2.1 cut to $8.00/$24.00 and Instant cut to $0.80/$2.40 — modest reductions on the existing lineup. Claude 3, launched in March 2024, reset the bottom of the range entirely: Haiku came in at $0.25/$1.25, well below anything Anthropic had offered before, while Sonnet and Opus launched at $3.00/$15.00 and $15.00/$75.00. Claude 3.5 Sonnet, three months later, launched at the exact same $3.00/$15.00 as Claude 3 Sonnet — a case of price staying flat while the model itself improved.
| Model | Input $/1M | Output $/1M | Priced as of | Source |
|---|---|---|---|---|
| Claude 1 / Claude 2 (same price) | $11.02 | $32.68 | Jul 2023 | archived pricing PDF |
| Claude Instant | $1.63 | $5.51 | Jul 2023 | archived pricing PDF |
| Claude 2.0 / 2.1 (after price cut) | $8.00 | $24.00 | Feb 2024 | archived pricing PDF |
| Claude 3 Haiku (launch) | $0.25 | $1.25 | Mar 2024 | archived API page |
| Claude 3 Sonnet (launch) | $3.00 | $15.00 | Mar 2024 | archived API page |
| Claude 3 Opus (launch) | $15.00 | $75.00 | Mar 2024 | archived API page |
| Claude 3.5 Sonnet (launch) | $3.00 | $15.00 | Jun 2024 | archived pricing snapshot |
| Haiku 4.5 (current cheapest) | $1.00 | $5.00 | current | Claude pricing |
| Sonnet 5 (current) | $2.00 | $10.00 | current | Claude pricing |
| Fable 5.1 (current flagship) | $10.00 | $50.00 | current | Claude pricing |
Here the curve does not simply fall. Claude 3 Haiku launched at $0.25 input in March 2024; Anthropic’s current cheapest model, Haiku 4.5, lists at $1.00 — four times more expensive, per token, than the cheap tier of two years earlier. Read alone, that looks like AI got more expensive. Read next to the rest of this article, it looks like something else: Anthropic kept its lowest tier’s per-token price roughly flat to slightly rising while making that tier far more capable, and it also added a new, pricier flagship tier (Fable 5.1 at $10.00/$50.00) above what Opus used to be the top of. The lineup grew more tiers instead of just getting uniformly cheaper.
DeepSeek and the reasoning-price shock
DeepSeek entered this market already priced far below its two rivals. In May 2024, its archived pricing page shows deepseek-chat, backed by DeepSeek-V2, at a flat $0.14 input and $0.28 output — cheaper than anything OpenAI or Anthropic offered at the time, before either had shipped a token cheaper than $0.25. By December 2024, deepseek-chat had been upgraded in place to DeepSeek-V3, and DeepSeek’s pricing page introduced a cache-hit/cache-miss split: an archived snapshot shows the standard (post-promotion) rate at $0.07 cache-hit input, $0.27 cache-miss input and $1.10 output. Then, on 20 January 2025, DeepSeek released R1, a reasoning model, at $0.14 cache-hit input, $0.55 cache-miss input and $2.19 output — an archived snapshot from eight days later confirms that price, explicitly excluded from any promotional discount.
| Model | Input $/1M (cache hit / cache miss) | Output $/1M | Priced as of | Source |
|---|---|---|---|---|
| DeepSeek-V2 (deepseek-chat) | flat $0.14 | $0.28 | May 2024 | archived pricing page |
| DeepSeek-V3 (deepseek-chat, standard rate) | $0.07 / $0.27 | $1.10 | Dec 2024 | archived pricing page |
| DeepSeek-R1 (deepseek-reasoner, launch) | $0.14 / $0.55 | $2.19 | Jan 2025 | archived pricing page |
| deepseek-flash (current, off-peak) | $0.003 / $0.15 | $0.60 | current | DeepSeek pricing docs |
| deepseek-v4-pro (current, off-peak) | $0.022 / $0.66 | $1.98 | current | DeepSeek pricing docs |
R1’s launch is the single sharpest number in this entire article: $2.19 output against OpenAI’s o1, a comparable reasoning model, at $60.00 output — a 27x gap, both figures from the tables above. That comparison is exactly what made R1 a story beyond the usual developer audience in January 2025. It is also worth noticing what DeepSeek’s own numbers say about reasoning specifically: R1’s $2.19 output price is roughly double V3’s $1.10, inside the same vendor, at the same moment in time. Chain-of-thought reasoning does not just add a cost multiplier at the API level; DeepSeek’s own documentation notes that a reasoning model’s output token count includes every token of its reasoning trace, not just the final answer, so a reasoning query burns far more output tokens than a direct one even before the per-token price is applied. And DeepSeek’s current pricing adds a mechanism neither other vendor uses: peak-hour rates roughly double off-peak rates on the same model, a form of demand-based pricing closer to electricity billing than to a fixed price list.
The pattern the flat numbers hide
Put the three vendors’ cheapest-available-token figures on one timeline and the shape is not a smooth curve. OpenAI’s cheap tier fell from $20.00 (Nov 2022) toward $0.10 today, essentially monotonically. Anthropic’s cheap tier fell hard once, from $1.63 (2023, Claude Instant) to $0.25 (Mar 2024, Claude 3 Haiku), then rose back to $1.00 (current, Haiku 4.5) as that tier got more capable. DeepSeek entered already cheap and has stayed there, while adding time-of-day pricing on top. None of the three vendors’ curves look like each other, and none of them look like a simple exponential decline once you track a specific tier rather than "whatever the cheapest model is called this year."
Output costs more than input, and the ratio widened
At GPT-4’s March 2023 launch, output cost exactly twice input: $60.00 against $30.00. Look at current flagship and mid-tier models across all three vendors and the ratio has widened: gpt-6-astra is 5x ($50.00 against $10.00), Claude 3 Opus launched at 5x ($75.00 against $15.00) and stayed there through Haiku 4.5 (5x, $5.00 against $1.00), and DeepSeek-V3’s standard rate is roughly 4x ($1.10 against $0.27 cache-miss). The usual technical explanation is architectural rather than commercial: generating output tokens happens one at a time, autoregressively, while a prompt’s input tokens can be processed in parallel in a single forward pass, so a provider’s actual compute cost per output token is higher than per input token, and pricing has drifted toward reflecting that gap more accurately over time than GPT-4’s original, rounder 2x split did.
Cache and batch discounts
All three vendors now discount tokens the model has effectively already seen. OpenAI’s current pricing page lists gpt-6-astra’s cached input at $1.00 against a $10.00 standard input rate, a 10x discount for reusing an identical prompt prefix, and separately states that its Batch API processes requests within a 24-hour window at half the standard synchronous price across eligible models. Anthropic’s current pricing shows Fable 5.1’s prompt-caching read price at $0.25 against a $10.00 standard input rate — 40x cheaper — while a cache write costs $12.50, a premium over the $10.00 standard rate, since writing a new cache entry does real extra work the first time. DeepSeek’s cache-hit/cache-miss split, described above, is built into its base pricing rather than offered as a separate feature: R1’s $0.14 cache-hit rate against its $0.55 cache-miss rate is roughly a 4x difference on every single call, not just on requests that opt in. For any workload that repeats a long system prompt or a fixed set of tool definitions across many calls — a security triage pipeline classifying one log line after another against the same instructions is a clean example — these discounts change which vendor is cheapest more than the headline per-token price does.
The Jevons paradox: why cheaper tokens raised total spending
In 1865, the economist William Stanley Jevons published The Coal Question, arguing that Britain’s coal consumption rose, not fell, after James Watt’s more fuel-efficient steam engine made coal-powered work cheaper. As Wikipedia’s summary of the argument puts it, quoting Jevons directly:
It is a confusion of ideas to suppose that the economical use of fuel is equivalent to diminished consumption. The very contrary is the truth.
William Stanley Jevons, 1865, as quoted in Wikipedia: Jevons paradox
The argument, since generalised well beyond coal, is that making a resource more efficient to use lowers its effective price, and a lower price increases the quantity demanded by more than the efficiency gain saves — so total consumption rises. According to the same Wikipedia summary, this exact argument was widely invoked in early 2025 around DeepSeek’s cheap R1 model, with commentators including Microsoft CEO Satya Nadella and economist Erik Brynjolfsson cited discussing whether radically cheaper AI inference would shrink or grow the total resources spent on AI. We rely here on that secondary summary of named commentary rather than a primary economic study measuring the aggregate effect, and the claim that industry-wide AI spending is rising because of cheaper tokens — rather than in spite of a separate, unrelated demand boom — is a debated interpretation, not a settled statistic. What the price tables above do support directly is the precondition for the argument: the same query that cost tens of dollars in compute in 2022 can cost cents today, at every vendor, which is exactly the kind of price collapse that historically has preceded consumption rising rather than falling.
What this means for running AI locally instead of in the cloud
Every price in this article is a metered, per-call cost billed by a company that necessarily receives your prompt to answer it — that is what an API call is. A model that runs entirely on your own device has no per-token bill after you have paid for it or downloaded it once, and it has nothing to send anywhere, because there is no remote endpoint in the request path at all. That is not a claim that on-device models match the frontier tiers priced above; a model small enough to run comfortably on a laptop is not going to out-reason Fable 5.1 or gpt-6-astra. It is a different trade: a fixed cost and a firm privacy boundary, in exchange for whatever capability fits in the memory you actually have.
That trade-off is not abstract; it is the same one a small, narrowly scoped on-device model makes inside a piece of security software. FireAI, HisnLabs’ Mac firewall, ships an optional local model whose only job is reviewing a connection from an app it does not yet have a rule for — a task with a far smaller scope than anything priced in the tables above, and one where the model runs on the Mac itself rather than calling any of these APIs. The point of walking through OpenAI, Anthropic and DeepSeek’s price history in this much detail is not to sell that design; it is to make the trade concrete: cloud tokens bought a great deal more capability for a great deal less money between 2022 and 2026, and at every one of those price points, using them still meant sending your data to someone else’s server to get an answer back.
How FireAI and HisnLabs fit in
None of this makes FireAI a cost story — it is a firewall, not a rented model — but the same bet sits underneath its own on-device reviewer: no per-token bill and nothing sent off the Mac, because the review happens locally instead of being shipped to a server.
FireAI is HisnLabs’ own product: an on-device AI firewall for Mac. It shows every connection your apps make, in plain language, and lets you decide what leaves your Mac — its AI runs locally, so your traffic is never sent to us or anyone else. HisnLabs’ security research team is the group that keeps that decision-making accurate: cataloguing which domains are ordinary telemetry versus a real product, tracking the country and network behind a connection, and training the on-device model (its Autopilot feature) on real traffic patterns, all without any of it leaving your Mac.
You can read the technical decisions behind it, or try FireAI for 17 days, at FireAI, by HisnLabs.
