Skip to content
All exclusives

This Is the New Oil and Most People Have No Idea: Why Tokens Become the Currency of AI

AI & Automation

There is a currency that already controls the future of every industry on the planet, and almost nobody knows it exists. Not Bitcoin. Not the dollar. Not your favorite crypto ticker. It is called a token.

Every time you ask ChatGPT a question, every time Claude writes code, every time Gemini digests a document, you spend this currency. You just do not see the bill. Someone else is paying. That invisible transaction explains why companies committed over $600 billion in infrastructure in a single year, why Elon Musk spent about $18 billion packing three Memphis buildings with half a million chips, why SpaceX filed to put a million AI compute satellites in orbit, and why OpenClaw ripped to 157,000 GitHub stars in weeks. Every headline snaps back to tokens.

I am Farzad. I have tracked this for fourteen years. Once you get tokens, you get the AI economy.

What a token actually is

An AI model does not read words the way you do. It chops text into tokens — not exactly a word, not exactly a character. Roughly four characters. "Hello" is one token. "Understanding" often splits into two. Rule of thumb: about 750 words equals roughly 1,000 tokens.

A short email might burn 2,000 to 3,000 tokens. A simple Q&A: 200 to 500. Behind every token sits brutal hardware. Specialized chips run $25,000 to $40,000 each. You need thousands humming at once. A single server with eight Nvidia H100s costs $200,000 to $400,000. Each chip can pull up to 700 watts. Add cooling, networking, storage, buildings, power — then multiply by hundreds of thousands of chips. Ask what to cook for dinner and you tap a resource that took billions of dollars of infrastructure to create.

Input vs output

Input tokens are your prompt and context. The model chews them in parallel. Efficient. Output tokens are the answer, generated one at a time. Every output token needs a full pass through billions of parameters. The bottleneck is memory bandwidth, not raw FLOPs. That is why output costs three to five times more than input at every major provider — and why RAM prices went vertical.

When GPT-4.5 first shipped: about $75 per million input tokens and $150 per million output. Cheap models like Gemini Flash-Lite sit around 8 cents per million input. That is a 900x spread. History trivia does not need a frontier brain. Building an app does.

Concrete example: upload about a million tokens of YouTube transcripts and ask a top-tier model for roughly a million tokens of book prose — call it ten 250-page books. At GPT-4.5-era prices that is roughly $75 in and $150 out, about $225 before edits. Absurdly cheap next to a human writing team. Intelligence is becoming a metered commodity.

The fastest cost collapse in tech history

November 2022: GPT-3 API at about $20 per million tokens. March 2023: GPT-4 at $30 input / $60 output. Then the cliff. November 2023: GPT-4 Turbo at $10 / $30. May 2024: GPT-4o at $5 / $15. July 2024: 4o Mini at 15 cents / 60 cents — about 200x cheaper on input than original GPT-4. By October 2024 you could get GPT-3.5-level performance for 7 cents per million. From $20 to 7 cents. A 280x collapse in under two years.

After January 2024, median declines hit about 200x per year. Peak rates hit about 900x. Sam Altman has said AI usage costs fall roughly 10x every twelve months. Solar followed Wright's Law for four decades. AI tokens crushed more cost in eighteen months than solar did in forty years. Nothing else comes close.

Jevons paradox walked in

Cheaper should mean lower bills. It does not. In 1865, William Stanley Jevons watched Watt make steam engines far more efficient and expected England to burn less coal. England burned more. Efficiency unlocked factories, machines, and industries that did not exist before. "It is a confusion of ideas to suppose that using fuel more efficiently means you'll consume less of it," Jevons wrote. "The very opposite is the truth."

Apply that to tokens. Prices fell 280x. Enterprise AI spending still jumped from $11.5 billion in 2024 to $37 billion in 2025 — a 320% surge. When GPT-4 became 4o and got about 100x cheaper, usage went up about 1,000x. Average monthly AI budget per organization in 2025 hit about $85,000, up 36%. Forty-five percent of orgs now spend over $100,000 a month — double in a year. Eighty-three percent of AI leaders worry about cost and keep spending anyway.

Every price cut makes ten new things viable. Software that never got greenlit. Contracts never reviewed by a model. Medical research nobody would fund. Box CEO Aaron Levie: the vast majority of future AI tokens will be spent on things we do not even do today as workers. Satya Nadella literally said "Jevons paradox strikes again." A company gets a 50% cut, cheers for five minutes, finds ten new uses. The bill goes up. The coal paradox ran 150 years. The AI token paradox is in year two.

Agents pour gasoline on it

A normal chat burns a few hundred tokens. An agent thinks, acts, observes, rethinks, and loops — for hours, sometimes overnight. A coding agent fixing a bug can chew 50,000 to 500,000 tokens in one session. Multi-step research tied to building software: 500,000 to 2 million, easy. Overnight research runs have hit 2 million on a single step.

Researchers studying coding agents found a 10x cost variance between runs on the same task. A huge chunk is input: every loop the agent rereads growing history. A 10-step loop can 50x a single call. Claude Code Max users collectively burned about 10 billion tokens in a month. Developers report burning $20 a day when they budgeted $20 a month. Reasoning models make it worse — seven "thinking" tokens versus 255 for the same visible answer. Higher quality. Higher burn.

OpenAI's GPT-5.3 Codex and Claude Opus 4.6 lean on context compaction so long sessions waste fewer tokens. Efficiency rises. Jevons answers again: people point agents at more work. OpenClaw went from nothing to 157,000 GitHub stars in weeks, runs on your machine, and already threw a developer conference. One linked agent spun up Multibook — a social network where AI agents post and argue. Roughly 150,000 agents generating content. Count the tokens.

Where you put the compute

Hyperscalers spent more than $260 billion CapEx in 2024, then climbed toward $450 billion. Amazon alone aimed near $125 billion. Microsoft dropped $35 billion in a single quarter, up 74% year over year. For 2026 the total easily crosses $600 billion, with about 75% — roughly $450 billion — going straight into AI. Goldman Sachs sees about $1.15 trillion in hyperscaler CapEx from 2025 to 2027.

Stargate targets up to $500 billion over four years and nearly 7 GW of power. xAI's Colossus in Memphis: over 500,000 GPUs across three buildings, about 2 GW, roughly $18 billion, build timelines measured in weeks. Target: a million GPUs.

In 2026, money spent running models — inference — surpassed training for the first time. Inference is already about 55% of AI cloud spend and headed toward 70–80% by 2030. The market alone is projected from about $106 billion toward $255 billion. The industry moved from building factories to running them.

When land, power, cooling, and neighbors run out? Space. On January 31, 2026, SpaceX filed with the FCC for up to one million low-Earth-orbit satellites — not for internet, for orbital data centers — while folding xAI under the same roof. Sun-synchronous orbit: solar on one face, radiate heat into cold space on the other. Their claim: lowest-cost AI compute will be in space within a few years. Timing will slip — Elon time always does. Treat it as a five-to-ten-year bet. The direction still matters.

Stack the optimizations. Demand still wins.

Quantization: drop 16-bit to 8-bit or 4-bit, lose maybe 1–2% accuracy, cut up to ~70% cost. Speculative decoding: 2–3x faster with identical outputs. Prompt caching: Anthropic offers up to 90% off cached prompts; one developer went from $720 a month to $72. Stack the tricks and you can see 100x cuts versus naive serving. Cerebras signed a $10 billion OpenAI deal for 750 MW of inference through 2028. None of that flattens the curve. Cheaper tokens create new demand that outruns the savings.

Who wins

Follow the tokens. Winners: whoever produces the most intelligence per dollar — Nvidia, Cerebras, Groq, the big clouds, and eventually SpaceX if orbital compute lands. Usage platforms see their market expand with every price cut as long as quality holds. Data centers, power, and cooling ride multi-gigawatt demand. Losers: thin wrapper startups that slap a UI on someone else's API. When the wholesale commodity deflates 200x a year, your retail margin evaporates. Seat-license software is next. When one agent replaces three analysts, selling three seats is a joke. By 2028, pure seat-based pricing is on track to be obsolete for AI-enabled workflows.

There is no AI bubble. There is an AI explosion. Tokens are the atomic unit. Costs collapsed 280x. Spending exploded anyway. Agents multiply demand another 10–100x. Hyperscalers are spending like there is no tomorrow because, for anyone who fails to build this stack, there is not one. The coal paradox ran 150 years. You are in year two of the token cycle. Follow the tokens.

Check the video here.

Digest

Prefer the daily pulse?

Short, sharp breakdowns of what actually moved — every day.