AI Tokens: The Invisible Currency Reshaping Every Industry
Why Costs Are Plummeting While Spending Explodes AI is transforming how businesses operate, but the real driver is an overlooked unit of computation that's getting dramatically cheaper—yet fueling unprecedented investments. This shift unlocks new capabilities, from automated r…
Why Costs Are Plummeting While Spending Explodes
AI is transforming how businesses operate, but the real driver is an overlooked unit of computation that's getting dramatically cheaper—yet fueling unprecedented investments. This shift unlocks new capabilities, from automated research to complex software builds, while infrastructure races to keep up.
Key Takeaways
- Tokens represent the core unit of AI processing, where input (prompts and context) costs less than output (responses), with prices varying 900 times between basic and advanced models.
- AI token costs have dropped 280 times in under two years, faster than any technology in history, enabling tasks like converting video transcripts into books for under $225.
- Despite falling prices, total AI spending surges due to expanded applications, with enterprise budgets rising 320% to $37 billion in 2025.
- AI agents amplify token consumption by 10 to 100 times through looped thinking and actions, making complex tasks like market analysis or app development feasible.
- Infrastructure investments hit $600 billion in 2026, shifting focus from training models to running them, with inference now over 50% of costs.
- Space-based computing emerges as a solution to earthly limits, with plans for up to 1 million satellites providing endless power and cooling.
- Optimizations like quantization and speculative decoding cut costs further, but new demands ensure spending keeps climbing.
- Winners include efficient token producers and infrastructure builders; losers are thin AI wrappers and seat-based software models.
The Fundamentals of Tokens in AI
Tokens break down text into manageable pieces for AI models, roughly equating to four characters each or about 750 words per 1,000 tokens. This system allows AI to handle queries efficiently, but every token requires significant computational power from specialized hardware.
Input tokens, which include user prompts and context, process in parallel for lower costs. Output tokens, generated sequentially, demand more resources due to repeated passes through the model's parameters, often costing three to five times as much. Memory bandwidth limits this process, explaining the global RAM shortage and price hikes.
Model pricing reflects capability levels. Advanced options charge $75 per million input tokens and $150 for outputs, while basic ones start at 8 cents per million inputs. This range suits varied tasks: simple queries use affordable models, while deep reasoning like software creation demands premium ones.
A practical example illustrates the economics. Transforming a library of video transcripts—around 1 million tokens—into 10 books costs about $75 for inputs and $150 for outputs, totaling $225. Additional steps like editing or fact-checking add more, but this remains far cheaper than hiring human teams for similar work.
The Unprecedented Collapse in Token Costs
Token prices have fallen at an extraordinary rate, outpacing historical declines in solar panels, batteries, or microchips. From late 2022, costs dropped from $20 per million tokens to 7 cents, a 280-fold reduction in under two years. Recent trends show accelerations up to 900 times per year.
This stems from rapid model iterations. Early versions charged $30 for inputs and $60 for outputs per million. Subsequent updates halved those figures multiple times, with mini models offering similar performance at fractions of the cost.
Competition and scaling drive this, following patterns where production doublings consistently lower expenses. In AI, 18 months achieved what took solar four decades, promising 10-fold annual drops.
How Falling Costs Drive Explosive Spending
Logic suggests cheaper tokens mean lower bills, but the opposite occurs. As prices drop, new uses emerge that were previously unaffordable, expanding overall consumption. This mirrors 19th-century steam engine improvements, where efficiency boosted coal use by enabling new industries.
In AI, a 280-fold price collapse coincided with enterprise spending jumping from $11.5 billion in 2024 to $37 billion in 2025—a 320% increase. Average monthly budgets reached $85,000, with nearly half of organizations exceeding $100,000. Over 80% express cost concerns, yet investments rise.
Each price cut unlocks applications like AI-reviewed contracts, medical research, or marketing automation. Leaders note that most future tokens will fund novel tasks, not replacements for existing ones. A 50% reduction leads to 10 new projects, pushing totals higher.
The Role of AI Agents in Amplifying Demand
AI agents take token usage to new levels by automating multi-step processes. Unlike simple chats consuming hundreds of tokens, agents loop through thinking, acting, and observing—potentially using 50,000 to 2 million tokens per session.
Tasks like summarizing documents and emailing results, or conducting industry research and building apps, demonstrate their scope. Variability is high; the same task can cost 10 times more across runs due to growing context in loops.
Reasoning models enhance accuracy by generating extended internal streams, using up to 255 tokens where basics need seven. This improves outputs but raises costs, akin to humans expending more effort for better results.
Open-source agents run locally, handling emails, calendars, and web interactions. One such system even spawned a social network for agents, with 150,000 participants generating content autonomously. Developer events highlight rapid adoption, with users consuming billions of tokens monthly.
New models focus on efficiency, like compressing context to maintain quality while cutting waste. Yet, improvements enable more complex uses, perpetuating the cycle of rising demand.
The Infrastructure Boom Fueling Token Production
Meeting this demand requires massive builds. Hyperscalers invested $260 billion in 2024, rising to $450 billion in 2025 and over $600 billion in 2026. About 75% targets AI, with projections reaching $1.15 trillion from 2025 to 2027.
Mega-projects include joint ventures planning $500 billion over four years for 7 gigawatts of power across multiple sites. Single installations pack 500,000 GPUs into facilities costing $18 billion, built in weeks instead of years, with goals of 1 million GPUs.
A key shift: inference (running models) now exceeds training costs, comprising 55% of spending and projected to hit 70-80% by 2030. The inference market alone grows from $106 billion to $255 billion in that timeframe, signaling widespread deployment.
Earthly constraints—power, land, cooling—push innovation skyward. Plans for 1 million satellites in low-Earth orbit provide constant solar power and free heat dissipation. Operating at 500-2,000 kilometers, they eliminate terrestrial costs, potentially adding 100 gigawatts annually. While timelines stretch to 5-10 years, the direction addresses scaling limits.
Optimizations Extending Token Efficiency
Beyond hardware, software tweaks reduce costs. Quantization compresses data precision for 70% savings with minimal accuracy loss. Speculative decoding uses small models to predict tokens, speeding processes two to three times without quality drops.
Batching and attention management boost throughput 23 times. Caching reuses prompt segments for 90% reductions. Combined, these yield 100-fold efficiencies over basic runs.
Chip advancements include massive single processors with 4 trillion transistors, running models 21 times faster than clusters. Deals secure gigawatts of compute through 2028. The market diverges into specialized training and inference chips, though some integrate both for versatility.
Winners and Losers in the Token Economy
Efficient token producers thrive: infrastructure giants, cloud platforms, and satellite providers expand markets with each price drop. Data centers, power, and cooling firms ride 7-gigawatt demand waves.
Thin interfaces over APIs struggle as margins shrink with deflating costs. Seat-based software fades; usage pricing dominates by 2028, as one agent replaces multiple roles.
This cycle—price drops enabling new demands—echoes historical shifts but accelerates. Year two of the token era promises sustained growth, rewarding those grasping its dynamics for investments and careers.
Related
Keep reading
AUG 12, 2026
This Is the New Oil and Most People Have No Idea: Why Tokens Become the Currency of AI
A single server with eight Nvidia H100 chips costs $200,000 to $400,000. GPT-4.5 launched at $75 per million input tokens and $150 per million output. By…
JAN 20, 2026
The AI Edge: Battlefield-Proven Tech Reshaping Global Power
Why Defense-Born AI Could Define the Next Decade of Innovation and Inequality AI built under extreme conditions isn't just surviving—it's thriving in ways that…
AUG 12, 2026
SpaceX Just Made the Internet 95% Cheaper: What the Market Is Missing
Starship just crushed the cost of putting bandwidth in orbit from $6.55 per Mbps down to about 30 cents. That 95% collapse is the number that decides whether…