Skip to content
farzad.fm
AI & Automation

Why CPUs Are About To Be Sold Out Everywhere

Everyone is watching GPUs, but AI agents hand their searching, clicking and code runs to the CPU, the head chef of the data center, and with the CPU to GPU ratio moving from 1:8 to 1:4, I think the next chip crunch hits CPUs.

Loads from YouTube only after you press play.

Watch on YouTube

Nvidia became the largest company in the world by selling GPUs. So everyone who follows AI closely is obsessed with GPUs.

But the story is no longer just the GPU. The CPU, the chip in charge of executing actions, is quickly becoming critical. In the data centers that will power AI agents, I see the ratio of CPUs to GPUs heading toward one to one.

Nothing here is investment or financial advice. Watch the upload for the full walkthrough. This is the same case in writing.

The kitchen

Picture a data center as a kitchen. The GPUs are the line, rows of cooks chopping ingredients all at once. What they produce is intelligence: a decision on what to do next. The CPU is the head chef, who takes what the line hands over and executes on it.

Now add AI agents. An agent is an AI like Meta Muse, Grok Bot or ChatGPT dots that does a whole task for you, step after step. It takes the decision from the cooks and turns it into a finished meal. GPUs generate the intelligence. CPUs execute on it. Agents need both.

Jensen Huang, who runs Nvidia, estimated an agent needs 15 to 100 times the computing of a person using AI. That means a much longer line. But an agent does far more than math. It searches, clicks around, runs code, edits files and fills out forms. All of that goes to the head chef. That is where I see the next huge chip crunch hitting.

Quick question. While an agent works, who is stuck waiting, the line or the head chef? Hold that guess.

Why the line won round one

AI math is one giant pile of the same small multiplication, over and over. Picture a mountain of onions. One chef is too slow, so you put a whole line on it. That is a GPU, a chip packed with tiny, simple workers doing the same math at once. The head chef runs everything else. Your laptop runs on one too: your browser, your files, your email.

The line won AI's first wave because of training, the biggest onion pile there is. Training today's models took a gigantic number of GPUs. Whatever AI sits on your phone, the line chops through the math and writes your answer.

The GPU frenzy is real, and I expect it to run for a long time. But the GPU is one worker. The question is what changes when AI stops answering and starts doing, like replying to your email or filling out a form.

The dinner rush

A ChatGPT or Claude answer is one plate, one ticket. An agent is a Friday night dinner rush. Every ticket sends the head chef running to the pantry, the phone, the register and the angry customers.

Say you ask an agent to book a weekend trip. First it thinks and plans on the line. Then it acts. It opens a browser to check flights, searches your calendar and runs a little code to compare prices. Each of those is a tool, any program the agent calls for help. Then it waits for the answer. Then it checks. If a flight is too early, it goes around again. Think, act, wait, check. That is the loop.

The browser, the calendar lookup, the sandbox and deciding what is next are mostly head chef work. So back to your guess. In tool-heavy jobs, a lot of the weight sits on the head chef. One research team timed tool-heavy agent workloads on their test machines. In the most tool-heavy case, the head chef's tool work took up to 88% of the end-to-end time.

That costs real money. If the head chef falls behind, the line waits, and every idle second still burns electricity powering those GPUs. Picture hiring ten cooks by the hour and watching them sit around most of the shift. That restaurant goes bankrupt fast. Faster GPUs speed up generating intelligence. Executing it is CPU work.

The companies are saying it out loud

As Tom's Hardware reports it, Intel's finance chief says the ratio has gone from one CPU per eight GPUs to one per four. One head chef used to cover eight cooks. Now it is four. He says it could converge on one to one, a head chef for every cook.

Nvidia's finance chief says the rise of agents is speeding up demand for data center CPUs. So Nvidia is building them. In March it launched a CPU rack made for agents. Nvidia says one rack can sustain more than 22,500 separate CPU environments at once. Think of each one as a sealed room for an agent's code. Jensen's pitch: "We call it Vera. This is CPU for agents. All the CPUs of the past we built for humans. This CPU is built for agents."

Amazon's CEO, who runs one of Nvidia's biggest customers, says agent tool use runs mostly on CPUs instead of AI chips. Microsoft's CEO, as Arm's blog quotes him, says that for running agents, CPUs are just as important as GPUs. AMD CEO Lisa Su, whose company is one of the two biggest makers of server CPUs, expects the server CPU market to reach about $220 billion by 2030. She says agent sandboxes, the smallest piece today, have the fastest growth ahead, and she sees them becoming the largest.

These are companies with CPUs to sell, so stay skeptical. But running agents on my own local machines, it is obvious to me that CPUs become an extremely large bottleneck as more people use agents.

How you build a head chef

A head chef for a dinner rush needs cores, memory and power. A core is one of the head chef's hands, each working its own ticket. More cores, more sandboxes at once. Memory is the counter space where work in progress sits, and every sandbox needs its own spot. Then power. Lots of it.

Why do AI builders want CPUs near the GPUs? SemiAnalysis says that when models practice on coding and math, each round needs lots of CPUs to run code, check it and use tools. The CPUs are there to keep the GPUs busy and cut idle time.

Intel's CEO says that as AI moves into agentic systems, general-purpose server CPU density keeps increasing. In the quarter that ended in June, Intel's server chip volume rose 9% from a year earlier. Intel's filing says demand beat supply because of its own supply constraints.

Mercury Research says AMD now has about 1 in 3 of the server processors shipped, up from a year earlier. It also estimates servers built on Arm, a different family of CPU, just hit a record share. And the cloud companies build their own. Amazon's CEO calls its Graviton chip the strongest CPU chip. Arm's blog says Microsoft expected its Cobalt 200 racks in more than 25 data centers by the end of July. Nvidia says its Grace CPU brought in more than $5 billion over the last 12 months.

The price of your agent

My call: agents get cheaper from here. Part of the drop comes from intelligence getting far better and cheaper, so you need fewer cooks for the same work at higher quality. Go from 20 amateur cooks at $30 an hour to five professional cooks at $60 an hour. Each cook earns double. The total bill falls, and speed and quality climb. The rest comes from the head chef's side, because in tool-heavy jobs that is where the waiting is. More CPUs, less waiting.

For builders, the number that matters is cost per completed task, what one finished job costs you. Time every step of your loop and fix the slowest first. Sometimes it is the model. Sometimes it is the CPU.

Agents need more of both chips. More tickets need a longer line and more head chefs. And as models commoditize, getting cheaper and easier to reach, I see the money moving to the hardware that runs and executes them.

What to watch

Nvidia first. Its finance chief's preliminary expectation is that CPU revenue more than doubles next fiscal year. If Nvidia hits it, agents are pulling in a lot more CPUs.

Then supply. Intel expects industry-wide limits to last into next year. If the squeeze eases because supply catches up, CPUs get cheaper, and so do agents. If it eases because demand drops, throw this whole thesis in the trash.

Last, the CPU to GPU ratio. If it stops shifting toward more CPUs, agentic workloads are not catching on the way I expect.

Everyone is still counting cooks. Count the head chefs.

Watch

Prefer the tape?

Open the source video on YouTube if you want the full cut.

Digest

Prefer the daily pulse?

Short breakdowns of what actually moved, every day.