Skip to content
All exclusives

It's Not NVIDIA You Should Be Watching: Why Inference and Power Shift the Race

AI & Automation

Nvidia is the most valuable company in history. It blew past $5 trillion in 2026 — larger than the GDP of most countries — selling what is basically a specialized calculator for AI. The story you keep hearing is that every other chip company is racing to knock it off that mountain. Some are. The more interesting ones looked at the mountain and walked the other way.

That is the puzzle. If you cannot beat Nvidia head-on, why are these companies worth billions? And which lane is a real business versus a trap?

Two jobs, one forever market

Intelligence has always been scarce. For most of human history the bottleneck on getting things done was access to smart, capable thinking. These chips are factories for manufacturing that intelligence at a scale we have never had.

An AI chip does two completely different jobs.

Training is building the model. You feed something like ChatGPT basically the entire internet — books, websites, conversations — until it learns the patterns of language. That process is brutal. Months of work. Tens of thousands of chips running flat out, burning enough electricity to power a small city. You do it once. Think of writing a cookbook. Slow, expensive, painful. When you finish, the book exists.

Inference is every time you actually use the AI. You type a question, hit enter, and the model runs. That is cooking the dish from the cookbook. You write the recipe once. You cook the dish a billion times. Every question, image, line of code, chatbot reply — all day, all night, forever. The recipe is a one-time cost. The cooking never stops.

For years the AI boom was about training. Biggest model. Most horsepower. That was Nvidia's home turf. As billions of people start using these systems every day, spend is shifting hard toward inference. As of 2026, inference is already roughly two-thirds of all AI computing on the planet, and it is still climbing toward 70% and 80%. That is the forever money.

Why Nvidia's real weapon is not the chip

So why does nobody just build a better training chip and take the crown? Because Nvidia's real weapon is not silicon. It is CUDA — Compute Unified Device Architecture — the software layer that sits on top of the chips and lets programmers actually use them.

Nvidia has been building CUDA since 2006. Roughly 18 years. In that time it became the road network the entire AI world drives on. Every AI engineer learned on it. The tools, tricks, and course materials are built for it. It is taught in universities. It is how people think. Showing up with a faster chip and no CUDA is like inventing a better keyboard layout than QWERTY. Yours might be objectively better. Everybody already types on QWERTY.

To beat Nvidia head-on you would have to rebuild 18 years of software and retrain every engineer alive. Almost every smart player figured that out. Almost.

Five mountains, not one race

These companies are not all climbing the same mountain. Most are on completely different ones.

AMD is the one that looked at Nvidia's mountain and started climbing anyway. It is the only real number two in AI chips, and the bet is simple: build a better, cheaper version of Nvidia's product and outmuscle them on price and performance. The hardware is genuinely good. In one big independent benchmark in 2026, AMD's top chip came within a few percent of Nvidia's best on exactly the inference work that is becoming the whole game. Their next generation packs in more high-speed memory than Nvidia's current flagship, which matters when you run giant models.

On paper AMD is right there. In practice, one word: CUDA. AMD has spent years building ROCm, its own software stack. For a long time it was buggy and painful. It has gotten dramatically better. AMD lives in a strange place where the hardware is almost there and the software is the wall they keep slamming into. If Nvidia's moat ever cracks, it cracks here first. The day AMD takes real market share is the day we learn the king is finally beatable.

Groq (with a Q) was founded in 2016 by Jonathan Ross, the Google engineer who built Google's first TPU almost as a side project. He understood exactly where Nvidia was unbeatable — training — and where it was vulnerable — inference. His whole bet was the fastest inference chip in the world.

Groq's chip is an LPU, a language processing unit. Time to first token sits around 18 milliseconds. On a normal GPU cloud setup that wait is often 200 to 400 milliseconds. Ten to twenty times faster at responding. They did it by putting model memory directly on the chip as ultra-fast SRAM instead of constantly fetching from slower memory — the whole cookbook open on the counter versus walking to the pantry for every ingredient.

That on-chip memory is tiny. Each Groq chip holds about 230 megabytes. An Nvidia H100 holds 80 gigabytes. So running one big model means hundreds of Groq chips wired together. Blazing fast. Not cheap or simple to scale. But using a giant Nvidia GPU for quick-response work is like booking a freight train to deliver one pizza. Wrong tool. That gap is what Groq drove into.

Cerebras kept the whole wafer. Every other chip company slices a silicon wafer into hundreds of little chips. Cerebras builds one chip the size of a dinner plate — 4 trillion transistors, 900,000 cores, something like 57 times the size of a normal Nvidia GPU. The largest computer chip ever built by a wide margin.

Why? The memory wall. Spread a giant model across thousands of separate chips and they spend a ridiculous amount of time talking to each other, shuffling data. Keep it on one massive slab and there is almost no shuffling. Built for enormous models — training and inference at serious scale. The company is earlier in its history and recently went public. Watch whether it executes the enormous OpenAI inference deal with a backlog reportedly worth tens of billions, and whether it wins customers beyond the one or two it leans on. For years most revenue traced back to G42 in the UAE. Wafer-scale tech is real. Real tech with one or two customers is tough. If a whale swims away, the story changes overnight.

Tesla builds its own AI chips — Dojo, then AI4, with AI5 and others coming. It does not sell them. It builds them so its cars drive themselves, Optimus robots think, and upcoming orbital satellites from SpaceX have brains. Same play as Apple building chips for its own Macs instead of buying from Intel. AI5 taped out in April 2026 and ramps to real volume in 2027. Watch production yields and timelines through 2027 — and remember Tesla's history of missing deadlines. For a Tesla investor the chip is a piece of the FSD, Optimus, and inference-satellite thesis, not a standalone reason to own the stock. If it works, they control the entire cost structure.

The money moves. So does Nvidia.

Nvidia's data center business — just the AI chip part — did about $75 billion in revenue in the quarter that ended April 2026, up around 92% year over year. Gross margin near 75%. Most hardware companies would kill for 30. Blackwell is shipping in volume. Rubin lands in the second half of 2026. There is Nvidia, and then there is everybody else.

But those beautiful margins cannot last forever as cooking becomes the main event. The giant cloud companies are building their own cheaper chips to do inference inside their data centers without paying Nvidia's tax. Google has TPUs. Amazon has Trainium. Microsoft has Maia. Training is one business. Inference is another. Inference is becoming the main one.

Nvidia knows this. On Christmas Eve 2025 it struck a deal worth about $20 billion with Groq — not an outright buy. Nvidia paid roughly $20 billion for a perpetual, non-exclusive license to Groq's entire patent portfolio and software stack, and in the same move hired Jonathan Ross, president Sunny Madra, and somewhere around 80% of Groq's people. Acquire the company and regulators show up asking antitrust questions about the most powerful chip maker on Earth swallowing a rival. License the tech and hire the humans instead. Groq as an independent challenger is basically Nvidia now.

Stop asking who beats Nvidia

The mountain everybody stared at — training, recipe writing — is becoming the smaller business. The forever cash is in the cooking. The question stops being who has the most powerful chip. It becomes who collects the toll every single time an AI does anything for anyone.

Here is how the lanes look from here. Nvidia keeps the crown through this inference era, especially with Rubin and with one of the best inference teams on the planet now inside the building — but watch the margin squeeze. Seventy-plus percent margins in an increasingly competitive field feel unsustainable long term. Hard to say exactly when that rain ends while AI is still booming.

AMD will not topple Nvidia soon. Odds of staying a strong distant number two look pretty good. Real chips, real customers, and an industry that desperately wants a second supplier so it is not 100% hostage to Nvidia pricing. Watch whether ROCm gets good enough that engineers reach for it, and whether the big clouds buy AMD in real volume just to keep a backup alive.

Cerebras lives or dies on execution and customer concentration. Tesla's private mountain is a cost-structure play inside a bigger thesis. Groq's people and patents now sit under Nvidia's roof.

These companies are fighting on different battlefields — training versus inference versus building your own private brain. Knowing which lane each chip is actually in is the entire thesis. Stop asking who beats Nvidia. Start asking who owns inference.

Check the video here.

Digest

Prefer the daily pulse?

Short, sharp breakdowns of what actually moved — every day.