Skip to content
farzad.fm
AI & Automation

NVIDIA Absorbs Groq in a $20 Billion Move as the Chip War Shifts From Training to Inference

With inference already running two-thirds of the world’s AI compute, NVIDIA’s Christmas Eve licensing deal for Groq’s patents, software, and top engineers quietly removed its fastest challenger while AMD, Cerebras, and Tesla scatter across entirely different mountains.

With inference already running two-thirds of the world's AI compute, NVIDIA's Christmas Eve licensing deal for Groq's patents, software, and top engineers quietly removed its fastest challenger while AMD, Cerebras, and Tesla scatter across entirely different mountains.

Everyone keeps asking who beats NVIDIA. It's the wrong question, and asking it will cost you money. The company that trains almost every AI model on the planet blew past $5 trillion in 2026, larger than the economy of most countries, selling what amounts to a very specialized calculator. But the real story sits underneath that number: the money in AI is migrating from building models to running them, and whoever collects the toll on that second job walks away with the majority of the cash flow of the entire AI age. NVIDIA understands this better than anyone, which is why its most important recent move wasn't a chip at all. It was a checkbook.

Key Takeaways

  • NVIDIA's data center business booked roughly $75 billion in a single quarter ending April 2026, up around 92% year over year, at a gross margin near 75%. Most hardware companies would kill for 30.
  • Inference, the act of actually running a model, is already about two-thirds of all AI compute in 2026 and still climbing toward 70 to 80%. Training, the part everyone obsesses over, is becoming the smaller business.
  • On Christmas Eve of 2025, NVIDIA paid roughly $20 billion for a perpetual, non-exclusive license to Groq's patents and software stack, then hired founder Jonathan Ross, president Sunny Madra, and about 80% of Groq's people. It never bought the company, which is exactly the point.
  • CUDA, NVIDIA's software layer, has been under construction since 2006. Eighteen years of accumulated tooling is the moat, not the silicon.
  • Groq's LPU chip pushes around 750 words per second with a time-to-first-token near 18 milliseconds, versus 200 to 400 milliseconds on a typical cloud GPU. The catch: each chip holds only ~230 megabytes of memory against 80 gigabytes on an NVIDIA H100.
  • Cerebras builds a single chip the size of a dinner plate, four trillion transistors, 900,000 cores, roughly 57 times the area of a normal GPU, and just went public with an OpenAI backlog reportedly worth tens of billions.
  • Tesla's AI5 taped out in April 2026 and ramps to volume in 2027, but Tesla sells none of it. The chips exist to run its cars, Optimus robots, and SpaceX satellites.
  • Google's TPU, Amazon's Trainium, and Microsoft's Maia are all custom inference chips built to stop paying NVIDIA's tax inside their own data centers.

The Two Jobs Nobody Separates, and Why It Decides Everything

There are two completely different things an AI chip can do, and confusing them is how people misread this entire war. The first is training. That's where you build the model from scratch, feeding it basically the whole internet until it learns the patterns of human language. It's brutal. Months of tens of thousands of chips running flat out, burning enough electricity to power a small city. But you do it once. Think of it as writing a cookbook. Slow, painful, expensive, and then it's done.

The second job is inference. That's what happens every single time you actually use the thing. You type a question, hit enter, and the model runs. That's cooking the dish from the recipe. You write the recipe once and you cook the dish a billion times a day, forever. Every chatbot reply, every generated image, every line of code, all night, never stopping.

Here's what that means for the money. For years the boom was about training, and that was NVIDIA's home turf. Now billions of people use these tools daily, and spending is shifting hard toward the cooking. Inference is already about two-thirds of global AI compute. The recipe is a one-time cost. The cooking is the forever money. So the war stops being about who has the most powerful chip and becomes about who owns the toll booth.

CUDA Is the Real Moat, and It's Eighteen Years Deep

You'd think someone could just build a faster training chip and take the crown. They can't, and the reason has nothing to do with silicon. NVIDIA's weapon is CUDA, the software layer that lets programmers actually use its chips. The company has been building it since 2006. In those eighteen years CUDA became the road network the entire AI world drives on.

Every AI engineer alive learned on it. All the tricks, the tools, the little secrets are built for CUDA. It's taught in universities and baked into how people think, the way everyone already knows how to use a phone. So a competitor can show up with a chip that's technically faster and it doesn't matter, because nobody knows how to drive on their roads. It's the QWERTY problem. Your keyboard layout might be objectively better, but the whole planet already learned to type. To beat NVIDIA head-on you wouldn't just build a better chip. You'd have to rebuild eighteen years of software and retrain every engineer on Earth.

AMD Is the Only One Actually Climbing NVIDIA's Mountain

Almost everyone looked at that wall and walked away. AMD didn't. It makes the same kind of general-purpose chips and made the simplest bet in the business: build a better, cheaper version of NVIDIA's own product and out-muscle it on price and performance. The hardware is genuinely close. In one independent 2026 benchmark, AMD's top chip landed within a few percent of NVIDIA's best on exactly the inference work that's becoming the whole game, and its next generation packs more high-speed memory than NVIDIA's current flagship even has.

So why isn't it a fair fight? The software. AMD's answer to CUDA, called ROCm, spent years being buggy and painful. It's gotten dramatically better, but it's still the wall AMD keeps slamming into. That leaves AMD in a strange spot, hardware almost there and software lagging behind. If NVIDIA's moat ever cracks, it cracks here first. The day AMD starts taking real market share is the day we learn the king is beatable. I'd give AMD good odds of staying a strong, distant number two, mostly because the entire industry is desperate for a second supplier so it isn't 100% hostage to one company's pricing.

Groq Bet the Company on Pure Speed, Then Got Absorbed

Jonathan Ross saw NVIDIA's weakness from the inside. At Google he built the first TPU, the custom chip that let Google run AI without leaning entirely on NVIDIA, and it started almost as a side project before ending up in Google data centers worldwide. He knew training was unbeatable and inference was exposed. So in 2016 he left and founded Groq to attack that vulnerability directly.

Groq's LPU is stupid fast. Around 750 words per second, with the AI starting to answer in about 18 milliseconds versus 200 to 400 on a normal cloud GPU. It does this by putting the model's memory directly on the chip in ultra-fast SRAM, instead of running to slower memory for every ingredient. Keep the whole cookbook open on the counter with everything in arm's reach. That's the difference. The problem is that on-chip memory is tiny, about 230 megabytes per chip against 80 gigabytes on an H100, so running one big model takes hundreds of chips wired together. Blazing fast, but not cheap to scale. Then, on Christmas Eve of 2025, NVIDIA paid roughly $20 billion for a perpetual license to Groq's patents and software and hired Ross, his president, and about 80% of the team. It didn't buy the company, because buying it invites antitrust regulators. It licensed the technology and hired the humans. Groq is basically NVIDIA now.

Cerebras Built the Craziest Chip in the World, and It's Riding on Two Customers

Every other company slices a silicon wafer into hundreds of small chips. Cerebras keeps the whole wafer. One chip the size of a dinner plate, four trillion transistors, 900,000 cores, roughly 57 times the area of a normal GPU. It sounds insane until you understand the problem it solves. When you spread a giant model across thousands of separate chips, they waste enormous time just talking to each other, shuffling data back and forth. That's the memory wall. Cerebras keeps everything on one massive slab, so there's almost no shuffling.

The technology is real. The business risk is concentration. For years most of Cerebras revenue traced back to a single customer, a UAE group called G42, and now a huge chunk of its future rides on an OpenAI inference deal with a backlog reportedly worth tens of billions. Impressive technology with one or two whales is a fragile thing. If one of them swims away, the whole story changes overnight. Now that Cerebras is public, the number to watch is whether it actually executes that OpenAI deal and wins customers beyond the two it leans on today.

Tesla Is Building a Private Mountain Just for Itself

Tesla makes its own AI chips too, from the Dojo project through its current AI4 generation, with AI5 and more coming fast. But Tesla sells none of them. They exist to make its cars drive themselves, its Optimus robots think, and eventually to power SpaceX satellites in orbit. It's the Apple move, building your own brain for your own body instead of buying it from Intel. AI5 taped out, meaning the design was finalized and sent to manufacturing, in April 2026, and ramps to real volume in 2027.

For a Tesla investor, the chip isn't a standalone reason to own the stock. It's one piece of the FSD, Optimus, and satellite thesis. I'll be honest about the risk here, because this is the company that built Dojo, disbanded the team, then brought it back, run by a founder with a long history of missing timelines. So watch AI5 production yield and schedule through 2027. But if it works, controlling its entire chip cost structure would hand Tesla a margin position almost nobody else can touch.

The Cloud Giants Are Quietly Building the Toll Booth Themselves

Here's the strongest argument against NVIDIA, and it's not a startup. It's the customers. Those beautiful 75% margins depend on training being the main event. As inference takes over, the giant cloud companies have every reason to build their own cheaper chips to do it in-house. Google has the TPU. Amazon has Trainium. Microsoft has Maia. Each one is designed to run inference inside their own data centers without paying NVIDIA a cent of tax.

And the chain feeds itself. More people use AI, which means more cooking, which means more demand that never turns off, and the inference share keeps climbing toward 70 or 80%. Whoever collects the toll on all that cooking owns the majority of the cash flow of the AI age. That's why NVIDIA's Groq move was so sharp: it wasn't defending training, it was buying its way into the lane it had neglected.

How I'd Actually Rank These Bets

I think NVIDIA keeps the crown clean through this inference era, helped by the Ruben chip ramping in late 2026 and by the fact that it just bought one of the best inference teams alive. The thing I'd worry about is the margin. Seventy-five percent in an increasingly crowded field feels unsustainable over the long haul, but calling the exact moment that reign ends is a fool's errand while AI is still booming.

The honest summary is that these companies are fighting on different battlefields, and knowing which lane each chip is really in is the entire thesis. AMD grinds up the same mountain and probably stays a solid number two. Cerebras owns the giant-model lane but carries real customer-concentration risk. Tesla is off building a private brain for its own machines. And Groq, the one that scared NVIDIA most, is now part of NVIDIA. Stop asking who beats NVIDIA. Start asking who owns inference. That's where the next decade of money actually lives.