Nvidia Pays $20 Billion for Groq's Team as the AI Chip War Shifts to Inference
By licensing Groq’s entire patent portfolio and hiring founder Jonathan Ross plus roughly 80% of its staff on Christmas Eve, the $5 trillion chip king is fortifying the one battlefield - inference - where in-house silicon from Google, Amazon, and Microsoft most threatens its 7…
By licensing Groq's entire patent portfolio and hiring founder Jonathan Ross plus roughly 80% of its staff on Christmas Eve, the $5 trillion chip king is fortifying the one battlefield - inference - where in-house silicon from Google, Amazon, and Microsoft most threatens its 75% margins.
Everyone keeps asking who beats Nvidia. I think that's the wrong question, and Nvidia's own behavior proves it. The company crossed $5 trillion in value in 2026 selling what amounts to a very specialized calculator for artificial intelligence, and the sharpest chip companies on Earth looked at that mountain and decided not to climb it. They went somewhere else. The reason is money: spending in AI is quietly migrating from building models to running them, and whoever collects the toll on that second job ends up owning the cash flow of the entire AI age. Nvidia knows it. That is why it just spent $20 billion to make sure it owns that side too.
Key Takeaways
- Nvidia became the most valuable company in history, crossing $5 trillion in 2026, on the back of a data center business that did about $75 billion in a single quarter - the one ending April 2026 - up roughly 92% year over year at close to 75% gross margin.
- Inference, the work of running a model every time someone uses it, is already about two-thirds of all AI computing in 2026 and is still climbing toward 70 to 80%, while training keeps shrinking as a share.
- Cuda, the software layer Nvidia has built since 2006, is the actual moat; about 18 years of tooling means every AI engineer alive already codes on Nvidia's roads.
- On Christmas Eve 2025, Nvidia paid roughly $20 billion for a perpetual license to Groq's patents and software stack and hired founder Jonathan Ross, president Sunny Madra, and around 80% of Groq's people, without technically acquiring the company.
- Groq's LPU spits out about 750 words per second with an 18 millisecond time to first token, versus 200 to 400 milliseconds on a standard cloud GPU, but each chip holds only 230 megabytes of memory against an Nvidia H100's 80 gigabytes.
- Cerebras builds a single chip the size of a dinner plate - 4 trillion transistors, 900,000 cores, about 57 times the size of a normal Nvidia GPU - and just went public riding a reported multi-billion-dollar OpenAI inference backlog.
- AMD is the only credible number two, landing within a few percent of Nvidia's best on inference benchmarks in 2026, but its ROCm software still can't match Cuda's 18-year head start.
- Tesla taped out its AI5 chip in April 2026 and ramps it in 2027, but it never sells the silicon - it feeds its own cars, Optimus robots, and SpaceX satellites, the way Apple builds chips for its own Macs.
- Google's TPU, Amazon's Trainium, and Microsoft's Maia are all in-house chips built to run inference inside their own data centers without paying Nvidia's tax, which is exactly what threatens those 75% margins.
The Two Jobs Hiding Inside Every AI Chip
An AI chip does one of two things. One job builds the model. You take something like ChatGPT and feed it basically the entire internet - every book, website, and typed-out conversation humans have ever produced - and let it grind on that until it learns the patterns of language. That is training. It takes months, tens of thousands of chips running flat out, and enough electricity to light up a small city. But you do it once. Think of it as writing a cookbook: brutal, slow, expensive, and then it's done.
The second job is inference. That's what happens every single time you use the model. You type a question, hit enter, and it answers. That's the AI cooking the dish from the recipe. You write the cookbook once. You cook the dish a billion times a day, forever, every image generated and every line of code spat out. The recipe is a one-time cost. The cooking never stops. And that difference is the whole story, because the money is moving from the writing to the cooking.
Cuda Is the Wall Every Rival Slams Into
People imagine a rival building a faster chip and taking the crown. It doesn't work like that. Nvidia's weapon is a piece of software called Cuda - Compute Unified Device Architecture - the layer that sits on the silicon and lets programmers actually use it. Nvidia has been building Cuda since 2006. That's roughly 18 years, and in that time it became the road network the entire AI world drives on. Every engineer learned on it. It's taught in universities. It's baked into how people think about the problem.
So when a competitor shows up with a chip that's technically faster, it barely matters, because nobody knows how to drive on its roads. It's the QWERTY problem. You can invent a better keyboard layout, and it might be objectively superior, but everyone already learned to type on QWERTY. To beat Nvidia head-on you wouldn't just need a better chip. You'd need to rebuild 18 years of software and retrain every engineer alive. Almost everyone figured that out and walked away. One company didn't.
The Numbers Behind Nvidia's Throne
The scale of the lead is hard to overstate. In the quarter that ended in April 2026, Nvidia's data center business - just the AI chip part - pulled in about $75 billion. In one quarter. That was up around 92% from the year before, at a gross margin sitting near 75%. Seventy-five percent margin. For every dollar of chips out the door, about 75 cents is pure profit, while it sells out everything it makes. Most hardware companies on Earth would kill for 30. Nvidia does more than double that.
Its current chip, Blackwell, is already shipping in volume, and the next one, Rubin, lands in the second half of 2026. So the standard story - there's Nvidia, and then there's everyone else trying to climb the same mountain - is basically true. It's just not the whole map. Because most of these companies aren't climbing the same mountain at all.
Five Companies, Five Different Mountains
AMD is the only one that looked at Nvidia's mountain and started climbing it head-on. It makes the same kind of general-purpose chips and made a simple bet: build a better, cheaper Nvidia and win on price and performance. The hardware is genuinely close. In one independent 2026 benchmark, AMD's top chip came within a few percent of Nvidia's best on exactly the inference work that's becoming the whole game, and its next generation packs in more high-speed memory than Nvidia's current flagship. On paper, it's right there. The wall is software. AMD's answer to Cuda, called ROCm, spent years buggy and painful. It has gotten dramatically better. It's still not there. If Nvidia's moat ever cracks, it cracks here first.
The others picked different ground entirely. Groq went for raw speed. Cerebras went for enormous models. And Tesla went off and built a private mountain just for itself - the Dojo project and its current AI4 and AI5 chips, which it never sells. Those are for making its own cars drive, its own Optimus robots think, and its upcoming SpaceX satellites compute. It's Apple building its own chips for its own Macs instead of buying from Intel. You build your own brain for your own body. AI5 taped out - the moment a design is finalized and sent to be manufactured - in April 2026, and it ramps to volume in 2027.
Groq's Speed Trick and Its Hard Ceiling
Groq was founded in 2016 by Jonathan Ross, the engineer who had built Google's first TPU - a tensor processing unit, Google's own custom AI chip - almost as a side project. Ross knew exactly where Nvidia was unbeatable, which was training, and exactly where it was soft, which was inference. He went straight at the soft spot.
Groq's chip is an LPU, a language processing unit, and in inference it is stupid fast: about 750 words per second of output, with a time to first token - how long before the AI starts replying - of roughly 18 milliseconds. On a normal cloud GPU that wait is often 200 to 400 milliseconds. So Groq answers ten to twenty times faster. It pulls this off by putting the model's memory directly on the chip, ultra-fast memory called SRAM, instead of running to slower memory for every ingredient. Keep the whole cookbook open on the counter versus walking to the pantry and back for each item.
The catch is brutal. That on-chip memory is tiny, about 230 megabytes per chip, against 80 gigabytes on an Nvidia H100. To run one big model, Groq needs hundreds of chips wired together. Blazing fast, not cheap, not simple to scale. It's the right tool for quick responses and the wrong tool for heavy lifting, and using a giant Nvidia GPU for quick responses is like booking a freight train to deliver one pizza. That gap is real. Which is exactly why what happened to Groq matters so much.
Inference Is Where the Forever Money Lives
Now remember the shift: money moving from writing the cookbook to cooking the dish a billion times a day. The training mountain, the one everyone's been staring at, is becoming the smaller business. The cooking is the forever money. As of 2026, inference is already about two-thirds of all AI compute, and that share is still climbing toward 70, 80%.
That's the strongest argument against Nvidia, and I take it seriously. Its beautiful 75% margins were built on training, and the giant cloud companies are now building their own cheaper chips to do inference in-house. Google has its TPU. Amazon has Trainium. Microsoft has Maia. Every one of them is designed to run AI inside their own data centers without paying Nvidia's tax. More people use AI, which means more cooking, which means demand that never switches off, and whoever collects the toll on all that cooking owns most of the cash flow of the AI age. So those margins probably can't hold. I just can't tell you exactly when they break.
The $20 Billion Move That Quietly Ended the Race
On Christmas Eve 2025, Nvidia struck a deal worth about $20 billion. With Groq. It didn't buy the company outright. It paid roughly $20 billion for a perpetual, non-exclusive license to Groq's entire patent portfolio and software stack, and in the same move hired Jonathan Ross himself, his president Sunny Madra, and somewhere around 80% of Groq's people.
The structure is the genius part. Acquire the company outright and regulators show up asking antitrust questions about the most powerful chip maker on Earth swallowing its most dangerous rival. License the technology and hire the humans, and you get the same result with far less heat. The scariest inference challenger just became Nvidia. So the useful question becomes which of these lanes you can actually win in, and which is a graveyard.
Which Lanes Win, and Which Are Graveyards
My read: Nvidia keeps the crown straight through this inference era, helped by the Rubin ramp in late 2026 and by the fact that it just bought one of the best inference teams alive. The thing I'd watch is that margin. Seventy-plus percent profit against an increasingly crowded field feels unsustainable over the long run, but AI is booming, and calling the top is a good way to look foolish. I'd put maybe a 25 to 30% chance I'm wrong about how long the reign lasts.
AMD isn't toppling Nvidia anytime soon, but its odds of staying a strong, distant number two look good. It makes serious chips, has real customers, and the whole industry is desperate for a second supplier so it isn't 100% hostage to Nvidia's pricing. Watch whether ROCm finally gets good enough that engineers reach for it by choice, and whether the big clouds buy AMD in volume just to keep a backup alive.
Cerebras is the one that worries me. The wafer-scale technology is seriously impressive, and it just went public on a reported multi-billion-dollar OpenAI inference backlog. But for years most of its revenue traced back to a single customer in the UAE, a group called G42, and now a huge chunk of its future rides on OpenAI. Impressive technology with one or two whales is fragile. If one of them swims away, the whole story changes overnight.
Tesla is one piece of the FSD, Optimus, and satellite thesis, not a standalone reason to own the chip. I'd watch AI5 yields and timelines through 2027, and I'd remember this is the company that built Dojo, disbanded the team, then brought it back, and that Elon has never met a deadline he couldn't miss. But if the silicon lands, Tesla controls its own cost structure end to end, and that's an enormous margin edge nobody can tax. As for Groq? It's Nvidia now. That's the whole point.
Related
Keep reading
JUL 01, 2026
NVIDIA Absorbs Groq in a $20 Billion Move as the Chip War Shifts From Training to Inference
With inference already running two-thirds of the world’s AI compute, NVIDIA’s Christmas Eve licensing deal for Groq’s patents, software, and top engineers…
AUG 12, 2026
It's Not NVIDIA You Should Be Watching: Why Inference and Power Shift the Race
Nvidia crossed $5 trillion selling specialized AI calculators. The real fight is not who builds a better training chip — it is who collects the toll every time…
AUG 31, 2026
The $500 Billion Nvidia Financing Push Is the Real AI War
Nvidia announced compute-financing platforms with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs, and KKR designed to mobilize more than $500 billion…