Nvidia Vera Rubin Will Slash the Price of Everything
After Nvidia’s Aug 26, 2026 print (106% YoY revenue on a ~$5T company), Farzad walks Vera Rubin NVL72: agents burn 4–15x chat tokens, Nvidia claims up to ~30x agent throughput per megawatt vs GB300, and Musk said SpaceX is exclusive Nvidia for Earth factories and a Starmind orbital NVL72.
Loads from YouTube only after you press play.
Watch on YouTubeOn Aug 26, 2026, Nvidia printed what Farzad calls one of the best reports in stock-market history: more than 100% revenue growth on a company already worth about $5 trillion. His Sept 10, 2026 long-form says that print points one way. The AI boom is early. The core move is collapsing the price of generating intelligence while the companies that sell that intelligence get more profitable. The product line to learn is Vera Rubin. If Nvidia’s claims hold at scale, Farzad says the important output is not a slightly faster chatbot. It is a collapse in the cost of useful digital work. That is also why he points to Elon’s Aug 4, 2026 SpaceX earnings line that the company would build exclusively on Nvidia, and to SpaceXAI’s plan for an optimized Vera Rubin NVL72 as the first-generation Starmind computer in orbit.
Agents burn tokens. Chatbots do not.
The first wave of generative AI was mostly prompt and response. You ask. The model answers. Done.
An agent is different. You ask it to research a company, fix a bug, negotiate with suppliers, or run a marketing campaign. It plans, searches, calls tools, writes and tests code, checks its own work, and sometimes spins other agents in parallel. Every step can mean another trip through the model. Context grows. Memory has to keep what already happened.
Anthropic’s published research framing (as Farzad cites it): a normal agent in its full research system used about four times as many tokens as ordinary chat. Its multi-agent system used about fifteen times as many. Tokens are the small units models use to read and write. Every prompt, tool result, instruction, and saved context adds more.
Training builds the brain. Inference pays for every thought after that. Agent jobs can run for hours or days. So the economic unit shifts. Cost per token still matters to labs. Businesses will care about cost per completed task, how long it took, and whether the result was good enough to ship.
What Vera Rubin actually is
Vera is Nvidia’s new CPU. Rubin is the GPU that does most of the AI math. The official rack is Vera Rubin NVL72: 72 Rubin GPUs and 36 Vera CPUs tied so they behave more like one machine. (Transcript ASR said “VO72.”)
The GPU is only one worker. The CPU organizes the job. Memory holds the model and the growing task history. The network moves data between GPUs. Storage saves context that may be needed again. Cooling and power have to keep up. If any layer falls behind, the GPU waits. Every idle second still burns electricity.
Nvidia’s pitch is co-design: every major part built around reasoning models and agents, not bolted together later.
On Rubin, Farzad cites Nvidia’s published specs: up to about 5x low-precision inference versus Blackwell; up to 288 GB of HBM4 per GPU; up to about 22 TB/s of memory bandwidth, roughly 2.8x Blackwell. Those 22 TB/s and ~2.8x figures are Nvidia’s July 21, 2026 technical-blog / CES published specs. The live Vera Rubin NVL72 spec table as of Sept 10, 2026 lists 288 GB HBM4 at 19.2 TB/s per GPU. Low precision means smaller numbers where the model can tolerate them, so more work for less energy and less memory traffic. Agents carry context. Bandwidth is how fast that context reaches the math units.
Inside the rack, mixture-of-experts models send different tokens to different specialists that may live on different GPUs. Nvidia’s NVLink 6, in Farzad’s telling, gives each Rubin GPU about 3.6 TB/s of two-way communication with the rest of the rack so all 72 GPUs talk as one system. CES and NVLink 6 marketing still say ~3.6 TB/s all-to-all per GPU (about 260 TB/s across the 72-GPU rack). The current product-page table lists 3 TB/s per GPU / 216 TB/s rack. Some coordination work can happen in the switches instead of bouncing everything back through the GPUs.
Vera has 88 Nvidia-designed cores and up to about 1.2 TB/s of memory bandwidth. Nvidia says it completes a set of agent training and data-processing tasks up to about 1.8x faster than common Intel/AMD-style server CPUs. More important to Farzad: Vera can share memory with Rubin over a fast chip-to-chip link, so less copy traffic and more time on the same job. He says SpaceXAI plans to use standalone Vera processors for CPU-heavy work around Grok’s agents, while Rubin runs the giant model math. (Farzad said “Grok bots”; the Aug 24 Nvidia/SpaceXAI release does not name a Grok Bot product.)
The rest of the rack hits costs operators usually ignore. BlueField processors handle networking, storage, security, and data-center ops so the main processors do not. A context-memory layer can save and reuse pieces of prior agent work instead of recalculating them. Liquid cooling pulls heat off the hardware. Nvidia’s DSX MaxLPS power controls, Farzad says, can let operators install up to about 40% more GPUs inside the same megawatt limit in some configurations. Mechanical design drops compute-tray assembly from hours to about a minute by removing in-tray cables, fans, and liquid hoses a tech has to attach.
The numbers Nvidia wants you to remember
Nvidia tested Vera Rubin NVL72 on AgentX, SemiAnalysis’s open agentic-coding benchmark (long context, tool calls, pauses, memory reuse, sub-agents). On DeepSeek V4 Pro, Nvidia says Vera Rubin delivered up to about 30x more agent throughput per megawatt than GB300 NVL72 — at 160 tokens per second per user on Nvidia’s chart — and cost per million tokens up to about 35x lower. Nvidia measured those Rubin numbers; they were still pending SemiAnalysis review as of the Aug 24, 2026 posts. Farzad said GB200; the Nvidia primary is GB300.
On that same Aug 26 earnings story, Farzad says Nvidia framed revenue opportunity per gigawatt of power as roughly $18 billion with Hopper, $25 billion with Blackwell, and $40 billion with Vera Rubin. So while token cost falls for customers, Nvidia’s claimed revenue per gigawatt roughly doubles versus two generations ago. Both sides can win if useful output rises faster than equipment price: Nvidia sells a more valuable rack; the operator pays less per unit of useful work; cheaper work creates demand that did not make sense before.
Jensen’s line, as Farzad repeats it: compute is revenue. Once models clear a usefulness threshold, more compute means more completed work, not just more chat.
Rebound demand, AWS, and $500B of outside capital
When one unit of intelligence gets cheaper, developers use more of it. Longer thinking. More context. Several agents instead of one. Self-checks. Weekly jobs become daily, then continuous. Computers got cheaper and the world bought more computers. Bandwidth got cheaper and AOL became Netflix. Efficiency does not shrink Nvidia’s market if demand grows faster.
That is Farzad’s read of the AWS partnership Nvidia announced: plans to deploy about 2 million additional Nvidia GPUs across Amazon’s global infrastructure in 2027 and 2028, including Blackwell Ultra, Rubin, and Rubin Ultra. AWS also builds its own AI chips. It is keeping those and still expanding the Nvidia fleet because customers want both.
Separately (announced Aug 10, 2026, restated in the Q2 highlights), Nvidia announced partnerships with six firms — Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs, and KKR — aimed at mobilizing more than $500 billion of outside capital for AI infrastructure over time. Falling compute costs are not killing the buildout. They are helping turn AI factories into an asset class large funds want to fund. Economists call the pattern a rebound effect.
What gets cheap after intelligence gets cheap
Once two products can both solve a job, the winner is the one that answers faster at lower cost. Software goes first: decide, write, test, fix, secure, explain, sell, support. Agents can touch every step. If agent cost falls hard, a tiny team can ship what used to need a much larger company. Then research, law, accounting, medicine, engineering, logistics, education, customer service, product design, and every other digital workflow that is mostly read → decide → act.
The price of intelligence falls first. Then price, speed, and quality of everything built with that intelligence follow. Nvidia sits in the middle of that stack if Vera Rubin ships the way the slides claim.
Keep the hedges. Aug 26 revenue growth and market-cap framing are Farzad’s summary of the print, not a substitute for Nvidia’s SEC tables. Anthropic’s 4x / 15x token multiples are research-system figures he cites, not universal laws. Rubin specs, NVLink 6 numbers, Vera core counts, 30x / 35x AgentX results, $18 / $25 / $40B per GW, the ~2M AWS GPU plan, and the $500B+ capital mobilization are Nvidia (and partner) claims as of the video. SpaceX “exclusive” Nvidia is Elon’s Aug 4 earnings wording, not an Nvidia joint PR. Starmind in orbit and standalone Vera around Grok agents are SpaceXAI/Nvidia stated plans in this piece, not finished orbital capacity. DeepSeek V4 Pro and SemiAnalysis AgentX are Nvidia’s chosen comparison, measured by Nvidia and pending SemiAnalysis review; the 30x / 35x baseline is GB300 NVL72, not GB200.
This Exclusive is from the long-form at https://www.youtube.com/watch?v=XGmAxtAHle4.
Receipt
Event → Farzad proof → book beat → buy link.
- Event: Public Exclusive on Nvidia Vera Rubin and the collapse in cost of useful digital work (XGmAxtAHle4).
- Farzad proof: Aug 26 print + agent token burn + Nvidia rack/co-design claims + AWS / $500B capital rebound framing.
- Book: AoC — cheap intelligence rewrites who can build and ship. Master Plan Energy / compute — who owns the factories that mint tokens.
- Get the books: Abundance or Collapse · Master Plan