World Models and AI Chips: Why Physical AI Needs a Different Kind of Speed

The biggest constraint on physical AI is not chatbot cleverness but real-time speed under tight power limits, as when a self-driving car must see, remember, predict, and act in fractions of a second rather than the multi-second delays common in systems like ChatGPT. Large language models predict the next token in a stream of symbols, while world models must track occluded objects, simulate physics, and choose actions dozens of times per second—handling video that dwarfs ordinary text by orders of magnitude. Approaches already diverge: DeepMind's Genie 3 generates interactive environments, Meta's V-JEPA 2 predicts compressed scene embeddings from over a million hours of video, and Nvidia's Cosmos mixes language, vision, audio, and action. That disagreement is why a purpose-built world-model chip does not exist yet, even as data-center GPUs like Nvidia's Rubin and edge systems like Jetson Thor or Tesla's FSD stack pull opposite ways on memory and watts. Overall, this points to specialized silicon, memory makers, and vertically integrated players shaping robots and cars far more than today's chat interfaces suggest.

Check the video here.

Previous
Previous

Nobody Read the Warning Buried in SpaceX's $2 Trillion IPO: What the Market Is Missing

Next
Next

Humanoid Robots as Labor Utilities: Ten Ways Cheap Machine Hours Reshape Work