Skip to content
farzad.fm
AI & Automation

Tesla & SpaceX's New Product Is Elon's Bet Of A Lifetime

I gave my Grok Bot a raw video and no direction, and it came back with a decent edit built from the transcript alone. Macrohard, also called Digital Optimus, is Tesla and SpaceX trying to build AI agents that actually watch the screen. That takes Tesla's camera-based driving system plus SpaceX's chips and power, and it will burn far more tokens than text.

Loads from YouTube only after you press play.

Watch on YouTube

Recently I recorded a talking head, just me rambling at a camera the way I usually do. I gave the video file to my Grok Bot, which is an agent, and told it: "do whatever you want with this, just edit it for me." No other direction. The edit it came back with was honestly not bad. And it built the whole thing by reading the transcript. The agent knew what I said. It didn't really know what the video looked like.

Tesla and SpaceX are making a bet of a lifetime on closing that gap. About a year ago, in August 2025, Elon Musk, CEO of both companies, announced a project called Macrohard at his AI company, xAI. Everyone ignored it, because the name is kind of stupid. Seriously, who names a freaking company Macrohard? In March, Elon started calling it "Macrohard or Digital Optimus," a joint xAI-Tesla project. xAI is now part of SpaceX and goes by SpaceXAI, so today it's a Tesla and SpaceX project. Two names, basically the same thing. Nobody is paying attention, because everyone is distracted by the AI-is-going-to-kill-us-all story, with the CEOs of Anthropic and OpenAI scaring everybody, and by the bubble or not-a-bubble hysteria. Underneath that noise, the way AI gets executed is changing, and this joint project is going to redefine how all of us interface with it.

What my agent did with zero direction

The agent decided when to cut in B-roll, the scenes that aren't just me talking, and it designed what that B-roll looked like based on what I was saying. It decided when to animate letters and when to drop them in. It animated a little typing effect. All of it came from reading the transcript and overlaying graphics on the footage.

It worked because I have a computer running DaVinci Resolve that's also hooked up to AIs like Claude, ChatGPT, and Grok. Together they edited the video with literally zero direction from me.

Keep that in your head for the rest of this piece. There are AIs now that can essentially look at a computer screen and figure out what to do without being told. That is a really big deal. It's also why this has to be a joint project. Neither Tesla nor SpaceX can pull it off without the other one's piece.

What Digital Optimus is trying to do

At its foundation, Macrohard, or Digital Optimus, or whatever it's called these days, is an attempt to get AI agents to interpret what's on somebody's screen and act on it. Editing a video. Buying something online. Playing a video game, or designing one. In every case you want the AI to understand what the thing looks like and to make something that looks good and works well for humans. That takes an AI that can actually look at a screen and then act.

AI today works very differently. It's built mostly around code, or around what it thinks it sees in a still image, like a screenshot pulled from a video. It doesn't really understand motion. It doesn't understand taste either, that very human reaction where you look at something and go, "Hey, that's human. A human did that versus AI." That's very hard to capture.

We capture it because our eyes are always tracking what's in front of us, like a video. The brain runs on about 20 watts, and it figured out how to build this entire civilization. We're an unbelievably efficient system for executing on intelligence, using our eyes, our ears, how we feel the world, and how we move.

Tesla already taught a machine to see

The system for looking at something with "eyes" and then acting is something Tesla already runs in its cars. It's FSD, Full Self-Driving. New Teslas have spent the last few years collecting data as people drive. The car feeds images and video from its cameras through an AI system, and that system figured out how to drive like a human. Somebody's here, go left. An object's there, go right. That light-looking thing turns red, stop.

Tesla is putting that same system into Optimus, its humanoid robot. Two legs, two arms, two hands, ten fingers, and cameras for eyes. Instead of dodging pedestrians, it navigates the real world the way you or I would. The goal is for it to pick up cups, use tools, clean up junk, fold laundry, run a lawn mower.

Now Tesla is going to take that exact system and put it into the computer screen. Our eyes see because photons come back to them. I'm not stopping to look up the official definition of a photon. For the robot in the computer, this Digital Optimus, the input is pixels. It sees the screen through those pixels and works out how to move.

The blanket test

Say I'm on Amazon looking for a blanket with a very specific look or shape. Most of the time that detail isn't in the description or the metadata, the data that describes the data. The listing says blue blanket or red blanket. Today's AIs aren't great at finding a specific shape. They might grab a still image, run it through, and guess.

A system that works natively off the screen just looks. It comes back with something like, "Yo, this is the exact shape you were looking for. I went through a trillion pages. I found this one." Then it sends me the link.

Then it goes a step further and watches actual motion. Say I'm clicking through a complex process across a bunch of screens, like editing a video. Or it's work that takes a lot of taste, where I want a specific color here and a specific object there. The system studies how I do it and replicates it. I don't have to do that step anymore. I can do something else, or have a million of those robots doing it inside the computer.

Why today's agents miss

Compare that with the agentic harnesses we use today. Even a chatbot is technically one. You ask why your feet smell bad, it tells the model somebody's trying to talk to it, and the model answers pleasantly.

One step up you have Claude Code and Codex, plus harnesses like Hermes and OpenClaw. They're more technical and more customizable, and they take actions: booking something on your behalf, coding a whole website or program. But they work the way AI traditionally has. They read the code and the data behind the screen. They don't necessarily look at the screen itself.

That's why these AIs are often bad at taste. They don't really know what they're creating visually. They're still unbelievably useful. I talk to a lot of entrepreneurs who love them for the productivity. But without that visual layer, they're limited in what they can perceive and execute, and a lot of kinds of work stay out of reach. Editing is one.

It's also why my edit test blew my mind. If I'd had an AI that could actually look at the screen, that edit would have looked a thousand times better. The editors working on my videos are going to get tools that let their imagination go wild. Instead of figuring out some complex animation, they'll just tell the thing, "do this animation."

Video eats tokens

The downstream effect of this kind of AI: token use is going to absolutely explode.

Tokens are the units of intelligence, the data you feed a model to ask a question or request an action, plus what comes back out. Right now most tokens are text, and text is tiny.

Rough example. On camera I guessed four books might be 20 to 100 kilobytes. The real number is more like two or three megabytes of plain text. A half-hour video like this one is around three gigabytes, roughly a thousand times bigger.

Now think about time. Four books take me maybe five or six days to read if I devote whole days to it. A 30-minute video takes 30 minutes. The books are a fraction of a fraction of the size. The video is a fraction of the length, and the file is gigantic.

So an AI working from video has to chew through far more data just to figure out what it's looking at. Then it has to decide what to do and check its own work. Did I do this right? Did I do this right? It's absolutely bananas.

Why it takes Tesla and SpaceX

Tesla figured out the visual side with its cars. Cybercab doesn't even have a steering wheel or pedals. Tesla is betting the software is safe enough not to need them, though federal regulators are still questioning that design. Eight to nine cameras constantly taking in video, pushed through one computer inside the car, making decisions 24/7. That's an unbelievable feat of engineering. I'd call it probably the most intelligent piece of technology per unit of energy. Actually, scratch "probably." It's 1,000% true, given how much data that computer processes to make sure nobody inside or outside the car dies.

But doing this for 8 billion people on their screens takes an unbelievable number of chips and an unbelievable amount of power. I wonder what SpaceX is working on. Oh yeah, Terafab, the gigantic chip factory SpaceX is building with Tesla in Texas. The goal is more than a terawatt of compute a year, which Tesla says is more than every chipmaker in the world combined can make today. That's how vast demand for AI is.

SpaceX also just announced Starbase Louisiana. It's going to be a spaceport for all the Starship rockets they're building, with construction starting in 2027. Louisiana is one of the country's biggest natural-gas producers, and SpaceX plans to make its own power and rocket fuel on site. The way I see it, that kind of energy can power a lot of compute while they ramp solar and put data centers in space to soak up energy from the sun. Every one of those systems needs chips.

That's why these companies are getting so tight-knit, and why there are all these merger rumors. The two systems are extremely complementary. None of this works without mass scale. You either run out of data on the Tesla side or run out of chips on the SpaceX side to power something this token-hungry.

Physical labor and digital labor

Start with the physical side. Self-driving vehicles move goods and people from one point to another. I think there will be hundreds of millions of them in the next 10 to 20 years. It's unstoppable. When something is that cheap to run, with no driver to pay, and that safe, everyone is going to use one or buy one.

Then the humanoids. Physical robots doing useful work, like folding laundry. I hate it. I know you hate it too. Construction, underwater work, cleanup. Basically anything a human can do in the physical world, these robots will do.

That's two giant pillars on the physical side: transportation and labor. On the digital side you have Digital Optimus, or Macrohard, doing the same thing inside the computer.

What's being built is the foundation of the next generation of civilization, where the work gets done for us by AI. That's hard to wrap your head around, and I know it sounds scary. It's still where the technology is going. It doesn't matter whether it's Tesla or SpaceX. Somebody is going to do this. It'll probably go slower if someone else does, but it makes too much sense and there's too much economic incentive.

Who controls the robots?

Which brings up the biggest question of all. Who controls these robots, and who decides how they get allocated?

  • The richest, most powerful people?
  • Governments?
  • Some kind of democratic institution?
  • One person, or one company?
  • A conglomerate of companies, or of governments?

Once humans no longer have to do most of today's roles, you start asking what the future looks like. Do humans end up renting all labor? How do people get the resources to rent it? How is value generated from these systems, and how does it get distributed?

You can see why some people are freaked out. There are a lot of open-ended questions, and it gets really weird when the foundation is labor that isn't done by humans. That's what's going to happen. It comes down to the chips, the power running those chips, the robots running on both, whether humanoids or cars, and the digital systems in our computers doing digital work.

Baby AI

This whole project, the one nobody is paying attention to, is mind-blowing. Everything we have right now is "baby AI." Everyone's freaking out about it, and this is nothing. I promise you, this is nothing.

This is how technology moves. Scary and exciting at the same time. I think that's how the industrial revolution felt to people living through the late 1800s into the early 1900s, with the railroad and electricity, and even the internet to an extent. All of it happened because humans willed it into existence.

This Exclusive is from the long-form at https://www.youtube.com/watch?v=W-On1ast00k.

Watch

Prefer the tape?

Open the source video on YouTube if you want the full cut.

Digest

Prefer the daily pulse?

Short breakdowns of what actually moved, every day.