How AI Learned To Think
A wedding toast stays flat while AI writes computer-checked Olympiad proofs. The split is not heart. It is the answer key: reinforcement learning with verifiable rewards, jagged capability, and who owns the checker.
Loads from YouTube only after you press play.
Watch on YouTubeRelated
Keep reading
Picture a wedding. The best man taps a glass and reads a toast an AI wrote. It is polite. Nobody laughs. Nobody cries. The bride smiles the way you smile at a nice email.
Now the strange part. In 2024, an AI wrote math proofs a computer checked line by line at silver-medal level for the International Mathematical Olympiad. A kid at that wedding can tell the toast was flat. No kid can check that proof.
So why is AI brilliant at proofs and flat at toasts? Watch the upload if you want the full walkthrough. This is the same case in writing.
Imitation has a ceiling
Start with how these machines learn to write at all. Your phone keyboard offers "soon" after you type "see you." A language model is that same next-word trick grown enormous. It reads a huge pile of human writing and plays one game: guess the next token. Compare. Nudge. Repeat across more text than any person will ever read.
That is imitation. Feed it enough recipes and it writes a recipe. Feed it enough math homework and it writes math homework, mistakes included. Under imitation, math and toasts are the same job. Copy how people write proofs. Copy how people write toasts. To a copy machine, both are just text.
Imitation plus some polish from human ratings gave us the chatbots you already know. Fluent email in seconds. For everyday questions, that is enough. But imitation has a ceiling. Copy the best student in class perfectly and you are exactly as good as that student, never better. People also write wrong answers online. The copy inherits all of it.
There is a sneakier problem. Finished answers get published. The scratch paper with the wrong turns and the fixes mostly gets thrown away. The model learns what good answers look like. It copies the look of thinking. Ask it a hard problem nobody has worked through in writing and it can hand you something shaped like a proof: neat and wrong. Imitation teaches the look. It does not teach the work.
Practice against an answer key
To get past the best answer anyone ever wrote, the model has to stop copying and start practicing. Stop showing it answers. Hand it a problem. Let it try. Tell it only one thing: right or wrong.
Math has an answer key. Nobody scores a wedding toast. Spend a summer alone with a stack of math problems and a key, and by August you are better. Nobody had to teach you a thing. Spend the next summer writing toasts, finish one, shrug, and practice barely moves you, because nobody scores a wedding toast.
Labs built the machine version of that summer. The main part is an automatic checker: a small program that marks each answer right or wrong in a blink. For math, compare the final number with the key. For code, run the tests. Then the loop: attempt, check, reinforce. Whatever led to a right answer gets nudged. Run that day and night and small nudges become skill.
Reinforcement learning means learning from a score instead of from examples. A verifiable reward is a score a machine can check by itself with no person in the room. Researchers at the Allen Institute for AI coined a name for the whole recipe: RLVR, reinforcement learning with verifiable rewards. That coinage and framing are on my tape; treat the label as their research term as I used it in the video.
Look at which attempts win. A model that blurts the final number is guessing. A model that writes each step can catch a slip before it ruins the answer. Longer, careful attempts score more often, and the checker rewards them. Over time the model starts writing its own scratch paper. It checks its own steps. It backs up when something looks off. That is what people mean when they say a model thinks.
o1, R1-Zero, and the score jump
On tape I put OpenAI shipping o1-preview in September 2024, then o1 by December 2024: it works before it answers. Then a cleaner test. A model called R1-Zero practiced against nothing but a rule-based checker. Nobody hand-taught it to reason. Tested on AIME, a math contest, its first-try score rose from 15.6% to 77.9%, from mostly wrong to mostly right by practice against an answer key.
One note from the video: that jump is R1-Zero only. The released R1 got thousands of curated example solutions first. The training method behind it is called GRPO: the model answers the same problem several times, and the answers that beat the group's average get reinforced. It races itself. That breaks the imitation ceiling. It does not touch the toast.
Those percentages, the o1 timeline, and the R1-Zero versus released-R1 distinction are as I stated on tape. I am not upgrading them into a paper citation here.
Proofs that compile
Where does the practice pay off? AI thinks best where answers can be checked by a machine. Math has an answer key, so a machine can practice all night and every attempt gets scored. Nobody scores a wedding toast, so practice on toasts moves slowly because every score needs a person.
The cleanest checker of all is a proof written in a computer language that checks every step. If one step is wrong, the proof does not compile. Remember the Olympiad proofs from the start? On my tape, DeepMind's AlphaProof wrote most of those proofs in a language like that, and every step passed the computer's check.
Separately, I said OpenAI claimed its AI produced a computer-checked proof that the famous fluid equations can break down, only if you allow one specialized push from outside. Experts say it looks technically correct, but the Clay Mathematics Institute has not awarded the prize, and the harder version of the question is still open. That is a hedge, not a medal ceremony. Both proofs share one thing: in math there is a referee that cannot be fooled. A wedding toast has no compiler.
Jagged by design
Here is the part I want you to sit with. The research for this video, the script, my cloned voice, the animation, and the edit were all made by AI. The idea and the analogy are mine. You have been listening to that stack the whole time. A video like this has no answer key. The only score it gets is yours.
AI does fuzzy work too. It just improves there more slowly. Picture AI's skills as a mountain range. Tall peaks where a checker exists. Low ground right beside them where none does. That shape is jagged: brilliant at one task and clumsy at the task right beside it. AI is jagged by design. The checkers decide where the peaks go.
Anything with a cheap automatic checker improves fastest. Cheap matters. A checker that runs in a blink lets the model practice all night. Software is the obvious one: code passes its tests or fails them. Math written for a computer to check. Chip design where a simulator shows whether a circuit works before anyone builds it. Paperwork checked against written rules. Moving data between systems where every row lands in the right place or it does not. Fuzzy work, taste, a hard talk with your teenager, improves slowly. Slowly is not never. There the score comes from people, one judgment at a time.
Who owns the shovel
Picture a tireless student who practices all night against a checker that never gets bored. Which work does that student reach first? Easy to grade beats easy to do. A junior coder's work gets graded by tests every hour. A wedding planner's work gets graded once by the mother of the bride. Guess which one meets the tireless student first.
The frontier of AI follows the frontier of what can be checked. In a gold rush, the people selling shovels get paid no matter who strikes gold. Investors call that picks and shovels. Here the shovels are the checkers and the practice environments: a fake codebase or a fake inbox with a checker built in. The better the fake world, the better the practice.
For any company or job you are sizing up, ask two questions. Is there a checker? And who owns it? A company sitting on years of graded work, like approved insurance claims, owns an answer key nobody else has. Whoever owns the checker decides who gets to practice. That tells you where the next peaks rise.
Now take one last look at the wedding. The best man folds the paper to polite applause. That toast was flat for the same reason the proofs got great. One kind of practice got scored. The other did not. That flat toast is a map. It marks the exact spot where the checkers stop.
On one side sit proofs and passing tests, and the tireless student already lives there. Every new checker someone builds moves that border. On the other side sit the toasts. The score still comes from people, one judgment at a time. You walked in wondering if that toast needed a heart. Now you know what it was missing.