Video Part 1 of 5 3:30

How AI learns from text

Before an AI model can answer you, it reads a vast amount of text and sits one test, guessing the next word, trillions of times. What goes in, what gets thrown out, who marks the guesses, what the model's numbers are, what it costs, and why there's no text inside it at all.

The video's title card: the label How AI is made above the headline Nobody taught it to write, and in blue, It learnt by guessing.

Nobody sat an AI model down and taught it grammar or facts. It learnt them by repeating one simple exercise trillions of times. This video follows that exercise: a vast amount of text, from web pages and books to code, Wikipedia and a lot of Reddit, most of it thrown away before training begins, and then one test sat on every passage. The next word is covered up and the model guesses it. The text marks its own test, because the word that really comes next was there all along, so nobody has to label anything.

Inside the model are billions of numbers, called weights, drawn here as tiny dials. A wrong guess turns them a fraction, and trillions of those turns add up to a model that writes. Along the way you’ll see what that costs, for the biggest version of Meta’s Llama 3, why there is no text inside the model at all, only numbers, why it can miss recent news, and why what comes out of this stage still isn’t an assistant.

Every claim was checked against OpenAI’s GPT-2 and GPT-3 papers, Meta’s Llama 3 paper and model card, Google’s post on its Reddit partnership, and research on what models memorise, in October 2026. This is an independent explainer, not affiliated with any AI company.

Transcript

Nobody sat an AI model down and taught it grammar or facts. It learnt them by repeating one simple exercise trillions of times.

Training starts with a vast amount of text, from public web pages, books, code, Wikipedia, and yes, a lot of Reddit. Meta trained its Llama 3 model on about fifteen trillion tokens. A token is a small piece of a word. Reading that much text, all day, every day, would take a person thousands of lifetimes.

Most of that text is thrown away before training begins. For GPT three, OpenAI started with forty five terabytes of web text and kept only about one per cent of it. Meta says it removed repeated pages, adult sites, and sites full of personal information.

Then the model sits the same test on every passage. It sees a piece of text with the next word covered up, and it guesses the word underneath.

So who marks the guess?

The text itself marks the guess. The word that really comes next was there all along. Every sentence ever written is a ready-made test, and nobody needs to label anything. Unlike a school test, no teacher sets the questions, which is why this stage can learn from so much material.

Inside the model are billions of numbers, called weights. Picture each one as a tiny dial. Together, their settings decide which word the model picks next, and no single dial means anything on its own.

When a guess is wrong, the dials turn a tiny amount, so that next time, the real word stands a slightly better chance. A single turn does almost nothing on its own, but add up trillions of them and you get a model that writes fluently.

Something surprising comes out of all that guessing. To guess well, the model has to learn whatever helps it predict, even though nobody teaches it directly. Facts help with capital cities. Grammar helps a sentence fit together. Patterns in code help with the next line of a program.

Running that test at this size takes enormous computing power. For the biggest version of Llama 3, Meta used up to sixteen thousand specialist chips, and nearly thirty one million hours of chip time. On one chip, that would take about three and a half thousand years.

Many people believe the model keeps a copy of everything it has seen, and searches it for answers. In fact, there is no text inside the model at all. What it holds is numbers, the dials set during training. Meta's biggest model holds four hundred and five billion of them, after training on fifteen trillion tokens.

What the model keeps is mostly patterns. A passage it saw many times can still be repeated exactly.

The text is collected up to a date, and then training starts. Llama 3 learnt from text up to the end of twenty twenty three, so anything after that was never in its pile. That is why a model can miss recent news, unless the app adds a search of the internet.

What comes out of this stage is good at extending any piece of text, but it is not yet an assistant. Ask it a question and it may simply write more questions, because the web is full of quiz pages. Turning it into a helper takes another stage of training, with people showing it good answers.

Pre-training is one test, repeated trillions of times. The model reads a vast pile of text, guesses the next word, and the text itself marks every guess.

Found this useful?

Subscribe for the next one, or tell me what you want explained. I take requests.