Video Part 2 of 5 2:55

How AI learns to help

A model that has only read the internet can continue text but can't yet help you. How people turn it into an assistant, by writing good answers, ranking its own, and training a second model to score them the way they would. Reinforcement learning from human feedback, in plain words, and whose taste it learns.

The video's title card: the label How AI is made above the headline Reading isn't helping, and in blue, People teach it to help.

A model trained only on internet text is brilliant at continuing it, but poor at being helpful. Ask it to explain the moon landing to a six year old and it may write more requests of the same kind. This video follows the stage that turns it into an assistant: people writing good answers for it to learn from, then ranking the answers it writes, and a second model, the reward model, learning to score new answers the way those people would. The main model practises against that score, round and round. That loop is reinforcement learning from human feedback, or RLHF.

It ends on what that achieved and whose preferences it learns. OpenAI’s labellers preferred a small trained model over GPT-3, which is more than a hundred times bigger, at least on OpenAI’s own prompts, and OpenAI’s own paper says the model learns what a specific group of people preferred, not human values in general. Then how it’s done now, with other models helping to give the feedback in Anthropic’s Constitutional AI, and six rounds of it for Meta’s Llama 3.

Every claim was checked against OpenAI’s InstructGPT paper, Anthropic’s Constitutional AI paper, the DPO paper and Meta’s Llama 3 paper and model card, in October 2026. This is an independent explainer, not affiliated with any AI company.

Transcript

A model trained only on internet text is brilliant at continuing it, but poor at being helpful. A second stage of training, guided by people, turns it into an assistant.

Ask a model straight out of its first stage of training to explain the moon landing to a six year old, and it may simply write more requests of the same kind. It was trained to continue web pages, not to follow instructions. OpenAI's paper on this says plainly that making a model bigger does not, on its own, make it better at following what people intended.

The first step is to show it good answers. OpenAI hired about forty contractors to write answers by hand for about thirteen thousand prompts. The model was then trained on those examples. That step is called fine-tuning.

Writing every answer by hand is slow, so in the next step, the model writes several answers to the same prompt, and a person puts them in order, from the best answer down to the weakest.

People can rank a few thousand answers. Training needs feedback on millions. Where does the rest of the feedback come from?

It comes from a second model. That model studies the people's rankings until it can score a brand new answer the way they would have. It's called a reward model.

Then the main model practises against that score. It writes an answer, the reward model scores it, and the model is adjusted towards answers that score higher. It goes round that loop many thousands of times. The loop has a long name, reinforcement learning from human feedback, or RLHF for short.

The result surprised a lot of people. OpenAI's own labellers preferred answers from a trained model with one point three billion numbers over GPT three, a model with a hundred and seventy five billion numbers. A little feedback beat a model more than a hundred times bigger, at least on OpenAI's own prompts.

It's worth asking whose preferences these are. OpenAI's own paper is clear about it. The model learns what a specific group of people preferred, mostly its labellers and researchers, rather than human values in general.

These days, other models help to give the feedback. In Anthropic's Constitutional AI, a model compares two answers against a written list of principles. Its choices stand in for people's labels on which answers are harmful. People still judge which answers are most helpful.

The details keep changing, but the idea stays the same. Meta ran six rounds of this for Llama 3, with people comparing two answers at a time, on a scale from marginally better to significantly better.

Reading teaches a model to write. People teach it to help, by showing it good answers, ranking its own, and rewarding the ones they prefer.

Found this useful?

Subscribe for the next one, or tell me what you want explained. I take requests.