Training data often needs a person to say what each example is. Now an AI labels first, a check scores how far to trust each label, and people settle the hard cases. Where labels fit in making a model, what it saves, and why the AI's own confidence isn't enough.
To teach a computer to sort customer reviews into happy and unhappy, you need thousands of reviews that someone has already labelled, and doing that by hand is slow and expensive. These days, an AI often does the first round. This video starts by putting labels in their place: they aren’t how a chatbot learns language, which needs no labels at all, but they come in wherever someone has to say what is good, what is bad, or where something belongs, and even help choose the text a model reads.
Then the three steps. An AI suggests a label for every item, far faster and cheaper than people, as studies from Microsoft and others found. A confidence check, run by a separate model, scores each label, because a model asked how sure it is tends to be overconfident. People then review the low scores, the edge cases and anything ambiguous, like a review that says “great, another delay”.
Every claim was checked against the Llama 3 paper, Microsoft’s and Gilardi et al.’s labelling studies, research on how well models judge their own confidence, the CHI 2024 paper that describes the three steps, and Amazon’s documentation for automated labelling, in October 2026. This is an independent explainer, not affiliated with any AI company.
Transcript
To teach a computer to sort customer reviews into happy and unhappy, you need thousands of reviews that someone has already labelled. Doing that by hand is slow and expensive. These days, an AI often does the first round of labelling.
Labels aren't how a chatbot learns language. That first stage reads raw text, and the next word in the text is the only answer it needs. Labels come in afterwards, wherever someone has to say what is good, what is bad, or where something belongs.
Labels even help choose the text a model reads. Meta asked its Llama 2 model to judge a sample of web pages, trained small, fast filters on those judgements, and used the filters to pick the text for Llama 3.
First comes pre-labelling. A large language model reads each review and suggests a label for it. In the time a person labels a handful, it gets through thousands.
Researchers at Microsoft found that labels from GPT three worked as well as labels from people. They cut the cost by between fifty and ninety six per cent. A second study compared ChatGPT with crowd workers, people paid online for each small task. ChatGPT came out ahead on four out of five labelling tasks.
The catch is that some labels will be wrong, and you need a way of spotting the mistakes. Ask a model how sure it is, and research shows it tends to be overconfident. That makes its own confidence a poor guide.
So how would you decide which labels a person should look at?
Second comes a confidence check. A separate model, called a verifier, gives each label a score for how reliable it looks. Labels that score well go straight through. The low scorers are set aside for a person to review.
Third comes human review, where people look at the low scores, the edge cases and anything ambiguous, and fix the labels that are wrong. A review saying great, another delay, sounds happy, but the customer is clearly annoyed.
None of this is a new idea. Amazon's labelling service ran the same loop with a smaller model. It set an accuracy bar, and sent people every label that fell below it. What large language models changed is how good that first round has become.
The AI labels everything first, a check flags the labels to doubt, and people settle the hard cases.
Found this useful?
Subscribe for the next one, or tell me what you want explained. I take requests.