Video Part 4 of 5 2:56

What the B in 8B means

AI models come in sizes written like 8B, 70B or 405B. The B is billions of parameters, the numbers training sets. What that number tells you about the memory and hardware a model needs, how quantisation shrinks it, why bigger isn't always better, and why some huge models only use a slice at a time.

The video's title card: the label How AI is made above the headline 8B. 70B. 405B. and, in blue, What the number tells you.

AI models often carry a number in their name, like 8B, 70B or 405B. It’s tempting to read it as a score for how good the model is, but mostly it tells you about the hardware. The B means billion, and it counts the model’s parameters, the numbers training sets, drawn here as tiny dials.

Every one of those numbers has to sit in memory while the model runs, so the count decides what a model needs: one large graphics card for Llama 3.1 8B, more than one server for the 405B version. You’ll see how quantisation shrinks that, why a bigger model costs more for every word, why DeepMind’s smaller Chinchilla beat its bigger Gopher, and why some huge models, like DeepSeek R1, only use a slice of their dials at a time.

Every claim was checked against Hugging Face’s documentation and Llama 3.1 figures, the GPT-2, GPT-3, Chinchilla and scaling-law papers, Microsoft’s Phi-2 post, the DeepSeek R1 and Mixtral pages and the GPT-4 report, in October 2026. This is an independent explainer, not affiliated with any AI company.

Transcript

AI models often carry a number in their name, like eight B, seventy B or four hundred and five B. It's tempting to read that number as a score for how good the model is. Mostly, it tells you about the hardware.

The B stands for billion, and it counts the model's parameters. A parameter is one of the numbers that training sets, also called a weight. Picture each one as a tiny dial, so Meta's Llama eight B has eight billion of them.

Those counts have grown fast. OpenAI's GPT two had one and a half billion parameters in twenty nineteen, and GPT three had a hundred and seventy five billion a year later. DeepSeek's largest model, released in twenty twenty six, has one point six trillion.

Every one of those numbers has to sit in memory while the model runs. In the usual format, each one takes two bytes, so Llama eight B needs about sixteen gigabytes. That fits on one large graphics card.

The four hundred and five B model needs eight hundred and ten gigabytes, which takes more than one powerful server. That's before the conversation itself, which needs memory of its own.

To shrink a model, people store each number less precisely, which is called quantisation. Think of a dial with fewer clicks. Stored in four bits instead of sixteen, Llama eight B needs four gigabytes and fits on a laptop. It costs a little accuracy, and Hugging Face says the loss is usually small.

More parameters usually make a better model, because there's more room for what it learnt. The catch is that every word it writes takes about two sums for every parameter, so a bigger model is slower and costs more to run.

Size isn't the whole story, and DeepMind showed it in twenty twenty two. It trained Chinchilla, at seventy billion parameters, on far more text than its own Gopher, which had two hundred and eighty billion. The smaller model beat the bigger one.

Small models keep catching up, too. Microsoft says its Phi two, with under three billion parameters, matches or beats models up to twenty five times its size on difficult tests.

Some of the biggest models only move a slice of their dials at a time. DeepSeek's R1 has six hundred and seventy one billion parameters, but works out each word with thirty seven billion of them, so each word costs about what a much smaller model's would. All of them still have to fit in memory.

You won't find this number for the main models behind ChatGPT, Claude or Gemini. OpenAI's report on GPT four left the size out on purpose, and none of the three makers publishes it today.

So the B counts the dials, which tells you what a model needs to run and much less about the quality of its answers.

Found this useful?

Subscribe for the next one, or tell me what you want explained. I take requests.