Video Part 5 of 6 3:19

Chunking explained

Before a document can be searched by meaning it is cut into pieces, and where the cuts fall decides whether the right answer is ever found. Too big, too small, cut in the wrong place, the two rules on the model side, and why nobody's default was chosen for your documents.

A frame from the video: the heading Part-time staff only in orange above a dashed cut line, the rule Unused days can be carried over in a piece below it, and the caption Stored without its context.

You ask your company’s AI assistant whether you can carry over unused holiday. The handbook says yes, in a paragraph under a heading that reads part-time staff only, and the assistant says yes and leaves out the heading. Nothing was wrong with the search. The document had been cut in the wrong place. This video shows why documents are cut into pieces before they can be searched by meaning, and what goes wrong with the cuts: too big, and one set of numbers stands for several ideas; too small, and a piece stops making sense on its own; cut in the wrong place, and a rule is stored without the heading that limited it.

Then the fixes. Cut where the document already has edges, copy the headings into each piece as Google’s document parser does, or write a note of context onto every piece as Anthropic did, which in their published test cut the searches that missed the right piece from about one in eighteen to about one in twenty seven. It ends on the model side, where a question and the passage that answers it are different kinds of text and every piece must go through the same model, and on the only way to know your cut is right: run real questions and count how often the right piece comes back.

Every claim was checked against Qdrant’s, Pinecone’s and Weaviate’s chunking guides, Anthropic’s contextual retrieval post, Google’s Vertex AI and Gemini embedding documentation, the defaults in LangChain, LlamaIndex and Google’s RAG Engine, and Chroma’s 2024 chunking report, in September 2026. No product is recommended. This is an independent explainer, not affiliated with any vendor.

Transcript

You ask your company's AI assistant whether you can carry over unused holiday. The handbook says yes, in a paragraph under a heading that reads part-time staff only. The assistant says yes, and leaves out the heading. Nothing was wrong with the search, because the document had been cut in the wrong place.

An embedding model reads a limited amount of text at once and turns it into one list of numbers. Google's current models take about two thousand tokens, roughly three pages, and by default they quietly drop anything past that. That is why a document is cut into pieces first, and each piece gets numbers of its own.

Cut the pieces too big and one list of numbers has to stand for several ideas at once. Qdrant calls the result a semantic average. A question about one of those ideas then lands near none of them.

Cut them too small and a piece stops making sense on its own. Five words from the middle of a policy have a meaning, but not the one the policy meant. A useful test, from Weaviate, is whether the piece makes sense to you when you read it alone.

Then what went wrong with the holiday answer? The piece that was found had the whole rule in it. What was missing?

The missing part was the heading. A cut made every few hundred words had separated part-time staff only from the paragraph it governed, so the piece said yes with no condition on it. The rule was stored without its context.

The first fix is to cut where the document already has edges. Headings, paragraphs, table rows and code functions come first, and size comes last. Google's document parser adds one more step. It copies the headings above a piece into the piece itself, so the piece still carries its own heading.

Anthropic went one step further, and had an AI write a short note of context onto every piece before it was stored. In their published test, the right piece was missing from the top twenty results for about one search in eighteen. With the notes, it was about one in twenty seven searches.

The model side has two rules. First, a question and the passage that answers it are different kinds of text. Google's example pairs the question of why the sky is blue with a passage about sunlight scattering, and as sentences they mean different things. That is why Google, Cohere and Voyage ask you to label each one as a question or a document.

Every piece and every question must also go through the same model, because numbers from two different models cannot be compared. Change the model, and everything is embedded again.

There is no right size. The tools ship defaults from a few hundred to a few thousand tokens, and Chroma's twenty twenty four report found the cutting method alone moved recall by up to nine points. The only way to know is to run real questions and count how often the right piece comes back.

Chunking is where retrieval is won or lost. Give every piece one idea, with enough around it to make sense on its own, and cut at the document's own edges. Then measure it, because nobody's default was chosen for your documents.

Found this useful?

Subscribe for the next one, or tell me what you want explained. I take requests.