- Chunking is how a document is split into pieces before embedding. The retriever can only return what a chunk contains, so a bad split caps answer quality no matter which model you use.
- Fixed-size chunking ignores structure: it cuts tables, code and numbered steps in half and strips the heading that gives a passage its meaning.
- Structure-aware chunking splits on headings, sections, list and table boundaries, and prefixes each chunk with its heading path so it still makes sense on its own.
- Retrieve small chunks for precision, then hand the model the surrounding section (parent document retrieval) for context.
Most retrieval-augmented generation (RAG) pipelines start from a tutorial default: split every document into 512-token chunks with 50 tokens of overlap, embed them, done. It works on clean prose. It quietly fails on the documents companies actually have: manuals, policies, contracts, specs and runbooks full of headings, tables and steps.
01What is chunking, and why does it matter so much?
Chunking is the step that splits a document into the units you embed and retrieve. It sounds like plumbing. It decides what the model can ever see.
The retriever returns chunks, not documents. If the answer to a question lives across a chunk boundary, or the chunk that holds the answer lost the heading that explains it, retrieval cannot recover. A larger model does not help, because it never receives the missing context.
02How does fixed-size chunking break?
Take a refund policy with a table of limits:
| Plan | Refund window | Max amount |
|---|---|---|
| Starter | 14 days | $200 |
| Business | 30 days | $2,000 |
A 512-token splitter does not know this is a table. It can end one chunk after “Business | 30 days” and start the next with “$2,000”. Now the chunk that matches “business refund window” has no amount, and the chunk with the amount has no plan name.
The same thing happens to:
- Numbered procedures, where step 4 lands in a different chunk from steps 1 to 3.
- Code samples, split in the middle of a function.
- Claims and their conditions, where “this applies only to EU customers” sits in the previous chunk.
- Headings, which end up attached to the wrong body text or to none.
Overlap softens the cut but does not fix it. It duplicates text without restoring structure.
03What does a structure-aware RAG chunking strategy look like?
Parse the document into its real structure first, then split along it.
- Convert to a structured form. HTML, Markdown, or a parsed PDF with headings, lists and tables identified. This is where most of the effort goes, and it is worth it.
- Split on section boundaries. A section under a heading is the natural unit. Only split further if a section is too long, and then on paragraph or list boundaries.
- Keep atomic blocks whole. Tables, code blocks and numbered procedures are never split. A table too large to embed becomes one chunk per row group, each with the header row repeated.
- Prefix the heading path. Every chunk starts with where it lives, so it means something on its own.
def chunk_document(doc: ParsedDoc, max_tokens: int = 600) -> list[Chunk]:
chunks = []
for section in doc.sections(): # split on headings
path = " > ".join(section.heading_path) # "Billing > Refunds > Limits"
for block in pack_blocks(section.blocks, max_tokens):
chunks.append(Chunk(
text=f"{doc.title}\n{path}\n\n{block.text}",
doc_id=doc.id,
section_id=section.id, # for parent retrieval later
acl=doc.acl, # permissions travel with the chunk
))
return chunks
def pack_blocks(blocks, max_tokens):
"""Group whole blocks (paragraph, list, table, code) up to the budget. Never split a block."""The heading prefix does more than it looks. A chunk that says “Billing > Refunds > Limits” followed by “Business: 30 days, $2,000” matches a question about business refund limits even if the body never repeats the word “refund”.
04How big should a chunk be?
There is no universal number, but there is a useful split of responsibilities:
- Small chunks retrieve well. A focused chunk has a focused embedding, so it ranks higher for a precise question.
- Large context answers well. The model often needs the whole section to answer correctly.
So retrieve small and read large. Embed chunks of a few hundred tokens, and when one is retrieved, give the model its parent section. This is often called parent document retrieval or small-to-big retrieval.
The retriever can only return what a chunk contains. Chunking decides the ceiling; everything after it works under that ceiling.
05How do you know your chunking is better?
Do not judge it by eye. Measure retrieval directly, before any generation:
- Build 50 to 100 real questions with the passage that answers each one.
- Measure recall at k: how often the answering passage appears in the top k results.
- Compare chunking strategies on the same questions and the same embedding model.
Most teams find structure-aware chunking moves recall more than switching embedding models does, at a fraction of the cost. It sits in layer L1 of our stack, data and knowledge, next to permission-aware retrieval: see filter by permission before you rank.