AGENSPHERE/ JOURNAL
← JOURNAL
J-007NOTE4 MIN READ

Chunk by document structure, not by token count

Fixed 512-token chunks split tables mid-row and separate claims from the context that makes them true. A RAG chunking strategy that follows headings, sections and lists is the cheapest retrieval improvement most systems never make.

IN SHORT
  • Chunking is how a document is split into pieces before embedding. The retriever can only return what a chunk contains, so a bad split caps answer quality no matter which model you use.
  • Fixed-size chunking ignores structure: it cuts tables, code and numbered steps in half and strips the heading that gives a passage its meaning.
  • Structure-aware chunking splits on headings, sections, list and table boundaries, and prefixes each chunk with its heading path so it still makes sense on its own.
  • Retrieve small chunks for precision, then hand the model the surrounding section (parent document retrieval) for context.

Most retrieval-augmented generation (RAG) pipelines start from a tutorial default: split every document into 512-token chunks with 50 tokens of overlap, embed them, done. It works on clean prose. It quietly fails on the documents companies actually have: manuals, policies, contracts, specs and runbooks full of headings, tables and steps.

01What is chunking, and why does it matter so much?

Chunking is the step that splits a document into the units you embed and retrieve. It sounds like plumbing. It decides what the model can ever see.

The retriever returns chunks, not documents. If the answer to a question lives across a chunk boundary, or the chunk that holds the answer lost the heading that explains it, retrieval cannot recover. A larger model does not help, because it never receives the missing context.

02How does fixed-size chunking break?

Take a refund policy with a table of limits:

PlanRefund windowMax amount
Starter14 days$200
Business30 days$2,000

A 512-token splitter does not know this is a table. It can end one chunk after “Business | 30 days” and start the next with “$2,000”. Now the chunk that matches “business refund window” has no amount, and the chunk with the amount has no plan name.

The same thing happens to:

  • Numbered procedures, where step 4 lands in a different chunk from steps 1 to 3.
  • Code samples, split in the middle of a function.
  • Claims and their conditions, where “this applies only to EU customers” sits in the previous chunk.
  • Headings, which end up attached to the wrong body text or to none.

Overlap softens the cut but does not fix it. It duplicates text without restoring structure.

03What does a structure-aware RAG chunking strategy look like?

Parse the document into its real structure first, then split along it.

  1. Convert to a structured form. HTML, Markdown, or a parsed PDF with headings, lists and tables identified. This is where most of the effort goes, and it is worth it.
  2. Split on section boundaries. A section under a heading is the natural unit. Only split further if a section is too long, and then on paragraph or list boundaries.
  3. Keep atomic blocks whole. Tables, code blocks and numbered procedures are never split. A table too large to embed becomes one chunk per row group, each with the header row repeated.
  4. Prefix the heading path. Every chunk starts with where it lives, so it means something on its own.
ingest/chunk.py
def chunk_document(doc: ParsedDoc, max_tokens: int = 600) -> list[Chunk]:
    chunks = []
    for section in doc.sections():                       # split on headings
        path = " > ".join(section.heading_path)          # "Billing > Refunds > Limits"
        for block in pack_blocks(section.blocks, max_tokens):
            chunks.append(Chunk(
                text=f"{doc.title}\n{path}\n\n{block.text}",
                doc_id=doc.id,
                section_id=section.id,                   # for parent retrieval later
                acl=doc.acl,                             # permissions travel with the chunk
            ))
    return chunks

def pack_blocks(blocks, max_tokens):
    """Group whole blocks (paragraph, list, table, code) up to the budget. Never split a block."""

The heading prefix does more than it looks. A chunk that says “Billing > Refunds > Limits” followed by “Business: 30 days, $2,000” matches a question about business refund limits even if the body never repeats the word “refund”.

04How big should a chunk be?

There is no universal number, but there is a useful split of responsibilities:

  • Small chunks retrieve well. A focused chunk has a focused embedding, so it ranks higher for a precise question.
  • Large context answers well. The model often needs the whole section to answer correctly.

So retrieve small and read large. Embed chunks of a few hundred tokens, and when one is retrieved, give the model its parent section. This is often called parent document retrieval or small-to-big retrieval.

The retriever can only return what a chunk contains. Chunking decides the ceiling; everything after it works under that ceiling.

05How do you know your chunking is better?

Do not judge it by eye. Measure retrieval directly, before any generation:

  • Build 50 to 100 real questions with the passage that answers each one.
  • Measure recall at k: how often the answering passage appears in the top k results.
  • Compare chunking strategies on the same questions and the same embedding model.

Most teams find structure-aware chunking moves recall more than switching embedding models does, at a fraction of the cost. It sits in layer L1 of our stack, data and knowledge, next to permission-aware retrieval: see filter by permission before you rank.

Building something like this?DESCRIBE A SYSTEM →