Techniques & Methods

Chunking in plain English.

Also known as: document chunking,text splitting,chunk size

The one-sentence version

Splitting long documents into smaller pieces before embedding them, so a retrieval system can find and return the relevant passage rather than the whole file.

Chunking is the step in a retrieval-augmented generation (RAG) pipeline where documents are cut into pieces small enough to embed and retrieve individually. A 200-page manual is useless as a single vector; split into 500-token chunks, the right paragraph can be found and handed to the model. How you chunk matters more than most people expect. Fixed-size chunks are simple but can cut a sentence in half; overlapping chunks reduce that risk at the cost of duplication; semantic or structural chunking splits on headings, paragraphs, or topic shifts so each piece is self-contained. Smaller chunks retrieve precisely but lose context; larger chunks keep context but retrieve noisily. Frameworks such as LlamaIndex and LangChain ship several strategies, and parsers like LlamaParse chunk by document structure, keeping tables intact. A common upgrade is to attach metadata (title, section, page) to each chunk and to include a short summary of the parent document, so the model knows where a passage came from.

Read the full guide