← Back to blog
RAG Architectures·October 2, 2026·9 min read

Fixed-Size, Semantic, or Late Chunking: Where RAG Retrieval Actually Breaks

Most RAG recall failures are decided at the chunking step, long before retrieval or reranking — here is how fixed-size, semantic, and late chunking actually differ, and when each one silently loses the fact you need.

Split a 40-page onboarding manual into 512-token chunks with 50-token overlap, embed each chunk, and ask "what's the refund policy for enterprise customers?" Half the time you'll get the consumer refund policy instead, because the sentence that says "the following applies only to Enterprise tier" landed in the previous chunk and the one with the actual dollar figures landed in the next one. The retriever did its job. The chunker didn't.

Most RAG post-mortems blame the embedding model, the vector database, or the LLM's reading comprehension. In practice, a huge fraction of retrieval failures are chunking failures — decided before a single embedding is computed, and silently constraining everything downstream. This is the layer that gets the least design attention and causes the most damage.

Why chunking is a retrieval problem, not a preprocessing step

A chunk is the atomic unit your retriever can return. If the fact a query needs is split across two chunks, no amount of reranking or query rewriting recovers it — the retriever can hand the LLM chunk A or chunk B, never the synthesis of both. Chunking decides the ceiling on recall before retrieval even starts.

It also determines what gets embedded. An embedding model maps a chunk's entire content to one vector (or, for late interaction models, a set of token vectors scoped to that chunk). A chunk that mixes two topics — common when splitting hits an arbitrary token count mid-paragraph — produces a smeared embedding that matches neither topic well. A chunk that's too small loses the surrounding context that disambiguates it: "it processes refunds within 5 business days" is a different sentence when the previous chunk said "Enterprise" versus "Free tier," but a vector doesn't carry that neighbor.

This is why the three chunking strategies below are not formatting choices — they're architectural decisions about where meaning boundaries live and whether the embedding step can see them.

1. Fixed-size chunking raw document text (ignores sentence/section boundaries) equal-size token windows, fixed overlap ⚠ cuts mid-sentence, severs local context, embeds boundary noise

2. Semantic chunking sentence embeddings → cosine distance between neighbors ✓ variable-size chunks split at topic shifts, not token counts ⚠ still embeds each chunk in isolation — no document-level context

3. Late chunking whole document → long-context embedding model → per-token vectors

Fixed-size chunking: the default that quietly sets a ceiling

Fixed-size chunking is what you get from RecursiveCharacterTextSplitter or similar out of the box: pick a token budget (commonly 256–1024), a stride (overlap, usually 10–20% of the chunk size), and slice.

def fixed_size_chunks(tokens, size=512, overlap=64):
    chunks = []
    start = 0
    while start < len(tokens):
        end = min(start + size, len(tokens))
        chunks.append(tokens[start:end])
        start += size - overlap
    return chunks

It's cheap — no model calls, just a tokenizer — and it's predictable, which matters when you're provisioning a vector index and need to estimate storage and latency up front. For homogeneous prose (news articles, narrative text) with generous overlap, it's often good enough.

The failure mode is structural: the splitter has no idea where a sentence, a bullet point, a table row, or a code block ends. Overlap is a band-aid, not a fix — it increases the odds that a chunk contains the needed sentence, at the cost of near-duplicate content bloating the index and diluting ranking signal (two near-identical chunks competing for the same top-k slot). Token-count boundaries are also brutal to structured content: a chunker that doesn't know about Markdown headers or HTML tags will split a table in half and leave you with two chunks that are both unparseable fragments.

Semantic chunking: let the embeddings draw the boundaries

Semantic chunking inverts the order of operations: embed first (at the sentence level), then decide where to cut based on where the meaning actually shifts.

The common implementation (this is roughly what LlamaIndex's SemanticSplitterNodeParser and most "semantic chunking" libraries do): split the document into sentences, embed each one individually, then walk through consecutive sentence pairs computing cosine distance between their embeddings. Distances are converted to a distribution, and a breakpoint is inserted wherever the distance crosses a percentile threshold (often the 95th percentile of all pairwise distances in the document) — i.e., wherever consecutive sentences are more dissimilar than usual for this document.

def semantic_breakpoints(sentence_embeddings, percentile=95):
    distances = [
        cosine_distance(sentence_embeddings[i], sentence_embeddings[i + 1])
        for i in range(len(sentence_embeddings) - 1)
    ]
    threshold = percentile_of(distances, percentile)
    return [i for i, d in enumerate(distances) if d > threshold]

This produces variable-length chunks that track topic shifts instead of token counts — a tight three-sentence chunk where the topic changes quickly, a long twelve-sentence chunk where one idea is developed at length. For documents with real internal structure (technical docs, long-form reports, meeting transcripts), this measurably improves recall over fixed-size windows because each chunk is more likely to be topically coherent.

The cost is real: you're now making one embedding call per sentence just to decide where to cut, then re-embedding the resulting chunks for the actual index — roughly 1.5-2x the embedding cost of fixed-size chunking, plus the latency of a sequential distance-threshold pass. And it inherits a subtler problem from fixed-size chunking: each chunk is still embedded in isolation. A chunk that reads "it increased 40% year over year" is semantically coherent on its own, but the embedding has no way to know "it" refers to the thing described three chunks earlier. Semantic chunking fixes where you cut; it does nothing for what's lost at the cut.

Late chunking: embed the whole document, then cut

Late chunking (the technique Jina AI published in 2024, since adopted by several embedding providers) attacks the isolation problem directly by reordering the pipeline again: embed the entire document in one pass through a long-context embedding model, producing a per-token (or per-sentence) contextualized vector sequence — the same kind of hidden states a transformer produces internally, before any pooling collapses them into a single vector. Chunk boundaries are then applied after embedding, by mean-pooling the token vectors within each chunk span.

The key difference: because every token's vector was produced by a model that attended over the whole document, the chunk-level vector for "it increased 40% year over year" is computed from token representations that already encode what "it" resolved to. You get chunk-sized retrieval units with document-scale context baked into each vector — without having to literally repeat the document's context as text inside every chunk (the hacky alternative some teams use: prepending a summary to every chunk, which bloats token counts and still only approximates what late chunking gets for free from attention).

The constraint is the embedding model's context window: late chunking only works if the whole document fits in one forward pass, which rules it out for very long documents unless the model supports long context (8k–32k tokens is increasingly standard for current embedding models, which covers most individual documents but not full books or massive PDFs without a fallback split-then-late-chunk-per-section approach).

Structure-aware chunking: the boring fix that solves the common case

Before reaching for semantic or late chunking, it's worth asking whether the real problem is simpler: most "fixed-size chunking is bad" horror stories involve content that already has explicit structure — Markdown headers, HTML sections, code blocks, table rows — that a naive character splitter is ignoring. A structure-aware splitter (chunk on Markdown header boundaries first, then fixed-size within each section; never split inside a fenced code block or table) fixes a large share of real-world chunking complaints for a fraction of the cost of semantic or late chunking, because it doesn't need any embedding calls to find boundaries — the document format already told you where they are.

In practice, teams get the best cost/benefit by layering this under one of the other strategies: structure-aware splitting first (respect headers, don't break code blocks or tables), then fixed-size or semantic chunking within each structural section.

Comparing the strategies

StrategyBoundary signalExtra embedding costPreserves cross-chunk contextBest forTypical failure mode
Fixed-sizeToken/char countNoneNoHomogeneous prose, cost-sensitive pipelinesMid-sentence cuts, duplicate near-chunks from overlap
Structure-awareDocument format (headers, tags, code fences)NoneNoMarkdown/HTML docs, API references, code-heavy contentFlat unstructured text (no signal to use)
SemanticSentence-embedding distance~1.5–2x (embed sentences, then chunks)NoLong-form reports, transcripts, docs with real topic shiftsStill embeds each chunk in isolation
Late chunkingFull-document attention, pooled post-hoc~1x (single pass, no re-embedding)YesDocuments with pronouns/references spanning sections, within context windowBounded by embedding model's max context length

How to actually decide

Don't pick a chunking strategy by vibes — measure it the same way you'd measure a retriever change: build a small labeled set of (query, expected-source-passage) pairs from your actual corpus, and compute recall@k for each chunking strategy at the chunk sizes you're considering. The differences are often larger than the differences between embedding models, which is where most teams spend their tuning time instead.

A few defaults that hold up across most corpora: start with structure-aware splitting if your documents have any markup at all — it's free and fixes the worst offenders. Use overlap (10–20% of chunk size) on fixed-size chunking as a baseline, not a solution. Reach for semantic chunking when your documents are long-form and topic shifts are frequent but irregular. Reach for late chunking specifically when your failure mode is referential — pronouns, "the aforementioned," section-spanning comparisons — and your documents fit the embedding model's context window; it's solving a different problem than semantic chunking, not a strictly better version of it.

The uncomfortable takeaway: if your RAG system is hallucinating or missing facts that are clearly present in the source documents, check the chunk boundaries before you touch the prompt, the reranker, or the model. The fact was probably never retrievable in the first place.

#rag#chunking#embeddings#vector-search#late-chunking#semantic-chunking

Related reading

RAG Architectures
Naive RAG, Agentic RAG, and GraphRAG: What Actually Changes Architecturally
Agent Memory
Agent Memory Isn't RAG: Why Vector Retrieval Falls Apart for Stateful Agents