AI Prompt to Chunk Documents for a RAG Pipeline
This AI prompt for chunking documents in a RAG pipeline helps you split long PDFs, manuals, or knowledge-base articles into retrieval-friendly pieces before they're embedded and indexed. It's aimed at developers and data teams building a retrieval-augmented generation (RAG) system who are past the "just split every 500 words" stage and need chunks that actually preserve meaning.
Poor chunking is one of the most common causes of bad RAG answers β a chunk that cuts a table in half, or splits a definition from the term it defines, gives the retriever a fragment the model can't use correctly even if the embedding search works perfectly. The prompt asks the model to reason about document structure (headings, lists, tables) rather than applying a fixed character count blindly, and to flag sections that shouldn't be split at all.
If your source documents are long PDFs or CSVs and you're not sure which sections are even relevant before you chunk them, Context Extractor is worth running first β it pulls the relevant snippets out of a large document via RAG, which can cut down what you need to chunk and index in the first place.
Prompt template
ROLE: You are a document processing assistant preparing content for a retrieval-augmented generation (RAG) system.
CONTEXT:
- Document type: [MANUAL, FAQ, POLICY DOC, API REFERENCE, ETC.]
- Target chunk size: [APPROXIMATE TOKEN OR WORD COUNT]
- Embedding model or vector database being used: [MODEL/DATABASE NAME]
- Document content follows: [PASTE DOCUMENT TEXT]
TASK: Split the document above into chunks suitable for embedding, following these rules:
- Never split a table, code block, or numbered procedure across two chunks
- Keep each chunk to roughly the target size, but prioritize semantic completeness over hitting the size exactly
- Attach the relevant section heading to each chunk as metadata
- Add a short overlap of the previous chunk's last sentence at the start of each new chunk
- Flag any section that should not be split at all, with a one-line reason
OUTPUT FORMAT: Return a numbered list of chunks. For each chunk, show: [Heading], [Chunk Text], [Overlap Note].
Example input
Document type: internal API reference page Target chunk size: ~300 tokens Embedding model: text-embedding-3-large Document content: "## Authentication\nAll requests require a Bearer token in the Authorization header. Tokens expire after 24 hours.\n\n## Rate Limits\n| Plan | Requests/min |\n|------|---------------|\n| Free | 60 |\n| Pro | 600 |\nExceeding the limit returns a 429 response.\n\n## Error Codes\n401: Invalid token. 429: Rate limit exceeded. 500: Server error."
Example output
Chunk 1 Heading: Authentication Chunk Text: "All requests require a Bearer token in the Authorization header. Tokens expire after 24 hours." Overlap Note: None (first chunk)
Chunk 2 Heading: Rate Limits Chunk Text: "Rate limit table β Free plan: 60 requests/min, Pro plan: 600 requests/min. Exceeding the limit returns a 429 response." Overlap Note: Kept the full rate limit table intact rather than splitting rows into separate chunks, since the table only makes sense as a whole.
Chunk 3 Heading: Error Codes Chunk Text: "401: Invalid token. 429: Rate limit exceeded. 500: Server error." Overlap Note: Referenced the Rate Limits section context ("429 matches the rate limit error above") to preserve the connection between sections.
Flagged section: Rate Limits table β should not be split mid-table, since a partial table row is meaningless on its own to the retriever.
When to use it
- When preparing a knowledge base, help center, or internal wiki for embedding into a vector database
- When your current RAG answers are missing context because chunks are cut mid-sentence or mid-table
- Before re-indexing a document set after switching embedding models or chunk sizes
- When source documents mix long prose, tables, and code blocks that need different chunking rules
Best practices
- Chunk by semantic unit (a heading section, a table, a list) before falling back to a fixed token or character limit
- Keep a small overlap between adjacent chunks so context isn't lost at the boundary
- Preserve the section heading or title as metadata attached to each chunk, not just the raw text
- Test retrieval quality on a handful of real user questions after changing your chunking strategy, not just on chunk count
Common mistakes
- Splitting strictly by character count and cutting a table or code block in half
- Losing the document title or section heading, so retrieved chunks have no context on their own
- Using one fixed chunk size for every document type instead of adjusting for prose vs. tables vs. FAQs
- Never testing retrieval on real questions, only checking that chunking ran without errors
FAQs
What's the ideal chunk size for a RAG pipeline?
There's no single right answer β it depends on your embedding model and content type, but most teams start around 200β500 tokens per chunk and adjust based on how well retrieval performs on real questions, not a fixed rule.
Why does chunking strategy matter more than embedding model choice sometimes?
A strong embedding model can't fix a chunk that's missing context β if a chunk is cut mid-table or separated from its heading, the retriever may find it, but the model still won't have enough information to answer correctly.
Should I use overlapping chunks or non-overlapping chunks?
A small overlap (one or two sentences) between adjacent chunks is usually worth the extra storage, since it prevents context from being lost right at a chunk boundary, especially for prose-heavy documents.
How do I know if my chunking strategy is actually working?
The only reliable test is running real user questions through your retrieval pipeline and checking whether the returned chunks contain what's needed to answer them, not just measuring chunk count or size.
What tool can help me pull the right content out of large source documents before I chunk them?
Context Extractor β pulls only the relevant snippets from large PDFs, CSVs, or docs via RAG, which reduces what you need to chunk and index and helps cut hallucinations in the final answers.