Back to cookbook

AI Prompt to Chunk Documents for a RAG Pipeline

2 views Updated
Share

This AI prompt for chunking documents in a RAG pipeline helps you split long PDFs, manuals, or knowledge-base articles into retrieval-friendly pieces before they're embedded and indexed. It's aimed at developers and data teams building a retrieval-augmented generation (RAG) system who are past the "just split every 500 words" stage and need chunks that actually preserve meaning.

Poor chunking is one of the most common causes of bad RAG answers β€” a chunk that cuts a table in half, or splits a definition from the term it defines, gives the retriever a fragment the model can't use correctly even if the embedding search works perfectly. The prompt asks the model to reason about document structure (headings, lists, tables) rather than applying a fixed character count blindly, and to flag sections that shouldn't be split at all.

If your source documents are long PDFs or CSVs and you're not sure which sections are even relevant before you chunk them, Context Extractor is worth running first β€” it pulls the relevant snippets out of a large document via RAG, which can cut down what you need to chunk and index in the first place.

Prompt template

ROLE: You are a document processing assistant preparing content for a retrieval-augmented generation (RAG) system.

CONTEXT:

  • Document type: [MANUAL, FAQ, POLICY DOC, API REFERENCE, ETC.]
  • Target chunk size: [APPROXIMATE TOKEN OR WORD COUNT]
  • Embedding model or vector database being used: [MODEL/DATABASE NAME]
  • Document content follows: [PASTE DOCUMENT TEXT]

TASK: Split the document above into chunks suitable for embedding, following these rules:

  1. Never split a table, code block, or numbered procedure across two chunks
  2. Keep each chunk to roughly the target size, but prioritize semantic completeness over hitting the size exactly
  3. Attach the relevant section heading to each chunk as metadata
  4. Add a short overlap of the previous chunk's last sentence at the start of each new chunk
  5. Flag any section that should not be split at all, with a one-line reason

OUTPUT FORMAT: Return a numbered list of chunks. For each chunk, show: [Heading], [Chunk Text], [Overlap Note].

Example input

Document type: internal API reference page Target chunk size: ~300 tokens Embedding model: text-embedding-3-large Document content: "## Authentication\nAll requests require a Bearer token in the Authorization header. Tokens expire after 24 hours.\n\n## Rate Limits\n| Plan | Requests/min |\n|------|---------------|\n| Free | 60 |\n| Pro | 600 |\nExceeding the limit returns a 429 response.\n\n## Error Codes\n401: Invalid token. 429: Rate limit exceeded. 500: Server error."

Example output

Chunk 1 Heading: Authentication Chunk Text: "All requests require a Bearer token in the Authorization header. Tokens expire after 24 hours." Overlap Note: None (first chunk)

Chunk 2 Heading: Rate Limits Chunk Text: "Rate limit table β€” Free plan: 60 requests/min, Pro plan: 600 requests/min. Exceeding the limit returns a 429 response." Overlap Note: Kept the full rate limit table intact rather than splitting rows into separate chunks, since the table only makes sense as a whole.

Chunk 3 Heading: Error Codes Chunk Text: "401: Invalid token. 429: Rate limit exceeded. 500: Server error." Overlap Note: Referenced the Rate Limits section context ("429 matches the rate limit error above") to preserve the connection between sections.

Flagged section: Rate Limits table β€” should not be split mid-table, since a partial table row is meaningless on its own to the retriever.

When to use it

  • When preparing a knowledge base, help center, or internal wiki for embedding into a vector database
  • When your current RAG answers are missing context because chunks are cut mid-sentence or mid-table
  • Before re-indexing a document set after switching embedding models or chunk sizes
  • When source documents mix long prose, tables, and code blocks that need different chunking rules

Best practices

  • Chunk by semantic unit (a heading section, a table, a list) before falling back to a fixed token or character limit
  • Keep a small overlap between adjacent chunks so context isn't lost at the boundary
  • Preserve the section heading or title as metadata attached to each chunk, not just the raw text
  • Test retrieval quality on a handful of real user questions after changing your chunking strategy, not just on chunk count

Common mistakes

  • Splitting strictly by character count and cutting a table or code block in half
  • Losing the document title or section heading, so retrieved chunks have no context on their own
  • Using one fixed chunk size for every document type instead of adjusting for prose vs. tables vs. FAQs
  • Never testing retrieval on real questions, only checking that chunking ran without errors

FAQs

What's the ideal chunk size for a RAG pipeline?

There's no single right answer β€” it depends on your embedding model and content type, but most teams start around 200–500 tokens per chunk and adjust based on how well retrieval performs on real questions, not a fixed rule.

Why does chunking strategy matter more than embedding model choice sometimes?

A strong embedding model can't fix a chunk that's missing context β€” if a chunk is cut mid-table or separated from its heading, the retriever may find it, but the model still won't have enough information to answer correctly.

Should I use overlapping chunks or non-overlapping chunks?

A small overlap (one or two sentences) between adjacent chunks is usually worth the extra storage, since it prevents context from being lost right at a chunk boundary, especially for prose-heavy documents.

How do I know if my chunking strategy is actually working?

The only reliable test is running real user questions through your retrieval pipeline and checking whether the returned chunks contain what's needed to answer them, not just measuring chunk count or size.

What tool can help me pull the right content out of large source documents before I chunk them?

Context Extractor β€” pulls only the relevant snippets from large PDFs, CSVs, or docs via RAG, which reduces what you need to chunk and index and helps cut hallucinations in the final answers.

Found this prompt useful? Share it.

Share

More in RAG & Knowledge Retrieval

RAG & Knowledge Retrieval

AI Prompt to Make a RAG Assistant Say 'I Don't Know' When Context Is Missing

This is a RAG prompt for stopping a retrieval augmented assistant from guessing when the retrieved documents don't actually contain the answ…

ROLE: You are a question-answering assistant that must only use the provided context to answer questions. You must never use outside knowledge, even if you know the answer.

CONTEXT:
[PASTE RETRIEVED DOCUMENT CHUNKS HERE]

QUESTION: [USER'S QUESTION]

INSTRUCTIONS:
1. Check whether the context above contains enough information to fully answer the question
2. If it does, answer using only that information, and quote or reference the specific part of the context you used
3. If the context only partially answers the question, answer the part you can and explicitly state which part is not covered
4. If the context does not contain relevant information at all, respond exactly with: "[INSUFFICIENT CONTEXT RESPONSE, e.g. I don't have enough information in the provided documents to answer this.]"
5. Do not guess, infer beyond what the context states, or use any knowledge from outside the provided context

CONSTRAINTS:
- [ANY ADDITIONAL CONSTRAINT, e.g. keep answers under 150 words, cite chunk numbers if provided]

OUTPUT FORMAT: A direct answer to the question, or the insufficient-context response, followed by a short reference to which part of the context (if any) was used.
RAG & Knowledge Retrieval

Prompt for Answering Only From Provided Context

This is a grounding prompt for retrieval augmented generation setups: it forces ChatGPT, Claude, or Gemini to answer strictly from the text…

You are a question-answering assistant that only uses the provided context.

Context:
[PASTE RETRIEVED PASSAGES OR DOCUMENTS]

Question: [USER'S QUESTION]

Instructions:
1. Answer using only information found in the context above.
2. Quote or reference the specific part of the context that supports your answer.
3. If the context does not contain the answer, respond exactly: "Not found in the provided context."
4. Do not use outside knowledge, even if you know the answer.

Output format:
Answer: [your answer, or the not-found line]
Source: [the quoted passage, or "n/a"]