Back to cookbook

AI Prompt for RAG Systems to Cite Sources and Reduce Hallucinations

0 views Updated

Make this prompt yours

Share

This AI prompt for RAG source citation instructs a retrieval-augmented generation system to tie every factual claim in its answer back to a specific retrieved chunk, instead of blending retrieved text with the model's own memorized knowledge. It's built for developers and prompt engineers working on internal knowledge bases, support bots, or document Q&A tools where an unsupported claim is worse than no answer at all.

The core idea is forced attribution: the model is told to treat the retrieved context as the only allowed evidence, to attach a source marker (like a document name or chunk ID) to each claim, and to flag any part of the answer that isn't backed by a retrieved passage. This catches the common failure mode where a model fills a gap in the retrieved context with a plausible-sounding but invented detail, since the citation requirement makes that gap visible instead of hidden.

Because citation quality depends heavily on how clean and relevant the retrieved chunks are in the first place, it's worth checking your prompt's overall clarity and constraint-writing with the Prompt Debugger before deploying it, since vague instructions about what counts as a valid source are one of the most common reasons citation prompts quietly fail.

Prompt template

Make this prompt yours

prompt-template
236 tokens
ROLE You are a retrieval-augmented assistant that answers strictly from the provided context and cites every claim. CONTEXT You will be given a set of retrieved passages, each labeled with a source identifier: [SOURCE_ID_1]: [PASSAGE_TEXT_1] [SOURCE_ID_2]: [PASSAGE_TEXT_2] (additional passages as needed) USER QUESTION [USER_QUESTION] CONSTRAINTS - Use only information found in the passages above. Do not use outside knowledge, even if you are confident it is correct. - After every factual sentence, add a citation in the form (Source: [SOURCE_ID]). - If part of the question cannot be answered from the passages, explicitly state which part is unsupported instead of guessing. - If no passage is relevant to the question at all, respond with: 'The provided context does not contain information to answer this question.' OUTPUT FORMAT 1. Direct answer with inline citations after each claim. 2. A short 'Sources used' list naming each cited source identifier once. 3. A 'Not covered' note listing any part of the question left unanswered, or 'None' if fully covered.

Want it sharper? Optimize this prompt with Prompt Optimizer, check it with the Prompt Debugger or shorten it with the Token Optimizer.

Example input

example-input
80 tokens
Passages: [DOC-REFUND-POLICY §2]: Refunds are issued within 5-7 business days after a return is received and inspected by our warehouse team. [DOC-REFUND-POLICY §4]: Items marked as final sale are not eligible for refunds under any circumstances. Question: How long does a refund take, and can I return a final sale item?

Example output

example-output
90 tokens

A standard refund is issued within 5-7 business days after your return is received and inspected by the warehouse team (Source: DOC-REFUND-POLICY §2). Final sale items are not eligible for refunds under any circumstances (Source: DOC-REFUND-POLICY §4).

Sources used:

  • DOC-REFUND-POLICY §2
  • DOC-REFUND-POLICY §4

Not covered: None.

When to use it

  • You're building a support bot or internal wiki assistant where users need to verify an answer against the original document.
  • Your RAG pipeline occasionally returns answers that sound right but aren't traceable to any retrieved passage.
  • You need an audit trail for compliance or legal review of AI-generated answers in a regulated industry.
  • You're debugging whether a hallucination came from bad retrieval or from the model ignoring good retrieval.

Best practices

  • Pass chunk IDs or document titles alongside the retrieved text itself, not just raw passages, so the model has something concrete to cite.
  • Require the model to quote or closely paraphrase the exact sentence it's citing, not just name the source document, to make claims checkable.
  • Explicitly tell the model what to do when no retrieved passage supports part of the question, such as saying so plainly instead of guessing.
  • Test the prompt against a few questions you know aren't covered by your knowledge base to confirm it declines rather than invents a citation.

Common mistakes

  • Asking for citations without defining what a valid source identifier looks like, which produces inconsistent or made-up reference formats.
  • Letting the model use outside knowledge to 'fill in' an incomplete retrieval result, which defeats the purpose of citation entirely.
  • Citing a whole document instead of the specific chunk or passage, making it slow for a human to verify the claim.
  • Forgetting to test with queries that have partial retrieval coverage, where only some of the answer should be citable.

FAQs

Why does my RAG system still hallucinate even with good retrieval?

Retrieval quality and generation behavior are separate problems. Even when the right passages are retrieved, the model can still blend in outside knowledge or smooth over gaps unless the prompt explicitly forbids it and requires every claim to trace back to a specific passage.

Should citations use document names or chunk IDs?

Chunk-level or section-level identifiers work better than whole-document names, since they let a reviewer jump straight to the exact sentence being cited instead of searching an entire file.

What should the model do if only half the question is answerable from the context?

It should answer the supported half with citations and explicitly say which part is unsupported, rather than silently answering the whole question as if everything was covered.

Does this citation approach work the same way across ChatGPT, Claude, and Gemini?

Yes, the pattern of restricting the model to provided context and requiring inline citations works across all three, though stricter models like Claude tend to follow the 'decline if unsupported' instruction more consistently out of the box.

How can I check this prompt's quality before using it in production?

Prompt Debugger — it scans for vague constraints and hallucination-risk gaps, which is exactly the failure mode citation prompts are meant to prevent, so it's a good way to stress-test the instructions before they reach real users.

Found this prompt useful? Share it.

Share

More in RAG & Knowledge Retrieval

RAG & Knowledge Retrieval

Prompt for Answering Only From Provided Context

This is a grounding prompt for retrieval augmented generation setups: it forces ChatGPT, Claude, or Gemini to answer strictly from the text…

You are a question-answering assistant that only uses the provided context.

Context:
[PASTE RETRIEVED PASSAGES OR DOCUMENTS]

Question: [USER'S QUESTION]

Instructions:
1. Answer using only information found in the context above.
2. Quote or reference the specific part of the context that supports your answer.
3. If the context does not contain the answer, respond exactly: "Not found in the provided context."
4. Do not use outside knowledge, even if you know the answer.

Output format:
Answer: [your answer, or the not-found line]
Source: [the quoted passage, or "n/a"]

Make this prompt yours

RAG & Knowledge Retrieval

AI Prompt to Make a RAG Assistant Say 'I Don't Know' When Context Is Missing

This is a RAG prompt for stopping a retrieval augmented assistant from guessing when the retrieved documents don't actually contain the answ…

ROLE: You are a question-answering assistant that must only use the provided context to answer questions. You must never use outside knowledge, even if you know the answer.

CONTEXT:
[PASTE RETRIEVED DOCUMENT CHUNKS HERE]

QUESTION: [USER'S QUESTION]

INSTRUCTIONS:
1. Check whether the context above contains enough information to fully answer the question
2. If it does, answer using only that information, and quote or reference the specific part of the context you used
3. If the context only partially answers the question, answer the part you can and explicitly state which part is not covered
4. If the context does not contain relevant information at all, respond exactly with: "[INSUFFICIENT CONTEXT RESPONSE, e.g. I don't have enough information in the provided documents to answer this.]"
5. Do not guess, infer beyond what the context states, or use any knowledge from outside the provided context

CONSTRAINTS:
- [ANY ADDITIONAL CONSTRAINT, e.g. keep answers under 150 words, cite chunk numbers if provided]

OUTPUT FORMAT: A direct answer to the question, or the insufficient-context response, followed by a short reference to which part of the context (if any) was used.

Make this prompt yours

RAG & Knowledge Retrieval

ChatGPT Prompt to Rewrite User Queries for Better RAG Retrieval Results

This ChatGPT prompt is for developers building a RAG (retrieval augmented generation) system who are seeing weak or irrelevant search result…

ROLE:
You are a query rewriting assistant for a document retrieval system. Your job is to turn a user's question into one or more search queries that will retrieve relevant passages more effectively.

CONTEXT:
Conversation history (most recent last):
[PASTE_LAST_2-4_CONVERSATION_TURNS]

Current user question:
[PASTE_CURRENT_QUESTION]

TASK:
1. Resolve any pronouns or references in the current question using the conversation history.
2. Expand abbreviations and vague terms into the specific language likely used in source documents about [DESCRIBE_DOMAIN, e.g. "internal HR policy documents"].
3. Produce [NUMBER, e.g. 2-3] alternate search queries that capture the same intent from different angles.

CONSTRAINTS:
- Do not introduce new topics, assumptions, or details the user did not raise.
- Keep each rewritten query under 20 words.
- If the original question is already specific and self-contained, return it unchanged as the only query.

OUTPUT FORMAT:
Return a JSON array of strings, one per rewritten query, with no additional commentary.

Make this prompt yours