Back to cookbook

ChatGPT Prompt to Rewrite User Queries for Better RAG Retrieval Results

2 views Updated
Share

This ChatGPT prompt is for developers building a RAG (retrieval-augmented generation) system who are seeing weak or irrelevant search results even when the answer clearly exists somewhere in the knowledge base. The problem is usually the query itself: a short, vague, or pronoun-heavy user question ("what about the second one?") retrieves poorly against a vector index, no matter how good the underlying embeddings are.

The prompt uses an LLM as a query rewriting step that runs before retrieval: it expands abbreviations, resolves references from prior conversation turns, and restates the question in the kind of specific, keyword-rich language that matches how information is actually phrased in source documents. This is a standard technique in production RAG pipelines, often called query expansion or query reformulation, and it typically improves recall without touching the retrieval index or embedding model at all.

Once you've rewritten a query, the actual retrieval step still has to pull the right passages out of your source material — a tool like Context Extractor is built for that half of the pipeline, pulling only the relevant snippets from large documents based on the query you feed it.

Prompt template

ROLE: You are a query rewriting assistant for a document retrieval system. Your job is to turn a user's question into one or more search queries that will retrieve relevant passages more effectively.

CONTEXT: Conversation history (most recent last): [PASTE_LAST_2-4_CONVERSATION_TURNS]

Current user question: [PASTE_CURRENT_QUESTION]

TASK:

  1. Resolve any pronouns or references in the current question using the conversation history.
  2. Expand abbreviations and vague terms into the specific language likely used in source documents about [DESCRIBE_DOMAIN, e.g. "internal HR policy documents"].
  3. Produce [NUMBER, e.g. 2-3] alternate search queries that capture the same intent from different angles.

CONSTRAINTS:

  • Do not introduce new topics, assumptions, or details the user did not raise.
  • Keep each rewritten query under 20 words.
  • If the original question is already specific and self-contained, return it unchanged as the only query.

OUTPUT FORMAT: Return a JSON array of strings, one per rewritten query, with no additional commentary.

Example input

ROLE: You are a query rewriting assistant for a document retrieval system. Your job is to turn a user's question into one or more search queries that will retrieve relevant passages more effectively.

CONTEXT: Conversation history (most recent last): User: Can employees carry over unused vacation days into the next year? Assistant: Yes, up to 5 days can be carried over, based on current policy. User: What about sick days, does the same rule apply?

Current user question: What about sick days, does the same rule apply?

TASK:

  1. Resolve any pronouns or references in the current question using the conversation history.
  2. Expand abbreviations and vague terms into the specific language likely used in source documents about internal HR policy documents.
  3. Produce 2-3 alternate search queries that capture the same intent from different angles.

CONSTRAINTS:

  • Do not introduce new topics, assumptions, or details the user did not raise.
  • Keep each rewritten query under 20 words.
  • If the original question is already specific and self-contained, return it unchanged as the only query.

OUTPUT FORMAT: Return a JSON array of strings, one per rewritten query, with no additional commentary.

Example output

[
"sick day carryover policy to next year",
"can unused sick leave roll over annually",
"sick leave carryover limit compared to vacation day carryover rule"
]

When to use it

  • A RAG chatbot returns irrelevant or no results for questions that reference earlier conversation turns ("what's the refund policy for that?")
  • Users type short, underspecified queries that don't contain the keywords your source documents actually use
  • Retrieval recall is noticeably worse on multi-part or comparative questions than on single, direct ones
  • You want to add query expansion to an existing RAG pipeline without retraining or re-embedding anything

Best practices

  • Feed the rewriting step the last few turns of conversation, not just the current message, so it can resolve pronouns and references correctly
  • Ask the model to produce 2-3 alternate phrasings of the same query rather than a single rewrite, then retrieve against all of them and merge results
  • Keep the rewritten query grounded in the original intent — instruct the model not to introduce new topics or assumptions the user didn't raise
  • Log the original query alongside the rewritten one during testing so you can see exactly what changed when retrieval quality shifts

Common mistakes

  • Letting the rewrite step change the meaning of the question instead of just clarifying its phrasing
  • Rewriting every query the same way regardless of whether it's a follow-up or a fresh, already-specific question
  • Not testing the rewritten queries against real retrieval results, so a rewrite that reads better doesn't actually retrieve better
  • Skipping conversation history entirely, which leaves pronoun-based follow-up questions unresolved and poorly retrieved

FAQs

Why does my RAG chatbot fail on follow-up questions but work fine on first questions?

Follow-up questions often rely on pronouns or implied context from earlier turns ("what about that one"), and a retrieval system searches on the literal text of the query. Without a rewriting step that resolves those references first, the retriever has no way to know what "that one" refers to.

Does query rewriting replace the need for good embeddings or chunking?

No, it's a complementary technique. Query rewriting improves the input to retrieval, while embedding quality and chunking strategy determine how well that retrieval actually performs once it has a clear query to work with.

Should I always generate multiple rewritten queries instead of just one?

Generating 2-3 alternate phrasings and merging the retrieved results generally improves recall more than a single rewrite, especially for ambiguous or broad questions, though it does add a small amount of latency and retrieval cost per query.

Can query rewriting introduce hallucinated details into the search query itself?

Yes, if the rewriting prompt isn't constrained carefully. An unconstrained rewrite step can add assumptions the user never stated, which then biases retrieval toward the wrong passages, so explicitly instructing the model not to introduce new topics is important.

How can I check whether this prompt's constraints are tight enough to avoid that kind of drift?

Prompt Debugger — it's built to catch exactly this kind of risk in RAG and query-reformulation prompts, flagging vague constraints that could let the model introduce unintended assumptions.

Found this prompt useful? Share it.

Share

More in RAG & Knowledge Retrieval

RAG & Knowledge Retrieval

AI Prompt to Chunk Documents for a RAG Pipeline

This AI prompt for chunking documents in a RAG pipeline helps you split long PDFs, manuals, or knowledge base articles into retrieval friend…

ROLE: You are a document processing assistant preparing content for a retrieval-augmented generation (RAG) system.

CONTEXT:
- Document type: [MANUAL, FAQ, POLICY DOC, API REFERENCE, ETC.]
- Target chunk size: [APPROXIMATE TOKEN OR WORD COUNT]
- Embedding model or vector database being used: [MODEL/DATABASE NAME]
- Document content follows: [PASTE DOCUMENT TEXT]

TASK:
Split the document above into chunks suitable for embedding, following these rules:
1. Never split a table, code block, or numbered procedure across two chunks
2. Keep each chunk to roughly the target size, but prioritize semantic completeness over hitting the size exactly
3. Attach the relevant section heading to each chunk as metadata
4. Add a short overlap of the previous chunk's last sentence at the start of each new chunk
5. Flag any section that should not be split at all, with a one-line reason

OUTPUT FORMAT:
Return a numbered list of chunks. For each chunk, show: [Heading], [Chunk Text], [Overlap Note].
RAG & Knowledge Retrieval

Prompt for Answering Only From Provided Context

This is a grounding prompt for retrieval augmented generation setups: it forces ChatGPT, Claude, or Gemini to answer strictly from the text…

You are a question-answering assistant that only uses the provided context.

Context:
[PASTE RETRIEVED PASSAGES OR DOCUMENTS]

Question: [USER'S QUESTION]

Instructions:
1. Answer using only information found in the context above.
2. Quote or reference the specific part of the context that supports your answer.
3. If the context does not contain the answer, respond exactly: "Not found in the provided context."
4. Do not use outside knowledge, even if you know the answer.

Output format:
Answer: [your answer, or the not-found line]
Source: [the quoted passage, or "n/a"]
RAG & Knowledge Retrieval

AI Prompt to Make a RAG Assistant Say 'I Don't Know' When Context Is Missing

This is a RAG prompt for stopping a retrieval augmented assistant from guessing when the retrieved documents don't actually contain the answ…

ROLE: You are a question-answering assistant that must only use the provided context to answer questions. You must never use outside knowledge, even if you know the answer.

CONTEXT:
[PASTE RETRIEVED DOCUMENT CHUNKS HERE]

QUESTION: [USER'S QUESTION]

INSTRUCTIONS:
1. Check whether the context above contains enough information to fully answer the question
2. If it does, answer using only that information, and quote or reference the specific part of the context you used
3. If the context only partially answers the question, answer the part you can and explicitly state which part is not covered
4. If the context does not contain relevant information at all, respond exactly with: "[INSUFFICIENT CONTEXT RESPONSE, e.g. I don't have enough information in the provided documents to answer this.]"
5. Do not guess, infer beyond what the context states, or use any knowledge from outside the provided context

CONSTRAINTS:
- [ANY ADDITIONAL CONSTRAINT, e.g. keep answers under 150 words, cite chunk numbers if provided]

OUTPUT FORMAT: A direct answer to the question, or the insufficient-context response, followed by a short reference to which part of the context (if any) was used.