ChatGPT Prompt to Rewrite User Queries for Better RAG Retrieval Results
This ChatGPT prompt is for developers building a RAG (retrieval-augmented generation) system who are seeing weak or irrelevant search results even when the answer clearly exists somewhere in the knowledge base. The problem is usually the query itself: a short, vague, or pronoun-heavy user question ("what about the second one?") retrieves poorly against a vector index, no matter how good the underlying embeddings are.
The prompt uses an LLM as a query rewriting step that runs before retrieval: it expands abbreviations, resolves references from prior conversation turns, and restates the question in the kind of specific, keyword-rich language that matches how information is actually phrased in source documents. This is a standard technique in production RAG pipelines, often called query expansion or query reformulation, and it typically improves recall without touching the retrieval index or embedding model at all.
Once you've rewritten a query, the actual retrieval step still has to pull the right passages out of your source material — a tool like Context Extractor is built for that half of the pipeline, pulling only the relevant snippets from large documents based on the query you feed it.
Prompt template
ROLE: You are a query rewriting assistant for a document retrieval system. Your job is to turn a user's question into one or more search queries that will retrieve relevant passages more effectively.
CONTEXT: Conversation history (most recent last): [PASTE_LAST_2-4_CONVERSATION_TURNS]
Current user question: [PASTE_CURRENT_QUESTION]
TASK:
- Resolve any pronouns or references in the current question using the conversation history.
- Expand abbreviations and vague terms into the specific language likely used in source documents about [DESCRIBE_DOMAIN, e.g. "internal HR policy documents"].
- Produce [NUMBER, e.g. 2-3] alternate search queries that capture the same intent from different angles.
CONSTRAINTS:
- Do not introduce new topics, assumptions, or details the user did not raise.
- Keep each rewritten query under 20 words.
- If the original question is already specific and self-contained, return it unchanged as the only query.
OUTPUT FORMAT: Return a JSON array of strings, one per rewritten query, with no additional commentary.
Example input
ROLE: You are a query rewriting assistant for a document retrieval system. Your job is to turn a user's question into one or more search queries that will retrieve relevant passages more effectively.
CONTEXT: Conversation history (most recent last): User: Can employees carry over unused vacation days into the next year? Assistant: Yes, up to 5 days can be carried over, based on current policy. User: What about sick days, does the same rule apply?
Current user question: What about sick days, does the same rule apply?
TASK:
- Resolve any pronouns or references in the current question using the conversation history.
- Expand abbreviations and vague terms into the specific language likely used in source documents about internal HR policy documents.
- Produce 2-3 alternate search queries that capture the same intent from different angles.
CONSTRAINTS:
- Do not introduce new topics, assumptions, or details the user did not raise.
- Keep each rewritten query under 20 words.
- If the original question is already specific and self-contained, return it unchanged as the only query.
OUTPUT FORMAT: Return a JSON array of strings, one per rewritten query, with no additional commentary.
Example output
["sick day carryover policy to next year","can unused sick leave roll over annually","sick leave carryover limit compared to vacation day carryover rule"]
When to use it
- A RAG chatbot returns irrelevant or no results for questions that reference earlier conversation turns ("what's the refund policy for that?")
- Users type short, underspecified queries that don't contain the keywords your source documents actually use
- Retrieval recall is noticeably worse on multi-part or comparative questions than on single, direct ones
- You want to add query expansion to an existing RAG pipeline without retraining or re-embedding anything
Best practices
- Feed the rewriting step the last few turns of conversation, not just the current message, so it can resolve pronouns and references correctly
- Ask the model to produce 2-3 alternate phrasings of the same query rather than a single rewrite, then retrieve against all of them and merge results
- Keep the rewritten query grounded in the original intent — instruct the model not to introduce new topics or assumptions the user didn't raise
- Log the original query alongside the rewritten one during testing so you can see exactly what changed when retrieval quality shifts
Common mistakes
- Letting the rewrite step change the meaning of the question instead of just clarifying its phrasing
- Rewriting every query the same way regardless of whether it's a follow-up or a fresh, already-specific question
- Not testing the rewritten queries against real retrieval results, so a rewrite that reads better doesn't actually retrieve better
- Skipping conversation history entirely, which leaves pronoun-based follow-up questions unresolved and poorly retrieved
FAQs
Why does my RAG chatbot fail on follow-up questions but work fine on first questions?
Follow-up questions often rely on pronouns or implied context from earlier turns ("what about that one"), and a retrieval system searches on the literal text of the query. Without a rewriting step that resolves those references first, the retriever has no way to know what "that one" refers to.
Does query rewriting replace the need for good embeddings or chunking?
No, it's a complementary technique. Query rewriting improves the input to retrieval, while embedding quality and chunking strategy determine how well that retrieval actually performs once it has a clear query to work with.
Should I always generate multiple rewritten queries instead of just one?
Generating 2-3 alternate phrasings and merging the retrieved results generally improves recall more than a single rewrite, especially for ambiguous or broad questions, though it does add a small amount of latency and retrieval cost per query.
Can query rewriting introduce hallucinated details into the search query itself?
Yes, if the rewriting prompt isn't constrained carefully. An unconstrained rewrite step can add assumptions the user never stated, which then biases retrieval toward the wrong passages, so explicitly instructing the model not to introduce new topics is important.
How can I check whether this prompt's constraints are tight enough to avoid that kind of drift?
Prompt Debugger — it's built to catch exactly this kind of risk in RAG and query-reformulation prompts, flagging vague constraints that could let the model introduce unintended assumptions.