Back to cookbook

AI Prompt to Re-Rank Retrieved Passages for RAG Relevance Before Generation

0 views Updated

Make this prompt yours

Share

This is an AI prompt for re-ranking retrieved passages in a RAG pipeline, built for engineers and prompt builders who already have a retriever returning a shortlist of chunks but need a cheap, model-based pass to push the truly relevant ones to the top before they ever reach the generation step. A vector or hybrid retriever optimizes for similarity, not relevance, so the top-k list it returns almost always contains near-miss chunks that happen to share vocabulary with the query but don't actually answer it. Asking an LLM to score and reorder that shortlist against the specific question catches those near-misses before they crowd out the passage that matters.

The prompt works as a relevance re-ranker: it takes the user's query and a numbered list of retrieved chunks, then asks the model to judge each chunk's actual usefulness for answering that query rather than its surface-level topical overlap. It returns a ranked order plus a short justification per chunk, so a downstream step can keep the top N, drop the rest, and hand a tighter, higher-signal context window to the generation call. This matters most in RAG pipelines, documentation Q&A bots, and any retrieval step that feeds a context window a model will reason over.

Because re-ranking is only as good as what gets retrieved in the first place, it pairs naturally with the retrieval step itself. If the chunks being fed into this prompt are already too broad, pulled from oversized documents, or mixed in with irrelevant sections, running them through Context Extractor first to pull only the relevant snippets from the source PDFs or CSVs will give this re-ranking prompt a cleaner shortlist to work from.

Prompt template

Make this prompt yours

prompt-template
358 tokens
ROLE You are a retrieval relevance judge for a RAG (retrieval-augmented generation) pipeline. Your job is to re-rank a list of retrieved passages by how useful each one is for answering a specific query, not by how similar the wording looks. CONTEXT User query: [USER_QUERY] Retrieved passages (numbered, in original retrieval order): 1. [PASSAGE_1] 2. [PASSAGE_2] 3. [PASSAGE_3] [ADD_MORE_PASSAGES_AS_NEEDED] TASK For each passage: 1. Judge whether it contains information that directly helps answer the query above. 2. Assign a relevance score from 0 to 10, where 10 means the passage directly and completely answers the query, 5 means it is partially relevant or provides supporting context, and 0 means it is unrelated despite any surface keyword overlap. 3. Write a one-sentence justification for the score, naming the specific information (or lack of it) that drove your judgment. CONSTRAINTS - Judge relevance based on whether the passage answers the query, not on how many query keywords it shares. - Do not invent information that is not present in the passage text. - If two passages are genuinely equally relevant, you may give them the same score, but explain why in the justification. - If none of the passages are relevant, say so explicitly rather than forcing a ranking. OUTPUT FORMAT Return a ranked list from most to least relevant, in this format: Rank [N] - Passage [ORIGINAL_NUMBER] - Score: [0-10] Justification: [one sentence] After the ranked list, add a line stating the top [NUMBER_OF_PASSAGES_TO_KEEP] passage numbers to pass forward to generation.

Want it sharper? Optimize this prompt with Prompt Optimizer, check it with the Prompt Debugger or shorten it with the Token Optimizer.

Example input

example-input
179 tokens
ROLE You are a retrieval relevance judge for a RAG pipeline. User query: What is our refund policy for digital subscriptions canceled mid-cycle? Retrieved passages: 1. Our refund policy for physical products allows returns within 30 days of delivery, provided the item is unused and in original packaging. 2. Digital subscriptions can be canceled at any time from the account settings page. Cancellation stops future billing but does not generate a prorated refund for the current billing cycle. 3. Subscription tiers include Basic, Pro, and Enterprise, each with different feature sets and monthly pricing. 4. If a customer is charged in error due to a billing system bug, support can issue a manual refund for the erroneous charge within 90 days. TASK Score each passage 0-10 for relevance, with a one-sentence justification. Return a ranked list and the top 2 passages to pass to generation.

When to use it

  • Your vector retriever returns a top-k list where 2-3 chunks look topically related but don't actually contain the answer
  • You're building a RAG chatbot and want a cheap relevance filter between retrieval and generation to cut down on irrelevant context
  • You need to compare results from two different retrieval methods (e.g., dense vs. hybrid search) and want a consistent way to judge which chunks are actually useful
  • You're debugging a RAG pipeline that keeps citing the wrong section and want to isolate whether the problem is retrieval ranking or generation

Best practices

  • Always include the exact user query the chunks were retrieved for, not a paraphrase, since relevance judgments depend on the specific wording and intent of that query
  • Keep the chunk list to a manageable size (roughly 5-15 passages); asking a model to re-rank 50 chunks in one pass degrades judgment quality and increases the chance it skims rather than reads
  • Ask for a one-line justification per ranked chunk, not just a score, so you can spot-check whether the model is ranking on genuine relevance or superficial keyword overlap
  • Before relying on this in production, run the prompt through Prompt Debugger to check for vague relevance criteria or edge cases, like tied scores or empty chunk lists, that could cause inconsistent rankings

Common mistakes

  • Feeding the re-ranker chunks that are already poorly segmented (cut off mid-sentence or missing headers), which makes relevance judgments unreliable no matter how good the prompt is
  • Treating the re-ranker's output as a hard filter without a fallback, so a single bad ranking call drops the one chunk that actually had the answer
  • Omitting the original query and only giving the model the chunks, which forces it to guess at what counts as relevant instead of judging against a concrete question
  • Re-ranking on topic similarity alone instead of asking the model to check whether each chunk contains information that actually answers the query

FAQs

Why does a RAG pipeline need re-ranking if the retriever already returns the top-k results?

Vector and hybrid retrievers rank by embedding similarity or keyword overlap, which measures how related a chunk's wording is to the query, not whether that chunk actually answers it. A re-ranking pass judges each chunk against the specific question, which catches passages that are topically close but factually useless and would otherwise crowd out the one that matters.

Can this re-ranking prompt replace a dedicated re-ranking model?

No. A dedicated re-ranker (like a cross-encoder model) is trained specifically for relevance scoring and runs faster and cheaper at scale. This prompt is useful for prototyping, debugging a pipeline, or low-volume use cases where spinning up a separate re-ranking model isn't worth the engineering cost, but it won't match a purpose-built model's latency or throughput in production.

How many passages should I send to the re-ranker at once?

Keep it to roughly 5-15 passages per call. Sending too many in one pass makes it harder for the model to apply consistent judgment across all of them, and you lose the benefit of careful per-passage reasoning, which is the whole point of re-ranking over raw retrieval order.

What should I do if the re-ranker says none of the retrieved passages are relevant?

Treat that as a signal to fix retrieval, not generation. It usually means the retriever's top-k didn't actually contain the answer, so the fix is adjusting your chunking strategy, expanding the retrieval query, or pulling from a different source, not forcing the generation step to answer from passages the re-ranker already flagged as unhelpful.

Which Cuelara tool can help me clean up the passages before re-ranking them?

Context Extractor — pulls only the relevant snippets from large PDFs or CSVs first, so the re-ranking prompt works with cleaner, better-segmented passages instead of noisy raw chunks. Prompt Debugger — scans the re-ranking prompt itself for vague relevance criteria or edge cases, like tied scores, that could cause inconsistent rankings in production.

Found this prompt useful? Share it.

Share

More in RAG & Knowledge Retrieval

RAG & Knowledge Retrieval

AI Prompt to Make a RAG Assistant Say 'I Don't Know' When Context Is Missing

This is a RAG prompt for stopping a retrieval augmented assistant from guessing when the retrieved documents don't actually contain the answ…

ROLE: You are a question-answering assistant that must only use the provided context to answer questions. You must never use outside knowledge, even if you know the answer.

CONTEXT:
[PASTE RETRIEVED DOCUMENT CHUNKS HERE]

QUESTION: [USER'S QUESTION]

INSTRUCTIONS:
1. Check whether the context above contains enough information to fully answer the question
2. If it does, answer using only that information, and quote or reference the specific part of the context you used
3. If the context only partially answers the question, answer the part you can and explicitly state which part is not covered
4. If the context does not contain relevant information at all, respond exactly with: "[INSUFFICIENT CONTEXT RESPONSE, e.g. I don't have enough information in the provided documents to answer this.]"
5. Do not guess, infer beyond what the context states, or use any knowledge from outside the provided context

CONSTRAINTS:
- [ANY ADDITIONAL CONSTRAINT, e.g. keep answers under 150 words, cite chunk numbers if provided]

OUTPUT FORMAT: A direct answer to the question, or the insufficient-context response, followed by a short reference to which part of the context (if any) was used.

Make this prompt yours

RAG & Knowledge Retrieval

AI Prompt for RAG Systems to Cite Sources and Reduce Hallucinations

This AI prompt for RAG source citation instructs a retrieval augmented generation system to tie every factual claim in its answer back to a…

ROLE
You are a retrieval-augmented assistant that answers strictly from the provided context and cites every claim.

CONTEXT
You will be given a set of retrieved passages, each labeled with a source identifier:
[SOURCE_ID_1]: [PASSAGE_TEXT_1]
[SOURCE_ID_2]: [PASSAGE_TEXT_2]
(additional passages as needed)

USER QUESTION
[USER_QUESTION]

CONSTRAINTS
- Use only information found in the passages above. Do not use outside knowledge, even if you are confident it is correct.
- After every factual sentence, add a citation in the form (Source: [SOURCE_ID]).
- If part of the question cannot be answered from the passages, explicitly state which part is unsupported instead of guessing.
- If no passage is relevant to the question at all, respond with: 'The provided context does not contain information to answer this question.'

OUTPUT FORMAT
1. Direct answer with inline citations after each claim.
2. A short 'Sources used' list naming each cited source identifier once.
3. A 'Not covered' note listing any part of the question left unanswered, or 'None' if fully covered.

Make this prompt yours

RAG & Knowledge Retrieval

Prompt for Answering Only From Provided Context

This is a grounding prompt for retrieval augmented generation setups: it forces ChatGPT, Claude, or Gemini to answer strictly from the text…

You are a question-answering assistant that only uses the provided context.

Context:
[PASTE RETRIEVED PASSAGES OR DOCUMENTS]

Question: [USER'S QUESTION]

Instructions:
1. Answer using only information found in the context above.
2. Quote or reference the specific part of the context that supports your answer.
3. If the context does not contain the answer, respond exactly: "Not found in the provided context."
4. Do not use outside knowledge, even if you know the answer.

Output format:
Answer: [your answer, or the not-found line]
Source: [the quoted passage, or "n/a"]

Make this prompt yours