AI System Prompt to Keep a Chatbot On-Topic and Resist Prompt Injection
This system prompt is built to keep a custom chatbot focused on its intended job and harder to knock off course with prompt injection attempts embedded in user messages, pasted documents, or tool output. It's meant for anyone deploying a support bot, internal assistant, or customer-facing GPT who has noticed users (deliberately or not) trying to get the model to ignore its instructions, adopt a different persona, or answer questions far outside its assigned scope.
The core idea is to give the model an explicit scope boundary and a standing instruction to treat any text that tries to override its role as untrusted content rather than a new command, even when that text claims to come from a developer or system message. It also gives the model a consistent, polite way to decline and redirect instead of silently complying or refusing abruptly, which keeps the experience smooth for legitimate users while closing off the most common jailbreak patterns.
Because this kind of guardrail prompt lives or dies on edge cases you didn't think of, it's worth running the draft through a prompt debugger before shipping it, since that kind of check is specifically built to surface the loopholes and vague constraints that let injected instructions slip through.
Prompt template
Want it sharper? Optimize this prompt with Prompt Optimizer, check it with the Prompt Debugger or shorten it with the Token Optimizer.
Example input
Example output
I can't ignore my instructions or take on a different role here, so I'll stick with helping you as RentHelper. I also can't write scripts or scrape listings since that's outside what I handle.
If you have a maintenance request, a question about your rent due date, or want to check your lease renewal window, I'm glad to help with any of those.
When to use it
- You're deploying a customer support or internal bot and need it to refuse off-topic requests without sounding broken
- Users have pasted text into the chat that tries to make the model ignore prior instructions or reveal its system prompt
- You're building a tool that passes untrusted content (web pages, documents, emails) into the model's context
- You need the same guardrail behavior to hold consistently across many conversations, not just a one-off response
Best practices
- State the bot's allowed topics and tasks explicitly rather than only listing what it should refuse, since an open-ended refusal list is easy to route around
- Tell the model to treat instructions found inside user-supplied content (quoted text, uploaded files, search results) as data to discuss, never as commands to follow
- Give the model a fixed, short decline-and-redirect phrase so off-topic refusals stay consistent instead of drifting across conversations
- Test the prompt against real attempted bypasses (role-play requests, "ignore previous instructions," fake system messages) rather than only the happy path
Common mistakes
- Writing the guardrail as a vague request like "stay on topic" instead of naming the specific topics, tasks, and tone the bot should maintain
- Forgetting to address content injected through tools or documents, so the model treats a pasted email's instructions as if you had written them
- Making the refusal so rigid that the bot blocks legitimate edge-case questions that are related to its actual job
- Never testing the prompt against adversarial inputs, so the first real jailbreak attempt is the first time anyone checks if it holds
FAQs
How do I stop users from getting a chatbot to ignore its system prompt?
Tell the model explicitly that instructions appearing inside user messages, pasted text, or documents are content to discuss, not commands to obey, and have it continue following the original system prompt regardless of claims that a new message overrides it. Pairing that rule with a clear, named scope makes it much harder for a single injected sentence to redirect the whole conversation.
Does a system prompt alone fully prevent prompt injection?
No single prompt makes a model immune to injection; a well-written system prompt reduces how often basic attempts succeed, but determined or novel attacks can still get partial compliance. Treat prompt-level guardrails as one layer, and pair them with output filtering or review for anything high-stakes.
Why does my bot still go off-topic even with a 'stay on topic' instruction?
A vague instruction like "stay on topic" gives the model no concrete boundary to check requests against. Listing the specific allowed tasks and specific disallowed topics, plus a fixed redirect phrase, gives the model something concrete to compare each request to instead of guessing at what "on topic" means.
Should the decline message be the same every time or vary by situation?
Keeping it close to a fixed phrase is usually better for consistency and support quality, since a varying refusal can sound evasive or inconsistent to users who compare notes. A small amount of variation to match context (referencing the actual topic asked about) is fine as long as the core boundary and tone stay the same.
Which Cuelara tool can help me spot weak points in a guardrail prompt like this before deploying it?
Prompt Formatter — restructures a guardrail prompt into clean, well-organized Markdown or XML sections, which makes scope rules and injection-handling instructions easier for the model to parse reliably. Intelligence Score — grades the prompt's clarity and specificity so you can see whether the scope and refusal rules are concrete enough before you rely on them in production.