Back to cookbook

AI Prompt to Segment Customer Data Into Meaningful Groups

0 views Updated

Make this prompt yours

Share

This AI prompt for customer segmentation takes a description of your customer dataset and asks the model to propose meaningful groups based on behavior, value, or lifecycle stage, along with the specific criteria that define each segment. It's built for analysts, marketers, and founders who have customer data in a spreadsheet or export but don't yet have a clear, actionable way to split it into groups worth treating differently.

Rather than asking for a vague list of "customer types," the prompt forces the model to name exact fields and thresholds for each segment, such as purchase frequency above a certain count or days since last order beyond a certain window, so the output can actually be turned into a filter or a query instead of staying a loose idea. It also asks for a short note on what action each segment suggests, since a segmentation that doesn't change what you do with a group isn't useful.

Once you have raw segmentation criteria back, pulling the relevant customer records together before running deeper analysis is often the next step; for large exports or multiple source files, Context Extractor can help pull only the relevant rows and fields out of bigger datasets before you build a full report.

Prompt template

Make this prompt yours

prompt-template
283 tokens
ROLE You are a data analyst helping design a customer segmentation scheme. CONTEXT Business type: [BUSINESS_TYPE, e.g. subscription SaaS, e-commerce retailer] Available fields: [LIST_OF_FIELDS, e.g. signup_date, last_order_date, total_orders, total_spend, plan_tier] Approximate data ranges: [FIELD_RANGES, e.g. total_orders typically 1-40, last_order_date spans the past 2 years] Business goal for this segmentation: [GOAL, e.g. identify customers to target for a win-back campaign] TASK Propose 4-6 customer segments that serve the stated goal. For each segment, provide: 1. A clear, specific name. 2. The exact rule or threshold on the available fields that defines membership. 3. A one-sentence description of what distinguishes this group. 4. One recommended action specific to this segment. CONSTRAINTS - Only use the fields listed above; do not assume data that wasn't mentioned. - Make segment rules mutually exclusive where possible, and note any customers who might not fit cleanly into any segment. - Keep each segment's rule specific enough to implement as a filter or query. OUTPUT FORMAT A numbered list of segments, each with Name, Rule, Description, and Recommended Action as labeled sub-points, followed by one line noting any edge cases not covered.

Want it sharper? Optimize this prompt with Prompt Optimizer, check it with the Prompt Debugger or shorten it with the Token Optimizer.

Example input

example-input
89 tokens
Business type: subscription SaaS Available fields: signup_date, last_login_date, monthly_spend, plan_tier, support_tickets_last_90_days Approximate data ranges: monthly_spend from $9 to $299, last_login_date spans the past 18 months, most accounts have 0-3 support tickets Business goal for this segmentation: identify accounts at risk of churning so the success team can reach out proactively

Example output

example-output
272 tokens
  1. Name: At-Risk High-Value Rule: monthly_spend >= $99 AND last_login_date more than 21 days ago Description: Paying customers on higher tiers who have gone quiet recently. Recommended Action: Trigger a personal outreach from the customer success team within 48 hours.

  2. Name: Disengaged Low-Tier Rule: plan_tier = 'basic' AND last_login_date more than 30 days ago Description: Lower-spend accounts that have stopped logging in but haven't cancelled. Recommended Action: Send an automated re-engagement email highlighting underused features.

  3. Name: Frustrated Active Users Rule: support_tickets_last_90_days >= 3 AND last_login_date within the past 7 days Description: Customers still actively using the product but hitting repeated issues. Recommended Action: Escalate to a senior support rep for a direct follow-up call.

  4. Name: Healthy Core Rule: last_login_date within the past 14 days AND support_tickets_last_90_days <= 1 Description: Regularly active accounts with minimal support friction. Recommended Action: Include in case studies or referral program outreach.

Edge cases: New signups within the last 14 days don't yet have enough login history to classify reliably and should be excluded from this segmentation until more data accumulates.

When to use it

  • You have a customer export with fields like order history, signup date, or plan tier, but no defined way to group them yet.
  • You're planning a marketing campaign and need distinct segments with different messaging rather than one message for everyone.
  • You want to identify your highest-value customers and your at-risk or lapsing customers as separate, actionable groups.
  • You need segment definitions specific enough to hand to someone else to turn into a database query or dashboard filter.

Best practices

  • List the actual columns available in your dataset so the model proposes segments based on data you really have, not fields you'd need to collect.
  • Ask for a specific threshold or rule for each segment, not just a label, so the definition can be turned into a filter later.
  • Request a short recommended action per segment so the output connects to a decision, not just a description.
  • Share rough counts or ranges for key fields (like typical order frequency) so segment thresholds are realistic for your actual customer base.

Common mistakes

  • Asking for segments without listing available fields, which produces generic marketing personas instead of criteria you can actually apply.
  • Accepting vague segment names like 'loyal customers' without a concrete rule for who qualifies.
  • Creating too many overlapping segments that don't lead to different actions, making the segmentation harder to use than no segmentation at all.
  • Forgetting to ask the model to flag edge cases, like customers who don't clearly fit any proposed segment.

FAQs

How many customer segments should I create?

Most useful segmentations land between 3 and 6 groups; more than that usually means segments stop mapping to genuinely different actions, which defeats the purpose of segmenting in the first place.

What fields are most useful for segmenting customers?

Recency, frequency, and monetary value (how recently they bought, how often, and how much) are the most common starting point, supplemented by lifecycle fields like signup date or plan tier when available.

Can this prompt work with a small customer base?

Yes, but with very few customers some proposed segments may end up with only one or two members; it helps to tell the model your approximate customer count so it can size segment rules appropriately.

How is this different from RFM analysis?

RFM (recency, frequency, monetary) is one common input to segmentation, but this prompt is broader and can incorporate any fields you have, including lifecycle stage, support history, or plan tier, not just the three RFM dimensions.

How can I check this prompt's clarity before running it on a large export?

Prompt Debugger — it's useful for catching vague constraints in the segment rules before you run the prompt against a full customer dataset and get back criteria that are too loose to apply.

Found this prompt useful? Share it.

Share

More in Data Analysis

Data Analysis

JSON Data Extraction Pipeline

This prompt converts messy, unstructured text (emails, articles, transcripts) into a predictable JSON array that your code can parse without…

Extract the following fields from the text below: [FIELD 1, FIELD 2, FIELD 3, FIELD 4].

Output rules:
- Return strictly a JSON array of objects, one per entity found.
- Use exactly these keys: [key_1, key_2, key_3, key_4].
- If a value is missing or unclear, use null. Never guess.
- Do not wrap the output in markdown code fences.
- Do not add any explanation before or after the JSON.

Text:
"""
[PASTE TEXT]
"""

Make this prompt yours

Data Analysis

AI Prompt to Clean and Standardize a Messy Spreadsheet Dataset

This is a data cleaning prompt for turning a messy spreadsheet or CSV export into a consistent, analysis ready dataset — built for analysts,…

ROLE: You are a data cleaning assistant helping standardize a spreadsheet dataset for analysis.

CONTEXT: Below is a sample of the dataset. Columns are: [LIST OF COLUMN NAMES]

[PASTE SAMPLE DATA HERE, INCLUDING KNOWN PROBLEM ROWS]

TASK:
1. Review each column and identify formatting inconsistencies (e.g. date formats, capitalization, whitespace, abbreviations, units)
2. Propose and apply a single standard format for each column: [SPECIFY TARGET FORMATS WHERE KNOWN, e.g. dates as YYYY-MM-DD]
3. Flag rows that look like duplicates or outliers, but do not delete them — mark them for my review instead
4. Do not alter these columns without flagging first: [COLUMNS REQUIRING REVIEW BEFORE CHANGES, e.g. customer name, ID numbers]

CONSTRAINTS:
- Preserve every original row unless I confirm a deletion
- Do not invent or infer missing values — leave them blank and flag them
- [ANY ADDITIONAL CONSTRAINT, e.g. keep a specific column's original casing]

OUTPUT FORMAT:
1. The cleaned dataset as a table
2. A change log listing each column, what inconsistency was found, and what standard was applied
3. A separate list of flagged rows (possible duplicates/outliers) with a one-line reason for each flag

Make this prompt yours

Data Analysis

AI Prompt to Find and Remove Duplicate Records in a Dataset

This AI prompt helps you find and remove duplicate records in a dataset, including near duplicates that differ only in casing, spacing, abbr…

Role: You are a data quality analyst reviewing a dataset for duplicate records.

Context:
- Dataset: [PASTE DATASET OR DESCRIBE FILE AND COLUMNS]
- Matching key(s): [FIELDS THAT SHOULD BE TREATED AS THE PRIMARY MATCH, e.g. EMAIL AND LAST NAME]
- Secondary fields to consider for near-matches: [E.G. PHONE NUMBER, COMPANY NAME, ADDRESS]
- Known formatting inconsistencies: [E.G. SOME ENTRIES HAVE EXTRA SPACES, MIXED CASE, ABBREVIATIONS]

Task:
1. Identify rows that are exact duplicates on the matching key(s).
2. Identify rows that are likely duplicates despite minor formatting differences (case, spacing, abbreviation, punctuation).
3. For each duplicate group, recommend which row to keep based on: [RULE, e.g. MOST RECENT DATE, MOST COMPLETE RECORD].

Constraints:
- Do not delete or merge any rows yourself — only recommend.
- Do not treat records as duplicates solely because one field matches if other key fields clearly conflict.
- Flag anything uncertain as "possible duplicate" rather than guessing.

Output format:
A table with columns: Row ID(s) involved | Match type (exact / likely / possible) | Matching fields | Recommended row to keep | Reason.
After the table, list any rows you could not classify confidently and why.

Make this prompt yours