Data Analysis Prompts
Prompts for data extraction, analysis, SQL generation, and visualization. This category has 5 ready-made Data Analysis prompts for ChatGPT, Claude and Gemini. Each one comes with a copy-ready template, a worked example, best practices and the mistakes to avoid. Open a prompt to fill in its blanks, or paste it straight into your AI chat.
Want one tailored to you? Fill in a prompt template and copy it.
AI Prompt to Segment Customer Data Into Meaningful Groups
This AI prompt for customer segmentation takes a description of your customer dataset and asks the model to propose meaningful groups based…
ROLE You are a data analyst helping design a customer segmentation scheme. CONTEXT Business type: [BUSINESS_TYPE, e.g. subscription SaaS, e-commerce retailer] Available fields: [LIST_OF_FIELDS, e.g. signup_date, last_order_date, total_orders, total_spend, plan_tier] Approximate data ranges: [FIELD_RANGES, e.g. total_orders typically 1-40, last_order_date spans the past 2 years] Business goal for this segmentation: [GOAL, e.g. identify customers to target for a win-back campaign] TASK Propose 4-6 customer segments that serve the stated goal. For each segment, provide: 1. A clear, specific name. 2. The exact rule or threshold on the available fields that defines membership. 3. A one-sentence description of what distinguishes this group. 4. One recommended action specific to this segment. CONSTRAINTS - Only use the fields listed above; do not assume data that wasn't mentioned. - Make segment rules mutually exclusive where possible, and note any customers who might not fit cleanly into any segment. - Keep each segment's rule specific enough to implement as a filter or query. OUTPUT FORMAT A numbered list of segments, each with Name, Rule, Description, and Recommended Action as labeled sub-points, followed by one line noting any edge cases not covered.
JSON Data Extraction Pipeline
This prompt converts messy, unstructured text (emails, articles, transcripts) into a predictable JSON array that your code can parse without…
Extract the following fields from the text below: [FIELD 1, FIELD 2, FIELD 3, FIELD 4]. Output rules: - Return strictly a JSON array of objects, one per entity found. - Use exactly these keys: [key_1, key_2, key_3, key_4]. - If a value is missing or unclear, use null. Never guess. - Do not wrap the output in markdown code fences. - Do not add any explanation before or after the JSON. Text: """ [PASTE TEXT] """
AI Prompt to Clean and Standardize a Messy Spreadsheet Dataset
This is a data cleaning prompt for turning a messy spreadsheet or CSV export into a consistent, analysis ready dataset — built for analysts,…
ROLE: You are a data cleaning assistant helping standardize a spreadsheet dataset for analysis. CONTEXT: Below is a sample of the dataset. Columns are: [LIST OF COLUMN NAMES] [PASTE SAMPLE DATA HERE, INCLUDING KNOWN PROBLEM ROWS] TASK: 1. Review each column and identify formatting inconsistencies (e.g. date formats, capitalization, whitespace, abbreviations, units) 2. Propose and apply a single standard format for each column: [SPECIFY TARGET FORMATS WHERE KNOWN, e.g. dates as YYYY-MM-DD] 3. Flag rows that look like duplicates or outliers, but do not delete them — mark them for my review instead 4. Do not alter these columns without flagging first: [COLUMNS REQUIRING REVIEW BEFORE CHANGES, e.g. customer name, ID numbers] CONSTRAINTS: - Preserve every original row unless I confirm a deletion - Do not invent or infer missing values — leave them blank and flag them - [ANY ADDITIONAL CONSTRAINT, e.g. keep a specific column's original casing] OUTPUT FORMAT: 1. The cleaned dataset as a table 2. A change log listing each column, what inconsistency was found, and what standard was applied 3. A separate list of flagged rows (possible duplicates/outliers) with a one-line reason for each flag
AI Prompt to Find and Remove Duplicate Records in a Dataset
This AI prompt helps you find and remove duplicate records in a dataset, including near duplicates that differ only in casing, spacing, abbr…
Role: You are a data quality analyst reviewing a dataset for duplicate records. Context: - Dataset: [PASTE DATASET OR DESCRIBE FILE AND COLUMNS] - Matching key(s): [FIELDS THAT SHOULD BE TREATED AS THE PRIMARY MATCH, e.g. EMAIL AND LAST NAME] - Secondary fields to consider for near-matches: [E.G. PHONE NUMBER, COMPANY NAME, ADDRESS] - Known formatting inconsistencies: [E.G. SOME ENTRIES HAVE EXTRA SPACES, MIXED CASE, ABBREVIATIONS] Task: 1. Identify rows that are exact duplicates on the matching key(s). 2. Identify rows that are likely duplicates despite minor formatting differences (case, spacing, abbreviation, punctuation). 3. For each duplicate group, recommend which row to keep based on: [RULE, e.g. MOST RECENT DATE, MOST COMPLETE RECORD]. Constraints: - Do not delete or merge any rows yourself — only recommend. - Do not treat records as duplicates solely because one field matches if other key fields clearly conflict. - Flag anything uncertain as "possible duplicate" rather than guessing. Output format: A table with columns: Row ID(s) involved | Match type (exact / likely / possible) | Matching fields | Recommended row to keep | Reason. After the table, list any rows you could not classify confidently and why.
ChatGPT Prompt to Convert Plain English Questions Into SQL Queries
This ChatGPT prompt converts a plain English question into a working SQL query , so analysts, product managers, and developers who don't wri…
Role: You are a SQL expert helping me write a query for [DATABASE SYSTEM, e.g. PostgreSQL 15]. Context: - Here is my table schema: [PASTE TABLE NAMES, COLUMNS, AND DATA TYPES] - Relationships between tables: [DESCRIBE FOREIGN KEYS / JOIN LOGIC] Task: Write a SQL query that answers this question: [PLAIN ENGLISH QUESTION, e.g. "which customers placed more than 3 orders in the last 30 days"] Constraints: - Use only the tables and columns listed above - Add a short comment above each major clause explaining what it does - If the question is ambiguous, state your assumption in a comment at the top of the query - Return results sorted by [SORT FIELD] in [ASCENDING/DESCENDING] order Output format: Return only the SQL query in a code block, followed by a 2-3 sentence plain-English explanation of what it returns.