Data Analysis & Data Science Prompts
Prompts for data cleaning, exploratory data analysis, visualization scripts, and SQL generation. This category has 5 ready-made Data Analysis & Data Science prompts for ChatGPT, Claude and Gemini. Each one comes with a copy-ready template, a worked example, best practices and the mistakes to avoid. Open a prompt to fill in its blanks, or paste it straight into your AI chat.
Want one tailored to you? Fill in a prompt template and copy it.
Claude Prompt to Analyze a CSV Dataset and Summarize Key Insights
This Claude prompt is for analysts, founders, and marketers who need to turn a raw CSV export into a plain language summary of what the data…
Role: You are a data analyst who explains findings in plain language for a non-technical audience. Context: - Dataset description: [WHAT THE DATA REPRESENTS, E.G., "MONTHLY SALES BY REGION"] - Columns: [COLUMN NAME: WHAT IT MEANS, repeat for each column] - Data: [PASTE CSV ROWS OR ATTACH THE FILE] - Business question: [WHAT YOU'RE TRYING TO LEARN, E.G., "WHICH REGION IS DECLINING"] Constraints: - State any assumptions you make about unclear columns before analyzing - Show the reasoning or calculation behind any number you report - Flag any data quality issues (missing values, small sample size, duplicates) you notice Task: Analyze the data to answer the business question above. Output format: 1. Key findings (3-5 bullet points, plain language) 2. Notable outliers or anomalies, with the specific rows or values involved 3. Caveats or data quality issues 4. Suggested next steps for further analysis
ChatGPT Prompt for Exploratory Data Analysis on a New Dataset
This ChatGPT prompt for exploratory data analysis (EDA) is built for analysts, students, and data scientists opening a dataset for the first…
Role: You are a data analyst performing an initial exploratory analysis on a new dataset. Context: - Dataset description: [WHAT THE DATA REPRESENTS AND WHERE IT CAME FROM] - Columns and data types: [LIST COLUMN NAMES AND TYPES] - Sample rows: [PASTE 5-10 SAMPLE ROWS] - Business question or goal: [WHAT YOU'RE ULTIMATELY TRYING TO UNDERSTAND OR DECIDE] Constraints: - Base every observation only on the data provided; do not invent values or trends not shown - Clearly separate factual observations from suggested next steps - Flag any columns with likely data-quality issues (missing values, inconsistent formatting, outliers) - If something cannot be determined from the sample provided, say so explicitly instead of guessing Output format: 1. Dataset overview (shape, column types, general structure) 2. Data-quality issues found (missing values, duplicates, inconsistent formats) 3. Key distributions and notable patterns 4. Three to five follow-up questions worth investigating further
AI Prompt to Detect Outliers and Anomalies in a Dataset
This AI prompt for detecting outliers and anomalies in a dataset is for analysts and data scientists who have a table of numbers — sales fig…
ROLE: You are a data analyst reviewing a dataset for outliers and anomalies before further analysis. CONTEXT: - Dataset (paste rows, CSV format or a table): [PASTE DATA] - Column descriptions and expected ranges/units: [DESCRIBE EACH RELEVANT COLUMN] - Known exceptions or seasonal patterns to account for: [E.G. "DECEMBER SALES ARE NORMALLY 3X HIGHER"] TASK: 1. For each numeric column, briefly describe what a normal/expected range of values looks like based on the data provided. 2. List specific rows that fall outside that range, referencing their row number or ID. 3. For each flagged row, classify it as either a likely data-entry error (e.g. impossible value, wrong format, wrong unit) or a statistical outlier (unusually high/low but plausible). 4. Give a one-sentence reason for each flag. CONSTRAINTS: - Only flag values you can point to directly in the provided data — do not assume rows exist beyond what was given. - Account for any known exceptions or seasonal patterns before flagging something as anomalous. - If a column has too few values to establish a reliable normal range, say so instead of guessing. OUTPUT FORMAT: Column: [name] - Expected range: [description] - Flagged rows: - Row [ID/number]: [value] — [Data-entry error / Statistical outlier] — [reason] (repeat per column)
ChatGPT Prompt to Check Data Quality Before Starting Analysis
This data quality check prompt is built for anyone who needs to validate a dataset before running any real analysis on it: analysts, data sc…
ROLE: You are a data quality analyst reviewing a dataset before it is used for analysis. CONTEXT: Dataset description: [WHAT THE DATASET CONTAINS AND ITS SOURCE] Intended use: [WHAT ANALYSIS OR DECISION THIS DATA WILL SUPPORT] Columns and expected types/ranges: [LIST COLUMN NAMES WITH EXPECTED FORMAT, e.g. "signup_date: YYYY-MM-DD", "amount: positive decimal"] Data sample or full dataset: [PASTE DATA OR DESCRIBE WHERE IT IS ATTACHED] TASK: Review the dataset against these quality dimensions: 1. Completeness - identify missing or null values and which columns/rows they affect 2. Consistency - flag mixed formats, inconsistent casing, or conflicting values for the same entity 3. Validity - flag values outside expected ranges, types, or formats 4. Uniqueness - identify duplicate records or duplicate keys that should be unique CONSTRAINTS: - For each issue found, cite the specific column and example row(s), not just a general statement - Rate each issue's severity as Blocking, Should Fix, or Cosmetic based on impact to [INTENDED USE] - Do not attempt to fix the data yourself; only identify and explain issues - If you cannot determine whether something is an issue without more context, say so explicitly rather than guessing OUTPUT FORMAT: A table with columns: Issue, Column(s) Affected, Example Row(s), Severity, Why It Matters Followed by a short summary paragraph stating whether the dataset is ready for analysis or needs cleanup first.
ChatGPT Prompt to Analyze A/B Test Results for Statistical Significance
This AI prompt for A/B test analysis walks a model through interpreting experiment results, so it tells you whether a difference between var…
Role: You are a data analyst evaluating the results of an A/B test for statistical significance. Context: - Test name/goal: [WHAT_WAS_BEING_TESTED] - Variant A: [CONVERSIONS_A] conversions out of [VISITORS_A] visitors - Variant B: [CONVERSIONS_B] conversions out of [VISITORS_B] visitors - Test duration: [NUMBER_OF_DAYS_OR_WEEKS] - Required confidence level: [CONFIDENCE_LEVEL, e.g. 95%] Task: 1. Calculate the conversion rate for each variant. 2. Determine whether the difference between variants is statistically significant at the stated confidence level, and show the calculation or reasoning. 3. Assess practical significance: is the size of the lift meaningful for the business, not just statistically detectable? 4. Flag any concerns about sample size, test duration, or external factors that could affect reliability. 5. Give a clear recommendation: ship variant B, keep variant A, or extend the test. Constraints: - Do not declare a winner based on conversion rate alone without checking significance. - State the result in plain English a non-statistician can act on. - Call out explicitly if the sample size is too small to trust the result. Output format: - Conversion rates for both variants - Significance result (significant or not, at what confidence level) - Practical significance assessment - Final recommendation with reasoning