Back to cookbook

ChatGPT Prompt for Exploratory Data Analysis on a New Dataset

2 views Updated
Share

This ChatGPT prompt for exploratory data analysis (EDA) is built for analysts, students, and data scientists opening a dataset for the first time and needing a structured first pass before any modeling starts. Instead of asking the model to "analyze this data" and getting a generic paragraph back, the prompt walks it through the same checklist a human analyst would follow: shape and structure, missing values, distributions, outliers, and relationships worth investigating further.

The prompt asks the model to separate observations from recommendations, so you get a clear read on what the data actually shows before any suggestion about what to do next. It also asks for follow-up questions the data raises, which is often the most useful output of an EDA pass, since it tells you where to dig deeper rather than treating the first summary as final.

When the dataset is large, comes from a CSV or PDF export, or is spread across multiple files, pasting the whole thing into a chat window either doesn't fit or buries the model in irrelevant rows; pulling just the relevant slice first with Context Extractor keeps the EDA prompt focused on the columns and records that actually matter.

Prompt template

Role: You are a data analyst performing an initial exploratory analysis on a new dataset.

Context:

  • Dataset description: [WHAT THE DATA REPRESENTS AND WHERE IT CAME FROM]
  • Columns and data types: [LIST COLUMN NAMES AND TYPES]
  • Sample rows: [PASTE 5-10 SAMPLE ROWS]
  • Business question or goal: [WHAT YOU'RE ULTIMATELY TRYING TO UNDERSTAND OR DECIDE]

Constraints:

  • Base every observation only on the data provided; do not invent values or trends not shown
  • Clearly separate factual observations from suggested next steps
  • Flag any columns with likely data-quality issues (missing values, inconsistent formatting, outliers)
  • If something cannot be determined from the sample provided, say so explicitly instead of guessing

Output format:

  1. Dataset overview (shape, column types, general structure)
  2. Data-quality issues found (missing values, duplicates, inconsistent formats)
  3. Key distributions and notable patterns
  4. Three to five follow-up questions worth investigating further

Example input

Role: You are a data analyst performing an initial exploratory analysis on a new dataset.

Context:

  • Dataset description: Monthly subscription cancellations for a SaaS product, exported from the billing system
  • Columns and data types: customer_id (string), plan_type (string), signup_date (date), cancel_date (date), cancel_reason (string, often blank), monthly_revenue (float)
  • Sample rows: cust_1042, Pro, 2024-03-01, 2025-01-15, "too expensive", 49.00 / cust_2210, Starter, 2024-11-20, 2025-01-18, "", 12.00 / cust_0087, Pro, 2023-06-05, 2025-01-20, "missing features", 49.00
  • Business question or goal: Understand which customer segments are cancelling most and why

Constraints:

  • Base every observation only on the data provided
  • Separate observations from next steps
  • Flag data-quality issues
  • Say explicitly if something can't be determined from the sample

Output format:

  1. Dataset overview
  2. Data-quality issues
  3. Key distributions and patterns
  4. Follow-up questions

Example output

  1. Dataset overview: Six columns covering customer ID, plan type, signup/cancel dates, cancellation reason, and monthly revenue. Each row represents one cancellation event, not the full customer base, so this sample only shows churned customers.

  2. Data-quality issues:

  • cancel_reason is blank for at least one row (cust_2210), which will undercount reasons in any breakdown
  • Only three rows are provided, which is too small a sample to draw statistically reliable conclusions β€” the patterns below should be treated as illustrative, not confirmed
  1. Key distributions and patterns:
  • Two of three sample rows are on the Pro plan ($49/mo), one on Starter ($12/mo)
  • Tenure before cancellation ranges from about 2 months (cust_2210) to over 19 months (cust_0087) in this sample
  • Stated reasons so far: "too expensive" and "missing features" β€” both suggest value perception issues rather than technical problems
  1. Follow-up questions:
  • What percentage of all cancellations have a blank cancel_reason, and can that gap be closed?
  • Does cancellation rate differ meaningfully between Pro and Starter plans across the full dataset?
  • Is there a tenure threshold (e.g., under 3 months) where cancellations cluster?
  • How does monthly_revenue at cancellation compare to the average revenue of active customers on the same plan?
  • Are "too expensive" cancellations concentrated in a particular signup cohort or pricing tier change?

When to use it

  • Opening a new dataset for the first time and needing a structured overview before modeling
  • Onboarding onto a dataset someone else collected, where the quirks and gaps aren't documented
  • Preparing a quick data summary to share with a non-technical stakeholder
  • Deciding which columns or segments deserve a deeper, more rigorous analysis

Best practices

  • Paste the actual column names, data types, and a few sample rows, not just a description of the dataset
  • Ask explicitly for missing-value counts and how the model would handle them, rather than letting it assume
  • Request the analysis in a fixed structure (summary, distributions, outliers, questions) so results are easy to scan
  • For very large or multi-sheet datasets, trim the input to the columns actually relevant to your business question before running the prompt

Common mistakes

  • Pasting an entire raw dataset with no context on what each column represents
  • Treating the model's summary as a finished analysis instead of a starting point for deeper investigation
  • Not asking for missing-value or data-quality issues, which then surface later during modeling
  • Accepting statistical claims (correlations, averages) without asking the model to show its reasoning or flag uncertainty

FAQs

How do I write a ChatGPT prompt for exploratory data analysis?

Give the model the actual column names, data types, and several real sample rows rather than a general description, then ask for a fixed structure: overview, data-quality issues, distributions, and follow-up questions. That structure is what turns a vague summary into a usable first pass.

Can ChatGPT do real statistical analysis on a dataset I paste in?

ChatGPT can describe patterns visible in the exact rows you provide and reason about likely relationships, but it cannot compute precise statistics across a full dataset it hasn't seen in full, and it can't verify claims against data outside what's pasted in. Treat its output as a starting hypothesis to verify with actual computation, not a finished statistical result.

What's the difference between EDA and full data analysis?

Exploratory data analysis is the first pass: understanding structure, quality, and obvious patterns before deciding what deeper analysis or modeling is worth doing. It's meant to generate questions and flag issues, not produce final conclusions.

How much sample data should I include in an EDA prompt?

Include enough rows to show the range of values and any quirks (blanks, inconsistent formatting, outliers) β€” typically 5 to 15 rows plus the full column list and data types. For large datasets, extracting a representative slice is more useful than pasting thousands of rows the model can't fully process anyway.

Which Cuelara tool helps when my dataset is too large to paste directly into a prompt?

Context Extractor β€” it pulls only the relevant snippets or columns from large CSVs, PDFs, or multi-file exports via RAG, which keeps the EDA prompt focused and cuts down on hallucinated patterns from incomplete context.

Found this prompt useful? Share it.

Share