ChatGPT Prompt to Check Data Quality Before Starting Analysis
This data quality check prompt is built for anyone who needs to validate a dataset before running any real analysis on it: analysts, data scientists, BI practitioners, and operators handed a spreadsheet or export they didn't produce themselves. Instead of jumping straight into charts or statistics, the prompt asks ChatGPT, Claude, or Gemini to systematically scan the data for the kinds of problems that quietly wreck downstream conclusions, like missing values, inconsistent formats, duplicate keys, impossible ranges, and mismatched data types across columns that should agree.
The model works from a data quality checklist rather than a vague "look this over" request, which matters because generic review prompts tend to catch only the most obvious issues and skip subtler ones, such as a date column silently mixing two formats or a category field with inconsistent casing that fragments what should be one group. Structuring the ask around specific checks (completeness, consistency, validity, uniqueness) gives the model a concrete rubric instead of leaving it to guess what "quality" means for this particular dataset.
Because a flagged issue list is only useful if someone can trust it, the prompt also asks the model to explain why each problem matters for the planned analysis, not just that it exists. When the dataset is large or split across multiple files, pulling only the relevant columns or sheets into context first with a tool like Context Extractor keeps the review focused and avoids burying real issues in irrelevant data the model has to wade through.
Prompt template
ROLE: You are a data quality analyst reviewing a dataset before it is used for analysis.
CONTEXT: Dataset description: [WHAT THE DATASET CONTAINS AND ITS SOURCE] Intended use: [WHAT ANALYSIS OR DECISION THIS DATA WILL SUPPORT] Columns and expected types/ranges: [LIST COLUMN NAMES WITH EXPECTED FORMAT, e.g. "signup_date: YYYY-MM-DD", "amount: positive decimal"] Data sample or full dataset: [PASTE DATA OR DESCRIBE WHERE IT IS ATTACHED]
TASK: Review the dataset against these quality dimensions:
- Completeness - identify missing or null values and which columns/rows they affect
- Consistency - flag mixed formats, inconsistent casing, or conflicting values for the same entity
- Validity - flag values outside expected ranges, types, or formats
- Uniqueness - identify duplicate records or duplicate keys that should be unique
CONSTRAINTS:
- For each issue found, cite the specific column and example row(s), not just a general statement
- Rate each issue's severity as Blocking, Should Fix, or Cosmetic based on impact to [INTENDED USE]
- Do not attempt to fix the data yourself; only identify and explain issues
- If you cannot determine whether something is an issue without more context, say so explicitly rather than guessing
OUTPUT FORMAT: A table with columns: Issue, Column(s) Affected, Example Row(s), Severity, Why It Matters Followed by a short summary paragraph stating whether the dataset is ready for analysis or needs cleanup first.
Example input
ROLE: You are a data quality analyst reviewing a dataset before it is used for analysis.
CONTEXT: Dataset description: Customer orders export from an e-commerce platform, 4,200 rows, covering the last 6 months Intended use: Monthly revenue trend analysis and customer segmentation Columns and expected types/ranges: order_id (unique string), customer_email (valid email format), order_date (YYYY-MM-DD), order_total (positive decimal, USD), status (one of: pending, shipped, cancelled, refunded) Data sample: order_id: ORD-1001, customer_email: [email protected], order_date: 2026-03-14, order_total: 89.99, status: shipped | order_id: ORD-1002, customer_email: [email protected], order_date: 03/15/2026, order_total: -12.50, status: Shipped | order_id: ORD-1001, customer_email: [email protected], order_date: 2026-03-16, order_total: 45.00, status: pending
TASK: Review the dataset against completeness, consistency, validity, and uniqueness. Rate severity and cite example rows.
OUTPUT FORMAT: Table as specified, followed by a readiness summary.
Example output
| Issue | Column(s) Affected | Example Row(s) | Severity | Why It Matters | |---|---|---|---|---| | Duplicate order_id | order_id | ORD-1001 appears twice ([email protected] and [email protected], different dates) | Blocking | order_id is expected to be unique; this will cause double-counting or incorrect joins in revenue totals | | Inconsistent date format | order_date | Row 2: 03/15/2026 vs. Row 1: 2026-03-14 | Should Fix | Mixed date formats will break date parsing and skew the monthly trend analysis | | Inconsistent email casing | customer_email | [email protected] vs. [email protected] (likely same customer) | Should Fix | Case differences will fragment one customer into two in segmentation unless normalized | | Negative order total | order_total | Row 2: -12.50 | Blocking | A negative total for a 'shipped' status is outside the valid range and will distort revenue sums unless it represents a refund that should be reclassified | | Inconsistent status casing | status | 'Shipped' vs. 'shipped' | Cosmetic | Will fragment status groupings in a group-by unless case is normalized first |
Summary: This dataset is not yet ready for revenue or segmentation analysis. Two blocking issues (a duplicate order_id and a negative total on a shipped order) need resolution before any aggregation, and the date and casing inconsistencies should be normalized to avoid silently incorrect groupings.
When to use it
- Before building a dashboard, report, or model on a dataset you received from someone else and haven't personally vetted
- When a dataset combines exports from multiple systems (CRM, spreadsheet, database dump) that may use different formats or conventions
- When stakeholders will make decisions based on the analysis, so undetected errors have real downstream cost
- When you're onboarding a new or unfamiliar dataset and need a fast read on how trustworthy it actually is
Best practices
- Tell the model the dataset's intended use (forecasting, reporting, joining with another table) so it flags issues that matter for that specific purpose, not quality in the abstract
- Ask for a severity rating on each issue found (blocking, worth fixing, cosmetic) so you can triage instead of treating every flag as equally urgent
- Share column definitions or a data dictionary if one exists, since the model can't know a value is invalid unless it knows the expected range or format
- Request examples of the actual rows that triggered each flag, not just a category label, so you can verify the issue is real before spending time fixing it
Common mistakes
- Asking the model to "check the data" with no criteria, which produces a shallow pass that misses domain-specific validity issues
- Treating the model's output as a guarantee of clean data rather than a first-pass review that still needs a spot check against the source
- Pasting a sample of rows and assuming the findings generalize to the full dataset, when a quality issue elsewhere may not appear in that sample
- Fixing every flagged issue mechanically without first confirming whether it actually affects the specific analysis you're about to run
FAQs
How do I write a ChatGPT prompt to check data quality before analysis?
Give the model the dataset (or a representative sample), the column definitions with expected formats, and the intended use of the analysis, then ask it to check specifically for completeness, consistency, validity, and uniqueness issues rather than a generic review. Asking for example rows and a severity rating for each flag makes the output actionable instead of just a list of concerns.
What's the difference between a data quality check and exploratory data analysis?
A data quality check is a narrower, earlier step focused on finding errors, missing values, and inconsistencies that would make an analysis unreliable. Exploratory data analysis comes after the data is trusted and focuses on understanding patterns, distributions, and relationships in the data itself.
Can AI models actually catch data quality issues reliably?
They're good at catching issues that are visible in the text of the data itself, such as mixed formats, obvious duplicates, out-of-range values, and inconsistent casing. They can't verify against ground truth outside the data they're given, so results should be treated as a strong first pass that still gets spot-checked, especially on a large dataset where only a sample was reviewed.
Should I paste raw data into the prompt or describe it?
Pasting actual rows (or a representative sample) works better than describing the data, because the model can only flag concrete issues it can see, like a specific malformed date or duplicate ID. A written description alone gives it nothing to check against.
Which Cuelara tool can help me work with large datasets in this kind of prompt?
Context Extractor — when the dataset is a large CSV, spreadsheet, or multi-file export, it pulls out just the relevant columns or sections so the quality check stays focused instead of getting lost in irrelevant data. Intelligence Score — useful for checking that your data quality prompt itself is specific enough (clear criteria, expected formats) before you run it on a real dataset.