SQL Data Wrangling
A broader look at why string cleaning matters — messy text is one of the most common reasons a dataset can't be trusted at face value.
What & Why
"Data wrangling" is the general term for transforming messy, inconsistent raw data into a clean, analysis-ready shape. Everything covered so far in this series — trimming, case normalization, splitting, pattern-based replacement — are the specific tools; wrangling is the overall practice of applying them systematically before trusting a dataset.
See How It Works
Marketing wants lead sources cleaned before a source-level count.
| id | campaign_id | created_at | qualified_at | converted_at | lead_score | source | country | |
|---|---|---|---|---|---|---|---|---|
| 301 | 1 | ana@example.com | 2024-01-21 09:10:00+00 | 2024-01-22 11:00:00+00 | 2024-02-02 10:00:00+00 | 86 | google_ads | US |
| 302 | 1 | ben@example.com | 2024-01-24 12:40:00+00 | NULL | NULL | 52 | google_ads | CA |
| 303 | 2 | chloe@example.com | 2024-02-16 08:30:00+00 | 2024-02-18 14:20:00+00 | NULL | 74 | GB | |
| 304 | 3 | dev@example.com | 2024-03-20 17:15:00+00 | 2024-03-21 09:00:00+00 | 2024-04-04 16:00:00+00 | 91 | US |
SELECT
LOWER(TRIM(REGEXP_REPLACE(source, '\s+', ' ', 'g'))) AS cleaned_source,
COUNT(*) AS leads
FROM marketing.leads
GROUP BY cleaned_source
ORDER BY leads DESC, cleaned_source;| cleaned_source | leads |
|---|---|
| google_ads | 2 |
| 1 | |
| 1 |
The complete cleanup pipeline preserves the actual normalized source counts.
Practice this concept
Marketing wants lead sources cleaned before a source-level count.
marketingPrefix tables with marketing.table_name.
cleaned_sourceleadsmarketing.leads| Column | Type |
|---|---|
| id | integer |
| campaign_id | integer |
| text | |
| created_at | timestamp with time zone |
| qualified_at | timestamp with time zone |
| converted_at | timestamp with time zone |
| lead_score | integer |
| source | text |
| country | text |
| archive_status | text |
Sign up free to try it on a real business scenario