SQL Data Wrangling

A broader look at why string cleaning matters — messy text is one of the most common reasons a dataset can't be trusted at face value.

What & Why

"Data wrangling" is the general term for transforming messy, inconsistent raw data into a clean, analysis-ready shape. Everything covered so far in this series — trimming, case normalization, splitting, pattern-based replacement — are the specific tools; wrangling is the overall practice of applying them systematically before trusting a dataset.

See How It Works

BUSINESS QUESTION

Marketing wants lead sources cleaned before a source-level count.

idcampaign_idemailcreated_atqualified_atconverted_atlead_scoresourcecountry
3011ana@example.com2024-01-21 09:10:00+002024-01-22 11:00:00+002024-02-02 10:00:00+0086google_adsUS
3021ben@example.com2024-01-24 12:40:00+00NULLNULL52google_adsCA
3032chloe@example.com2024-02-16 08:30:00+002024-02-18 14:20:00+00NULL74emailGB
3043dev@example.com2024-03-20 17:15:00+002024-03-21 09:00:00+002024-04-04 16:00:00+0091linkedinUS
EXAMPLE QUERY
SELECT
  LOWER(TRIM(REGEXP_REPLACE(source, '\s+', ' ', 'g'))) AS cleaned_source,
  COUNT(*) AS leads
FROM marketing.leads
GROUP BY cleaned_source
ORDER BY leads DESC, cleaned_source;
RESULT — exact output from the displayed Queryflo rows
cleaned_sourceleads
google_ads2
email1
linkedin1

The complete cleanup pipeline preserves the actual normalized source counts.

Now You Try

Practice this concept

Marketing wants lead sources cleaned before a source-level count.

Available schema
marketing

Prefix tables with marketing.table_name.

cleaned_sourceleads
marketing.leads
ColumnType
idinteger
campaign_idinteger
emailtext
created_attimestamp with time zone
qualified_attimestamp with time zone
converted_attimestamp with time zone
lead_scoreinteger
sourcetext
countrytext
archive_statustext
query.sql
Intermediate business practice

Sign up free to try it on a real business scenario