SQL JOIN Duplicate Rows
A join can multiply a row from the 'one' side once for every match on the 'many' side — genuinely surprising the first time you see it, and a common source of subtly wrong aggregate totals.
What & Why
When one marketing.campaigns row matches several marketing.leads rows, the campaign columns repeat once per matching lead. This is expected one-to-many behavior, but a later COUNT(*) or sum can be inflated if the intended grain is campaigns rather than joined rows.
See How It Works
Marketing compares the number of joined campaign-lead rows with the number of unique campaigns represented.
| id | name | channel | spend | start_date | end_date | status | target_segment |
|---|---|---|---|---|---|---|---|
| 1 | Spring Launch | google_ads | 55000.00 | 2024-01-15 | 2024-03-31 | active | smb |
| 2 | Retention Webinar | 45000.00 | 2024-02-10 | 2024-04-15 | active | enterprise | |
| 3 | Finance Retargeting | 50000.00 | 2024-03-12 | 2024-05-31 | active | enterprise | |
| 4 | Enterprise Search | google_ads | 60000.00 | 2024-04-01 | 2024-06-30 | active | enterprise |
| id | campaign_id | created_at | qualified_at | converted_at | lead_score | source | country | |
|---|---|---|---|---|---|---|---|---|
| 301 | 1 | ana@example.com | 2024-01-21 09:10:00+00 | 2024-01-22 11:00:00+00 | 2024-02-02 10:00:00+00 | 86 | google_ads | US |
| 302 | 1 | ben@example.com | 2024-01-24 12:40:00+00 | NULL | NULL | 52 | google_ads | CA |
| 303 | 2 | chloe@example.com | 2024-02-16 08:30:00+00 | 2024-02-18 14:20:00+00 | NULL | 74 | GB | |
| 304 | 3 | dev@example.com | 2024-03-20 17:15:00+00 | 2024-03-21 09:00:00+00 | 2024-04-04 16:00:00+00 | 91 | US |
| id | name |
|---|---|
| 1 | Spring Launch |
| 2 | Retention Webinar |
| 3 | Finance Retargeting |
| 4 | Enterprise Search |
| id | campaign_id | source |
|---|---|---|
| 301 | 1 | google_ads |
| 302 | 1 | google_ads |
| 303 | 2 | |
| 304 | 3 |
| campaign_name | lead_id | source |
|---|---|---|
| Spring Launch | 301 | google_ads |
Lead 301 adds one valid joined row without changing the source tables.
SELECT
COUNT(*) AS joined_rows,
COUNT(DISTINCT c.id) AS unique_campaigns
FROM marketing.campaigns c
JOIN marketing.leads l ON l.campaign_id = c.id;Practice this concept
Marketing wants to compare joined campaign-lead rows with the number of unique campaigns represented.
marketingPrefix tables with marketing.table_name.
joined_rowsunique_campaignsmarketing.campaigns| Column | Type |
|---|---|
| id | integer |
| name | text |
| channel | text |
| spend | numeric |
| start_date | date |
| end_date | date |
| status | text |
| target_segment | text |
| legacy_id | text |
marketing.leads| Column | Type |
|---|---|
| id | integer |
| campaign_id | integer |
| text | |
| created_at | timestamp with time zone |
| qualified_at | timestamp with time zone |
| converted_at | timestamp with time zone |
| lead_score | integer |
| source | text |
| country | text |
| archive_status | text |
Sign up free to try it on a real business scenario