How duplicates happen honestly
Duplicates are usually not sloppiness. The same homeowner fills in a form on Tuesday, does not hear back fast enough, and calls on Thursday. A spouse enquires separately about the same roof. Someone uses a work email in March and a personal one in June. Each is a genuinely separate contact event about one household.
So the question is never 'how do we stop duplicates' — it is 'when two records are the same person, what happens to what each of them knew'.
The merge rule decides your attribution model
Most systems default to the newest value winning, because that is the sensible rule for a phone number or an address. Applied to a source field, it is not hygiene, it is a model change: the most recent channel takes credit for a customer that a different channel introduced.
The effect compounds in the worst direction. Retargeting and branded search generate lots of second touches, so they progressively absorb credit from the channels that created the opportunities they are harvesting.
Rules that preserve the record
- First captured origin is immutable once set; later captures are stored as additional touches.
- Contact details follow last-write-wins — the newest phone number really is the better one.
- Merging is logged with both record identifiers, the winner, and who or what performed it.
- The surviving record keeps the earliest capture timestamp, since that is what attribution windows are measured from.
Duplicates distort cost figures too
Before dedup, one household counted three times makes cost per lead look a third of what it is. After dedup, the same spend divided by real people tells the truth, and often the truth is that a cheap-looking channel is producing the same buyers repeatedly.
Because of this, always compute cost per lead on deduplicated contacts, and say so on the report. Comparing a deduplicated figure to a raw one across quarters produces a phantom improvement nobody can explain.