Module 02Data migration and cleanup
CRM data migration that merges your duplicates before a single record goes in.
A CRM data migration service for sales teams moving off spreadsheets or an old CRM. We profile every file, map every column to a field you actually use, and run duplicate cleanup before import. Loading the mess and promising to tidy it later is the one thing we will not do.
Price basis
What a migration costs is set by the record count, not by how bad the data looks.
We charge per 10,000 records across every object you move: contacts, accounts, deals and activity history all count. Where in the band a job lands depends on how many sources feed it and how many columns need rules. One clean export from one CRM sits at the bottom. Four spreadsheets with free-text owner names sit at the top.
A smaller job shows where the minimum bites. 38,000 records at the rate comes to $570 to $950, so the $900 floor lifts the low end and the band becomes $900 to $950. Every figure here is an estimate band, not a quote. We quote in writing after a 45-minute scoping call, once we have seen a sample of the files. Run your own numbers in the cost estimator; migration is priced separately there and is not touched by the automation multiplier.
Field mapping
Every source column gets a destination field and a written rule, or it gets dropped.
The mapping sheet is the migration equivalent of our stage definition sheet. One row per source column. Each row names the CRM field it lands in, the transform applied on the way, and who signed it off. Nothing imports until that sheet is approved by someone on your side who knows the data.
Most exports we receive carry between 40 and 120 columns. Roughly half of them survive. The rest are empty, duplicated under another name, or a field somebody made required in 2019 and nobody has read since. If a field is required and nobody reads it, we delete the field rather than migrate it.
The table below is a shortened version of a real mapping sheet shape. Your column names will differ. The rule column is the part that matters.
| Source column | CRM field | Rule on the way in |
|---|---|---|
| Company | Account Name | Trim whitespace, normalize Inc, LLC and Corp suffixes so "Acme, Inc." and "ACME INC" meet as one account |
| Contact Email | Lowercase, reject values with no @ or no domain, flag role addresses like sales@ for review | |
| Phone / Mobile | Phone, Mobile | Convert to E.164 (+1 and ten digits for US numbers), keep extensions in a separate field |
| Rep / Owner | Record Owner | Match free-text names to active users through a lookup table; departed reps go to a named holding owner |
| Status | Stage | Old values mapped to the signed stage sheet; anything with no match lands in a review queue, not a default stage |
| Last contacted | Last Activity Date | Parse to ISO 8601; ambiguous dates like 03/04 are resolved by the file's own majority format |
| Lead source | Lead Source picklist | Collapse free-text variants to the picklist values you approve; unmatched values kept in a text field, not guessed |
| Notes | Activity history | Split into dated activity records attached to the contact rather than one 30,000-character text field |
| Fax, Pager, Twitter 2 | Not migrated | Blank on 98% or more of rows, or unread since the field was created |
Duplicate cleanup
The dedupe rules we run, and the matches we refuse to merge by machine.
CRM duplicate cleanup is only safe when the match rules are written down before the first merge. Ours are strict on purpose. A false merge destroys history you cannot get back; a missed duplicate only costs a second look.
Exact email, after normalizing, merges without a human.
Lowercased, trimmed, with plus-addressing kept as written. Two records that share one real inbox are one person. The surviving record takes the most recently modified non-empty value for each field, except the owner, which follows whoever holds an open deal.
Same company domain and a close name goes to a review sheet.
"Jon Reyes" at reyesco.com and "Jonathan Reyes" at reyesco.com are probably the same buyer. Probably is not enough to merge. These pairs go into a spreadsheet with both records side by side and a keep column that someone on your team fills in. We merge exactly what they mark.
Never merged automatically: matches on phone number alone (shared switchboards), on name alone, or across two different account domains. Accounts follow the same logic, keyed on website domain first and normalized legal name second.
We merged 38,000 contacts into 21,400. The other 16,600 were the same people typed twice.
Reconciliation report, line 01
The 38,000 to 21,400 merge
Where 16,600 duplicates come from, and why nobody noticed them.
The file arrived as an old CRM export plus two years of spreadsheets the SDR team kept on a shared drive. 38,000 contact rows. The client believed they had about that many people in their market. They had 21,400.
The duplicates were not exotic. A rep imported a conference list that overlapped the CRM by a third. Another rep kept a personal sheet and re-imported it every quarter. Capitalization differences on email addresses defeated the old system's own duplicate check, which compared strings exactly.
The cost of leaving them in was visible in the numbers. Two reps were emailing the same buyer from two records. Activity counts were split, so the account looked cold on one record and warm on the other. Pipeline reports counted a few deals twice because they hung off twin contacts.
Rule 01 merged most of the pairs. The Rule 02 review sheet was the slow part: the client's sales ops lead worked through it over three afternoons. The final reconciliation report showed 38,000 rows in, 21,400 records out and 16,600 merged, with every merge traceable back to its source rows.
We take CSV, XLSX and API exports. Send what you have.
- CSV, UTF-8 preferred, comma or semicolon delimited. We detect Windows-1252 and convert it rather than asking you to re-export.
- XLSX, one object per sheet. Merged cells and colour-coded rows are fine; we read the colour legend as a column if you tell us what it meant.
- API export straight from your current CRM, when it has an export endpoint and you can grant a read-only user. This keeps record IDs, which makes relationships and activity history far easier to rebuild.
Some sources we turn away at the door.
- Purchased or scraped lead lists. We do not buy them, scrape them, or load them for you. They are the largest single source of duplicates and bounced email we see.
- PDFs and scanned business cards. Transcribe them first or leave them out.
Files move through a shared folder you control, not email attachments. We delete our working copies when the reconciliation report is signed.
How the migration runs
Five steps, one sign-off each, and a sandbox before production.
Inside a full engagement the data audit sits in weeks 1 to 2, alongside the stage sheet. A stand-alone migration follows the same order. See the week-by-week process for where it meets the rest of the build.
-
01Profile
We read every file and produce a profile: row counts per object, blank rate per column, distinct values per picklist, and the obvious duplicate count under Rule 01.
You get the profile before we price the rest. It usually settles the argument about which columns matter.
-
02Map
The mapping sheet is drafted against your signed stage definition sheet and field list from CRM setup. Old statuses map to new stages or into a review queue.
Sign-off: your sales ops lead or whoever owns the data.
-
03Dedupe
Rule 01 merges run on the staged data. Rule 02 pairs go out as a review sheet with a keep column. Nothing is loaded yet.
The review sheet is the one task we need real hours for on your side.
-
04Test load
A 10% sample goes into a sandbox. Reps open twenty of their own accounts and check owner, stage and history against what they remember.
Anything wrong goes back to the mapping sheet, not into a manual fix.
-
05Load and reconcile
Full import to production, then a reconciliation report: rows in, records out, merged, rejected with reason. The counts must add up to the row.
The 30-day fix window in our terms covers any record that did not land as the mapping sheet says.
Objections
The questions buyers ask before they send us a file.
Migration usually sits next to workflow automation or integrations. Clean data first, then rules that act on it.
Do we lose activity history when records merge?
No. Emails, calls, notes and tasks from both records reattach to the survivor. What goes away is the second record, not what happened on it. Attachments come across when the source is an API export; from CSV they usually cannot, and we tell you that at the profile step.
We have more than 500,000 records.
The estimator stops at 500,000 because above that the work changes shape: batch windows, API allowances and a longer test load. The per-10,000 rate still applies as a starting point. We quote it in writing after the scoping call.
Can you import now and dedupe once we are live?
No. Deduping after go-live means reps have already logged calls against both twins, created deals on the wrong one and built reports that count people twice. It costs more, not less. Dedupe before import is the rule we sell this service on.
Who is this not for?
If you have under 2,000 records from a single clean source, your CRM's own import tool will do the job in an afternoon, and we will say so on the call. The $900 minimum exists because below it the work is not worth your money.
Next step
Tell us the record count and the sources.
A rough count per object and the number of files is enough to start. We reply during business hours, Pacific time, and book the 45-minute scoping call from there.
SAN FRANCISCO, CA 94105