Standardize phone formats across every capture point so SMS reaches the whole file, not just the records that happened to be entered cleanly
SMS is the highest-engagement channel most brands underuse, and the reason is often unglamorous: the phone numbers are a mess. They were captured across web forms, in-store point of sale, support calls, and paid lead-gen, each in its own format, with country codes sometimes present and sometimes not, with spaces, dashes, parentheses, and leading zeros applied inconsistently. The result is a customer file where a meaningful fraction of the numbers will not send, the same customer appears several times under different formats, and the SMS program reaches a subset of the audience it should while the team assumes the channel just underperforms.
This is a data-hygiene problem masquerading as a channel problem. The same person's number captured as "(555) 123-4567" in store and "+15551234567" online are the same number, and a stack that does not normalize treats them as two, sends to neither reliably, and counts the customer's consent inconsistently. Normalizing phone numbers to a canonical format across every capture point is what makes SMS work across the whole file.
This recipe standardizes phone formats at capture and in backfill so SMS activation reaches the full audience.
A customer file where phone numbers are in one canonical format, deduplicated, and reliably sendable, so SMS reaches the whole eligible audience rather than the cleanly-entered subset. The metric is SMS deliverable rate: the share of the file with a valid, normalized, consented number that will actually send, which rises as normalization and dedupe replace the inconsistent raw capture. The realistic effect depends on how messy the starting file is, which is usually messier than the team assumes because the problem is invisible until someone audits send failures against the file size.
The second outcome is consistent consent accounting: when the same number is one record rather than three, the SMS consent attached to it is unambiguous rather than scattered across format-variant duplicates.
Normalization at the point of capture (validating and canonicalizing as numbers enter) plus a backfill pass over the existing file, with deduplication on the normalized form. A phone-validation library or service handles the canonicalization and country inference. Composable stacks suit this because the normalization runs in the pipeline feeding the warehouse and the dedupe happens against the resolved customer. The capability that matters is applying normalization both at capture (so new data is clean) and in backfill (so the existing mess is fixed), because normalizing only new captures leaves the historical file broken.
Compare the tools on Martech Stack Builder
Data engineering owns the normalization at capture and in backfill, the validation, and the dedupe on the normalized form, which is the substance of the recipe. Marketing ops runs SMS activation against the normalized file and monitors deliverability. Privacy ops covers SMS consent capture per source and ensures the consent travels with the normalized number, which matters because SMS consent is tightly regulated. The recipe ships in weeks because normalization is well-trodden with mature libraries. The work is applying it consistently across every capture source and backfilling the existing file, plus keeping the consent attached correctly through the dedupe.
Work in this order. Normalize at capture, backfill the file, dedupe on the canonical form, carry consent with the number.
Signal quality monitoring covers steps 1 to 3, 7 and 8, anonymous-to-known stitching step 4, and consent-scoped signal collection steps 5 and 6.
The first failure is normalizing new captures but not the existing file. Cleaning data going forward while leaving the historical mess untouched means SMS still fails for the bulk of the file that was captured before the fix, and the team sees only marginal improvement and concludes normalization did not help. The backfill is as important as the capture-time normalization.
The second is losing consent in the dedupe. When three format-variant duplicates of one number merge into one record, the SMS consent attached to each has to reconcile correctly, and a careless merge can drop a consent or, worse, infer consent the customer did not give. SMS consent is heavily regulated, so the dedupe has to preserve the consent state accurately rather than treating the records as interchangeable.
The third is dropping the country-code ambiguity on the floor. A number captured without a country code is ambiguous, and a normalization that guesses wrong sends to the wrong country or fails. The rule has to infer country from available context (capture location, customer address) and flag the genuinely ambiguous rather than silently assuming, because a wrong country inference is a failed send at best and a misdirected message at worst.
The Workshop works out with your team which of these matter for your stack right now, and what to do first: a 90-minute session with the people who own the decision.
The canonical format and country inference, the backfill that fixes the historical mess rather than only new captures, and the consent reconciliation that survives the dedupe: those are the decisions that turn a messy multi-source phone file into one SMS can actually reach.
Did this recipe match your situation?Anonymous response. Sign up to leave a longer note tied to your account.
Resolve individual visitors to the account they belong to, so the buying committee shows up as one account rather than scattered anonymous leads
Connect the same person across phone, laptop, and tablet using signals you can stand behind, with an honest fallback when you cannot
Run identity resolution behind one service rather than each downstream system maintaining its own join logic, so a single resolution rule survives the next stack change