Connect the same person across phone, laptop, and tablet using signals you can stand behind, with an honest fallback when you cannot
A customer researches on their phone over lunch, then buys on their laptop that evening. To most stacks that is two people: one anonymous mobile visitor who abandoned, and one returning desktop customer who converted. The retargeting keeps chasing the phone for a purchase that already happened on the laptop, the attribution credits the wrong session, and the personalization treats a returning buyer like a stranger because it never connected the surfaces.
The instinct is to reach for a device graph vendor and probabilistically stitch everything to everything. That works in the demo and creates a different problem in production: probabilistic matches made silently, at scale, without disclosure, which is both a compliance exposure under GDPR and a trust collapse the moment a customer notices their phone behavior showing up on a shared family laptop.
This recipe builds the deterministic baseline first: connect the surfaces you can connect with signals you can defend, and treat probabilistic linking as a consented, disclosed extension rather than the default.
A customer who identifies on any device is recognized on the others, carrying their prior context forward. Realistic effect on the deterministic baseline: resolution of the logged-in or post-purchase share of the audience, which in most consumer businesses is somewhere between a third and two-thirds of meaningful sessions depending on how much of the journey happens behind a login. Measure your own share before building anything: the proportion of meaningful sessions in your analytics that carry a login or a purchase. The figure climbs in subscription and B2B contexts where authentication is routine, and stays lower in casual-purchase retail where most buyers never create an account. Probabilistic linking can extend coverage further, but the honest framing is that it trades match confidence for reach, and that trade should be a decision rather than a default.
The downstream payoff is continuity: suppression that holds across devices, attribution that credits the session that actually converted, and personalization that recognizes a returning customer rather than restarting the journey on each new screen.
A CDP or identity-resolution layer that maintains the graph and exposes a resolution service the rest of the stack can query at decision time. Composable stacks typically run this as a dedicated identity service over the warehouse, with reverse-ETL pushing the resolved graph to activation destinations. Packaged suites bundle the graph into the marketing cloud, which is simpler to stand up and harder to inspect when a merge goes wrong. The capability that matters most is not the matching algorithm but the merge-and-split handling: the system has to be able to undo a bad merge cleanly, because it will make some, and a graph that can only merge accumulates errors that compound over time.
Compare the tools on Martech Stack Builder
Identity or CDP ownership holds the design decisions: the structure of the graph, the merge rules, and the line between deterministic and probabilistic matching. Data engineering builds the writes and the resolution service that decisioning reads from, including the merge-and-split logic that keeps the graph correctable. Legal is critical here rather than advisory, because where probabilistic matching is permitted, what disclosure it requires, and how long the graph itself may be retained are questions with real regulatory weight that differ across the regimes the business operates in. Analytics measures the match rate, audits for false merges where two different people collapse into one, and validates that journeys actually stay continuous across devices rather than just appearing to in a dashboard. The recipe takes months rather than weeks because the graph is infrastructure that everything downstream comes to depend on, and because the merge rules need real traffic to tune. The deterministic baseline can ship sooner; the time goes into making it correctable and defensible.
Work in this order. Each step protects the next from a mistake that is expensive to undo.
Anonymous-to-known stitching sits in step 3, cross-device identity resolution in steps 2, 5 and 6, and identifier graph write in step 4.
The first and worst failure is the false merge. Two people who share a household laptop, or a hashed email reused across family members, collapse into one resolved identity, and now one person's behavior drives the other's experience. Beyond the personalization weirdness, this is a privacy incident: you have shown one customer's activity to another. The defense is conservative deterministic merge rules, a probabilistic layer that is gated rather than default, and merge-and-split logic that can undo the error once detected.
The second is staleness at the write layer. If the graph updates in nightly batch, a customer who logs in on a new device this morning is still anonymous to the activation layer this afternoon, and the welcome-back experience fires from the stranger branch. The load-bearing edges have to land synchronously for real-time decisioning to mean anything; a graph that is only eventually consistent quietly undermines every recipe built on top of it.
The third is treating consent as a checkbox rather than a constraint on the linking itself. Resolving behavior across devices is a use of personal data with a legal basis that is not automatic, and probabilistic matching in particular carries disclosure obligations in several regimes. A graph built without legal in the room from the start tends to get rebuilt after the first audit, which is the expensive way to learn this.
The Workshop works out with your team which of these matter for your stack right now, and what to do first: a 90-minute session with the people who own the decision.
The deterministic baseline, where probabilistic linking is permitted and what it requires, the synchronous-versus-batch write strategy, and the merge-and-split rules that keep the graph correctable: those are the decisions that separate an identity graph you can trust from one that quietly accumulates errors.
Did this recipe match your situation?Anonymous response. Sign up to leave a longer note tied to your account.
Resolve individual visitors to the account they belong to, so the buying committee shows up as one account rather than scattered anonymous leads
Run identity resolution behind one service rather than each downstream system maintaining its own join logic, so a single resolution rule survives the next stack change
Normalize email at capture so [email protected], [email protected], and [email protected] all resolve to the same person, before they become three records