Users have 4–8 digital identities across your stack.
Users have 4–8 digital identities across your stack. Until you unify them, your personalisation system is working from an incomplete picture, and the model performance reflects that.
A single customer typically exists in your stack as: an anonymous browser cookie on first visit, a registered account with an email, a loyalty programme member with a separate ID, a mobile app user with a device ID, a customer in your CRM with a support ticket history, and a subscriber to your email list with a different email than their account.
These are all the same person. Your personalisation model treats them as six strangers.
Personalisation models require signal: a history of what a user has browsed, clicked, purchased, skipped. When that history is fragmented across six identities, each individual identity looks like a new user. The model has no pattern to learn from. It falls back to popularity-based recommendations instead of true personalisation.
The result: the recommendation engine looks active but behaves as a general-purpose popularity engine with a "Recommended for You" label.
Solving this requires three things:
1. A deterministic matching layer: When the same email appears in two systems, link the records. Same device ID appearing across web and app. Login events that merge an anonymous session into a registered account. These are hard rules, no ML required, and they should be the foundation.
2. A probabilistic matching layer: For identities that share no exact-match keys, use probabilistic signals, same IP address plus similar browsing times, shared device fingerprint characteristics, behavioural similarity. This is where ML enters, but it's identity resolution ML, not recommendation ML.
3. A unified identity store: A canonical user ID that all downstream systems reference. Every event from any source gets linked to a canonical ID. The recommendation system reads from this unified view, not from individual source systems.
In almost every e-commerce recommendation project we've worked on, the first three weeks are identity resolution, not ML. The team wants to start with the model. We insist on starting with the data foundation. The models we train after a proper identity resolution work typically show 3–5x improvement in offline evaluation metrics compared to models trained on fragmented data, before we've changed anything about the model architecture.
The recommendation algorithm is the easy part. The data engineering is the work.
We run 2-hour feasibility calls at no cost. We'll tell you what applies to your project.