Why alumni records are a genuinely hard entity-resolution problem
Every alumni platform demo looks intelligent on the vendor's curated tenant. The tenant is curated because the linkage problem was solved once, by hand, for a clean base. Your base is not that base. Alumni data is entity resolution's hardest consumer-adjacent case for structural reasons: people change names (marriage, divorce, professional preference), personal emails churn every few years, employment history spans multiple stints with gaps, and the same individual exists in systems built a decade apart that disagree about who they are.
Add the migration-specific wrinkles and the problem sharpens. Multiple exit-and-rehire cycles mean one person legitimately holds several employee records. Legacy systems stored contractors and employees in one table; the current HRIS does not. Some vintages of data use a stable employee identifier, some use a national insurance or social-security fragment, some use only a name and a department code that was reassigned in 2017. Every one of these patterns is present, in some volume, in every enterprise alumni base we have examined. The question is never whether you have a linkage problem; it is whether you solve it deliberately or discover it as a data-quality incident after go-live.
Linkage rules that actually work
Mature linkage pipelines are a funnel, and the funnel's shape is stable across tools and vendors:
– Deterministic first. Stable employee identifiers surviving across systems resolve the easy majority — often 60–80% of the base — with no ambiguity. Exhaust deterministic keys before any fuzzy logic touches them; probabilistic matching applied to records that share an employee ID manufactures duplicates with confidence scores attached.
– Probabilistic second, on the right features. For the remainder, the features that discriminate are name similarity, date-of-employment overlap, department and location, and (where lawfully usable) date of birth. Personal email is a weak feature — it churns; phone is weaker. A vendor whose "AI dedupe" leans on contact details has told you it has not done this at scale.
– A human review queue for the ambiguous middle. Between auto-match and auto-separate sits the band where a person decides. Budget it explicitly: review throughput of a few hundred pairs per analyst-day is the industry norm, and the queue size is a function of your base's age and source count, not the vendor's cleverness. The projects that skip the queue either merge the wrong people or duplicate the right ones, and both failures are visible to alumni.
– Unmatch honestly. A pipeline that forces every record into a cluster is worse than one that leaves orphans flagged. Orphans are countable, fixable, and honest; false merges are none of those things.
Effective-dated history is the asset — flattening destroys it
The most destructive migration shortcut is flattening: importing "current role at exit" and discarding the effective-dated history — every prior role, promotion date, location, and manager relationship. The flattened record supports a directory. The effective-dated record supports the graph: career-path analysis, promotion velocity, "who worked with whom when," and every career-oriented capability the category now sells — natural-language query over alumni, AI career roadmaps, alumni-to-opportunity matching. Those features consume history, not current-state rows.
This is where migration discipline and platform evaluation meet. If your incoming platform's import schema cannot represent effective-dated employment history, it does not have the substrate its roadmap promises, whatever the deck says. Ask to see the data model's history representation during evaluation, not after import, and import the history even if year-one features do not use it — it is the cheapest insurance in the project, because the source systems that hold it will not keep it forever.
Why AI features amplify bad linkage rather than surviving it
A search engine over a duplicate-heavy base returns the same person three times with three job titles, and the alumnus who spots it is an alum who now doubts the program's competence with their data. Worse, natural-language query and matching features compound the damage: a model asked "find alumni now in fintech in Singapore" over a base where current-role data attaches to the wrong person of the same name produces answers that are confidently wrong, and confident wrongness in an AI feature is the fastest route to program credibility collapse. The vendor that says its AI "works on any data" is selling a rules engine in an AI costume; the credible vendor tells you its features degrade below stated linkage thresholds and helps you measure whether you are above them.
UAT acceptance criteria that survive contact with reality
Treat linkage as a tested deliverable with numbers attached, not a phase that "completed." The acceptance set we recommend for a migration UAT:
– Gold-set precision and recall. Hand-label a stratified sample (200–500 individuals across vintages and exit years) before import; measure the pipeline's match precision and recall against it. Set contractual thresholds; anything in the high nineties on precision for a directory people can browse is the honest minimum.
– Duplicate rate by stratum. Report residual duplicate rate per source system and per exit decade, not pooled. Pooled numbers hide the 2008-vintage disaster inside the 2023-vintage success.
– History integrity spot-checks. Sampled individuals verified end-to-end: every role, date, and location matches source systems. Twenty checks per import batch, by people who knew the individuals' careers.
– Suppression and deletion propagation. Objected and erased individuals verified absent from search indexes and export files, not just the primary table.
– The review-queue burn-down. The ambiguous-middle queue must reach a defined steady state before launch; a migration that "passes" with 40,000 unresolved pairs pending has deferred the hard part, not completed it.
The cost line nobody budgets
Across every phase of alumni platform work, linkage is where projects quietly overrun, because the vendor's statement of work prices the software and the human review queue is yours. Name a data owner for the migration before signature — someone with authority over match rules and access to the analysts who work the queue. The firms whose migrations succeed are not the ones with cleaner data; they are the ones that treated identity resolution as the program-defining work package and staffed it like one. Everything the platform will claim to know about your alumni is downstream of whether these rules ran honestly.