How to Build a Doctor 360 (HCP 360) View
Search for how to build an HCP 360 view and you will find a great deal of agreement about why you should. Unified doctor profiles improve targeting, reduce duplicate outreach, make field and marketing activity legible in one place, and let a brand team answer questions that currently take three analysts and a week. All true, and none of it is the reason these projects fail.
The reason they fail is that the interesting part — deciding what a doctor 360 contains — takes about a day, and the boring part — deciding that two records refer to the same human being — takes six months and is where the entire value is created or destroyed. Most published guidance covers the first and gestures at the second.
This article is written the other way round. It assumes you already believe in the destination and need to know how to get there, on Indian doctor data, where the identity problem is genuinely harder than it is in the markets most of this guidance was written for.
An HCP 360 view is a single reconciled record per healthcare professional, assembled from every internal and external source that touches them, with one stable identifier that every downstream system uses. It is not a report and not a dashboard — it is a data structure with a resolution layer underneath it. The build has three parts: ingest every source without cleaning it first, resolve the records that refer to the same person, and survive conflicting field values according to explicit rules. Roughly 80% of the effort sits in the resolution step, and the quality of that step determines whether everything above it is worth having.
What data goes into a doctor 360 profile?
Seven layers: identity (who they are and how you recognise them), professional (qualification, specialty, registration), affiliation (where they practise, and when that changed), behavioural (what they have done with you across every channel), commercial (prescribing, potential, segmentation), influence (publications, trials, networks, KOL status), and governance (consent state, data source lineage, last-verified date). Most teams build identity, affiliation and behavioural, skip governance, and discover the gap during their first audit.
The field lists below are the working version, not the aspirational one. Build the first three layers properly before adding the rest — a 360 with three complete layers is useful, and one with seven half-populated layers is not.
| Layer | Core fields | Typical sources | Build priority |
|---|---|---|---|
| 1. Identity | Internal doctor ID (stable, meaningless) · full name and name variants · transliterated forms · date of birth where available · gender · mobile numbers · email addresses · source system IDs as aliases | CRM, field app, event registrations, data vendors, web forms | First. Nothing works without it |
| 2. Professional | Primary qualification · additional qualifications · state medical council · registration number · registration year · primary specialty · sub-specialty · years in practice | Council registers, data vendors, rep-collected data, NMR | First. Registration number is your strongest match key in India |
| 3. Affiliation | Current institution · institution type · department · designation · city, district, state, PIN · consultation address · secondary affiliations · affiliation start and end dates | CRM, field visits, hospital directories, data vendors | Second. The layer that decays fastest |
| 4. Behavioural | Calls and visits · emails sent, opened, clicked · WhatsApp sent, delivered, read, replied · event invitations and attendance · content viewed and duration · sample requests · portal logins | CRM, marketing platform, WhatsApp provider, event system, web analytics | Third. High volume, low ambiguity |
| 5. Commercial | Prescribing or sales indicators where lawfully available · brand potential · segment and tier · territory · assigned rep · target list membership · call frequency plan | Sales data, analytics, territory alignment | Fourth. Governance-sensitive |
| 6. Influence | Publications · trial involvement · society memberships · speaking history · guideline authorship · advisory board participation · peer network position · KOL tier | PubMed, trial registries, congress agendas, medical affairs records | Fifth. Medical affairs, separate governance |
| 7. Governance | Consent state per channel and per purpose · consent source and timestamp · data source per field · last-verified date per field · confidence score · merge history · do-not-contact flags | Consent platform, MDM, verification process | Build with layer one, not last. Retrofitting is painful |
The layer most teams leave until it is too late Layer seven is not administrative metadata. It is what makes the other six defensible. Without source lineage per field you cannot explain why the record says what it says. Without a last-verified date you cannot tell a stale record from a fresh one, so every record is treated as equally trustworthy, which means none of them are. Without consent state on the record itself, campaign systems have to ask a different system for permission — and that question gets skipped under deadline pressure. Under the DPDP framework this stopped being good practice and became structural. Build layer seven alongside layer one. It costs a fraction as much then as it does after the fact, when it means backfilling provenance for records whose origin nobody can reconstruct. |
The hard part — identity resolution without a national identifier
Every published HCP 360 methodology assumes a step that does not exist in India. Understanding why is the difference between a project that works and one that produces a confident, wrong answer.
What other markets have that India does not
| Market | National identifier | What it gives a builder |
|---|---|---|
| United States | NPI — National Provider Identifier, public, unique, universal | A deterministic join key. Two records with the same NPI are the same person. Match rates start in the nineties |
| United Kingdom | GMC number — public, unique, searchable | The cleanest public match key in Europe. Verification is a lookup |
| France | RPPS — national shared directory identifier | A national key, though less openly queryable than NPI |
| Netherlands | BIG register number | National, verifiable, structured |
| India | None currently in general use. Registration sits with individual state medical councils, historically without centralised electronic synchronisation | No deterministic join key. Matching must be built on name, registration number, qualification, location and contact data — every one of which is noisy |
This single row is why an HCP 360 build in India is not the same project as one in Boston, and why match rates quoted in US-authored guidance are not achievable targets here. A team benchmarking itself against a 95% deterministic match rate will conclude its data is uniquely bad. It is not — the identifier is simply absent.
What is changing, and why it should affect your design today
The National Medical Commission published a draft amendment — the Registration of Medical Practitioners and Licence to Practice Medicine (Amendment) Regulations, 2026, in August 2026 — proposing a Unique Identification (UID) Number for registered medical practitioners, issued under the National Medical Register, structured to embed the state or union territory code together with the state register number. Once assigned, a practitioner would be eligible to practise in any state without fresh registration, and disciplinary history would be centrally tracked. This is a draft proposal, not implemented law. It should not be treated as an existing identifier — but it should absolutely change how you design your identity layer this quarter.
The architectural decision this forces — and the cost of getting it wrong
If a national doctor identifier arrives in India, every HCP 360 built without a slot for it will need surgery. Retrofitting a new primary key into a live 360 — with downstream systems, historical behavioural data and campaign histories all keyed on something else — is the single most expensive avoidable rework in this category.
The design that survives either outcome costs nothing extra today:
Use a stable, meaningless internal ID as the primary key. Not the registration number, not the CRM ID, not an email. A surrogate key that never changes and carries no meaning.
Store every source identifier as an alias, not as the key. State council registration number, CRM ID, vendor ID, event system ID — all aliases against the internal key, each with its source and date.
Leave the UID slot open now. One nullable field with a source and verification date. When and if the UID lands, it becomes the strongest alias and eventually the preferred match key, and nothing above it has to change.
This is a five-minute schema decision that either costs nothing or saves a rebuild. Make it before you ingest your first source.
How to build a 360 view of doctors — the nine-step method
Written in the order the work actually has to happen. Steps two and five are where projects are won and lost.
- Define the entity before you define the fields. Decide precisely what one record represents — a person, or a person at a place of practice. A doctor consulting at three hospitals is one person and three affiliations, not three doctors. Teams that skip this decision discover it later as a duplicate problem that is actually a modelling problem, and no amount of match tuning will fix it.
- Design the identity model with a surrogate key and an alias table. Stable internal ID as primary key; every source identifier stored as an alias with its source and date; a nullable slot for a future national identifier. Do this before ingestion, not after.
- Inventory every source honestly. CRM, field application, event registrations, marketing platform, WhatsApp provider, sample management, web forms, purchased datasets, medical affairs records, and the spreadsheets on regional managers' laptops. The last category is always larger than expected and is often the most current.
- Ingest raw, then standardise — never the reverse. Land every source unaltered in a staging layer, then standardise into a common structure. Cleaning during ingestion destroys the evidence you will need when a match looks wrong six months later, and it always looks wrong eventually.
- Standardise the fields that matter for matching, aggressively. Names into components with transliteration variants; qualifications into a controlled vocabulary; specialties into a fixed taxonomy; mobile numbers into E.164; addresses into components with PIN; institutions into a canonical list. Matching quality is decided here, not in the algorithm.
- Resolve identity in two passes — deterministic first, probabilistic second. High-confidence exact matches on strong keys, then scored fuzzy matching on the remainder. Never one pass. The detail is in the next section.
- Apply survivorship rules to build the golden record. For every field, an explicit rule for which source wins and why. Written down, versioned, and reviewable — not embedded in a transformation script nobody can read.
- Publish the record with provenance attached. Every field carries its source, its verification date and a confidence indicator. A 360 that cannot explain itself will not be trusted by the brand teams it was built for, and untrusted data does not get used.
- Instrument it and set a refresh cadence. Match rate, false-merge rate, coverage and freshness measured monthly. Affiliation is the fastest-decaying layer and needs the shortest cycle.
The step teams skip and always regret - Step four — ingest raw before you standardise.
- The temptation is to clean at the door, because staging raw data feels like storing garbage. Six months later a brand manager insists a doctor has been merged with the wrong person. If you cleaned on ingestion, you cannot reconstruct what the source actually said, so you cannot tell whether the merge was wrong or the source was.
- Raw staging is cheap. Unreconstructable history is not.
Match logic that works on Indian doctor data
The section that most guidance omits, written concretely enough to hand to an engineer.
Pass one — deterministic matching on strong keys
Exact matches on keys strong enough that agreement means identity. Run these first, accept them without scoring, and remove matched records from the probabilistic pass.
| Key | Strength in India | Caution |
|---|---|---|
| State council code + registration number | Strongest available key. Unique within a council | Only as good as the capture quality. Reps transcribe these by hand and digits get transposed |
| Verified mobile number | Strong. Personal and rarely shared between doctors | Shared clinic numbers appear in older data. Exclude any number appearing against more than three distinct names |
| Professional email on an institutional domain | Strong where present | Coverage is low. Personal domains are much weaker — several doctors may share a clinic address |
| Future: NMR UID | Would become the strongest key if implemented as proposed | Draft only. Treat as an alias slot, not as an available key |
Pass two — probabilistic matching on the remainder
Everything unmatched goes to scored comparison. Three design decisions determine whether this works.
- Blocking keys, to make the problem computable. Comparing every record against every other is quadratic and unnecessary. Block on something cheap and reliable — first letter of surname plus city, or PIN plus specialty — and compare only within blocks. Use two or three different blocking schemes and union the results, so a record mis-keyed on one scheme still gets a chance on another.
- Field-level comparison with weights that reflect Indian data. Name similarity carries less weight here than in markets with stable spelling, because transliteration variance is genuine rather than erroneous — the same doctor legitimately appears with several spellings across sources. Registration number similarity carries more. Address similarity should be scored on components, since the same clinic is written five ways.
- Two thresholds, not one. Above the upper threshold, auto-merge. Below the lower, auto-reject. Between them, queue for human review — and staff that queue. Teams that set a single threshold either merge distinct doctors or leave obvious duplicates unmerged, and the first failure is far more damaging and far harder to detect.
The transliteration problem, specifically
Indian doctor names arrive in multiple romanisations across sources — the same person as Krishnan and Krishnann, Mohammad and Mohammed, Chatterjee and Chattopadhyay, with initials expanded in some sources and abbreviated in others, and given-name and surname order inverted between systems. Southern Indian naming conventions frequently place an initial before the given name in one source and after it in another.
Generic string-distance functions handle none of this well. What works is a preprocessing layer that generates name variants rather than a single canonical form — phonetic encodings tuned for Indian names, an expansion and contraction pass on initials, and an order-invariant comparison for given name and surname. Score against the best-matching variant rather than against one normalised string.
This is also why buying a verified doctor universe often outperforms building one. The variant handling is a specialised problem, and a provider who has solved it across a national universe has amortised work that is genuinely expensive to repeat. We set out how to evaluate that in doctor data validation for pharma.
Survivorship — deciding which source wins
Once records are resolved, conflicting values must be reconciled into one. This is where a 360 either becomes authoritative or becomes an average of everyone's mistakes.
The failure mode to avoid is a single global rule — "CRM always wins" or "most recent wins". Neither is true field by field. The rep who visited last week knows the current institution better than any vendor file; the vendor file knows the registration number better than a rep transcribing it into a phone.
| Field | Rule | Why |
|---|---|---|
| Registration number | Council register or verified vendor wins over rep-entered | Transcription error is the dominant failure mode. The register is authoritative by definition |
| Current institution | Most recent verified field observation wins, then vendor file | Affiliation changes constantly and the field team sees it first. This is the one field where recency genuinely beats authority |
| Specialty | Council or qualification-derived wins over self-reported and rep-entered | Reps record what a doctor practises; registers record what they are qualified in. Both are useful — store the qualified specialty and the practising specialty as separate fields rather than resolving them |
| Mobile number | Most recently verified as reachable wins. Retain previous numbers as inactive aliases | Numbers change. Never delete an old one — it is match evidence |
| Institutional domain beats personal. Most recently engaged beats older | Engagement is the only real proof an address is live | |
| Name | Do not resolve. Store a display name plus all observed variants | Collapsing to one spelling destroys match evidence and guarantees future duplicates |
| Consent state | The most restrictive value wins, always. Never the most recent | A withdrawal recorded anywhere must dominate a permission recorded elsewhere. Fail closed |
| Address | Most recent verified visit wins for practice address; vendor for registered address | They are different facts. Keep both |
The one survivorship rule with legal consequences Consent state must resolve to the most restrictive value, never the most recent. Every other field in this table resolves toward accuracy. This one resolves toward safety, and the distinction matters. If a doctor withdrew WhatsApp consent in one system last month and a stale opt-in exists in another system dated yesterday, recency logic will reinstate a permission that was revoked. Under the DPDP framework a withdrawal must be honoured and must cascade. A survivorship rule that lets a stale opt-in win is a compliance defect implemented in SQL, and it will be invisible until someone asks why a doctor who unsubscribed received a campaign. Write this rule explicitly, test it explicitly, and have it reviewed. |
How to know whether it worked
Four measures, reported separately. Collapsing them into one number is how a broken 360 passes review.
| Measure | What it asks | What good looks like | The trap |
|---|---|---|---|
| Match rate | What share of source records resolved to a golden record? | Rises over time as standardisation improves. Compare against your own baseline, never against US benchmarks built on NPI | Easy to inflate by loosening thresholds. Meaningless without the next row |
| False-merge rate | How many merges joined two different doctors? | Measured by sampling merged clusters and checking them by hand. Should be near zero and must be actively tested | The most damaging failure and the least visible. Nothing in the system reports it. You only find it by looking |
| Coverage | What share of the addressable universe do you hold, by specialty, city and tier? | Stated honestly per segment, including where it is thin | A high national number hiding near-zero coverage in tier-two cities |
| Freshness | What share of records were verified within the last N months, by field? | Reported per layer. Affiliation needs the shortest cycle | Treating a record verified once in 2023 as equivalent to one verified last month |
The relationship between the first two is the one to internalise. A 95% match rate with an 8% false-merge rate is materially worse than an 80% match rate with none — the unmatched records are a visible gap you can work on, while wrongly merged records are invisible corruption that propagates into targeting, territory alignment and incentive compensation. One under-delivers. The other actively misleads, and does so with total confidence.
Sample and hand-check merged clusters every month. It is dull work and it is the only way this failure is ever caught.
Consent state is a field, not a separate system
A short section on a decision that has become architectural rather than optional.
The common pattern is a consent platform beside the 360, with campaign tools querying it separately. It looks clean and it fails under pressure, because it makes permission a lookup that a system can skip. Every skipped lookup is a compliance event.
The alternative is to hold consent state on the doctor record itself, per channel and per purpose, with source and timestamp — layer seven above. Then a campaign selecting an audience from the 360 cannot select a doctor whose state does not permit that purpose on that channel, because the ineligible records are not returned. The restriction is enforced by the query, not by a person remembering.
The practical consequence for this build is that your 360 becomes the enforcement point, which raises the bar on the resolution step. If two records for the same doctor fail to merge and one carries a withdrawal, the unmerged twin remains contactable. An identity resolution failure becomes a consent failure. That is a strong argument for the human review queue described earlier, and for the monthly false-merge sampling.
We covered the obligations themselves in why pharma CRMs fail at consent tracking and the cascade requirement in consent withdrawal under DPDP.
Build it or buy the universe underneath it?
The honest framing is that this is not one decision. Almost nobody should build all of it, and almost nobody should buy all of it.
| Component | Build or buy | Reasoning |
|---|---|---|
| The doctor universe — who exists, where, in what specialty | Buy | Assembling and maintaining a national verified universe is a continuous operation, not a project. The variant and registration handling is specialised, and the cost of doing it badly is invisible until it is expensive |
| Identity resolution engine | Buy or use a platform | Deterministic plus probabilistic matching with blocking, thresholds and review queues is a solved problem. Building it in-house is six months of work to reach parity |
| Your own behavioural and commercial layers | Build | Nobody else has your CRM history, your campaign engagement or your sales data. This is your proprietary layer and the reason the 360 is yours |
| Survivorship rules | Build | These encode your commercial judgement about which of your sources to trust. A vendor default will be wrong for you in ways you will not notice |
| Consent layer | Buy or platform, integrate tightly | Regulatory surface, changing requirements, and the cost of an error is asymmetric |
| The serving layer — how brand teams actually query it | Build | Adoption is a product problem. A technically perfect 360 nobody queries has produced nothing |
The pattern that works most often: buy the universe and the resolution capability, build the layers that are genuinely yours, and spend the saved time on adoption. The most common reason an HCP 360 fails is not that it was built badly — it is that it was built correctly and nobody used it, because querying it required a data analyst and brand teams have deadlines.
Seven ways this goes wrong
Failure modes, in roughly the order they appear.
- Modelling the entity wrong at the start. Deciding late that a doctor at three hospitals is three records. Everything downstream inherits the error, and no matching improvement fixes a modelling mistake.
- Cleaning on ingestion. Destroys the evidence needed to adjudicate disputed merges later. Always land raw first.
- One match threshold instead of two. Guarantees either false merges or unmerged duplicates, and false merges are the expensive half.
- Never measuring false merges. The system will not tell you. Sampling is the only detection method, and teams that skip it discover the problem through a brand manager's complaint about a specific doctor — by which point the corruption has propagated for months.
- Global survivorship rules. "CRM always wins" is wrong for registration numbers; "most recent wins" is wrong for consent. Rules belong at field level.
- Treating governance metadata as optional. Source lineage, verification dates and consent state added later cost several times what they cost built in, and some of it cannot be reconstructed at all.
- Building for analysts instead of brand teams. If using the 360 requires SQL, it will be used by three people. Adoption is the deliverable, not the data model.
The failure that dwarfs the others - False merges, undetected.
- Every other item on this list produces a visible problem — a gap, a delay, an argument. This one produces a confident wrong answer that flows into targeting lists, territory alignment, call planning and incentive compensation, and it looks exactly like correct data until somebody notices that a cardiologist in Pune appears to have attended an event in Kochi.
- By then it has been in the numbers for two quarters. Sample your merged clusters every month. It is the least interesting recurring task in this build and the one with the highest expected value.
Reference architecture — how the pieces fit together
Five zones. The discipline that matters is that each zone has exactly one job and never reaches backwards, because the moment a downstream zone starts patching upstream data, you lose the ability to reason about where anything came from.
| Zone | Its one job | What lives here | The rule that keeps it honest |
|---|---|---|---|
| 1. Landing | Receive source data exactly as sent | One table per source, raw, append-only, with an ingestion timestamp and a batch identifier | Nothing is ever modified or deleted here. This is your evidence layer and your only defence when a merge is disputed |
| 2. Standardised | Reshape each source into one common structure | Same schema across all sources. Names decomposed, phones in E.164, specialties mapped to the taxonomy, addresses componentised | Transformations are deterministic and re-runnable. If you cannot rebuild this zone from landing, you have hidden logic somewhere |
| 3. Resolved | Decide which records are the same person | Match candidates, scores, decisions, the review queue, and the cluster-to-internal-ID assignment | Every merge decision is stored with its score and its reason. A merge you cannot explain is a merge you cannot defend |
| 4. Golden | Produce one record per doctor | The 360 itself — surviving field values, each carrying source, verification date and confidence | Survivorship rules are declarative and versioned. No field arrives here without provenance |
| 5. Serving | Make it usable by people who are not analysts | Filtered views, segment builders, exports, API endpoints, and the consent-enforced audience selector | Consent state is applied here as a filter that cannot be switched off, not as a warning a user can dismiss |
Two design consequences follow from this shape, and both are worth defending in a design review.
- The golden record is derived, never edited. When someone finds an error, the correction goes into a source or into the survivorship rules — never into the golden layer directly. A hand-edited golden record is unreproducible, and the next rebuild silently reverts it, which is how teams lose faith in the whole system.
- The serving layer is where consent is enforced, not where it is displayed. If a user can build an audience and then see a warning about consent, the warning will be ignored under deadline. If ineligible records are simply not returned by the query, there is nothing to ignore.
Making brand teams actually use it
Worth its own section, because a technically excellent 360 that nobody queries has produced nothing — and this is the most common way these programmes quietly fail after a successful launch.
The pattern is recognisable. The data team delivers on time, the architecture is sound, the match rates are respectable. Six months later usage is three analysts, brand teams are still requesting lists by email, and the programme is reviewed as a partial success that will not be extended. Nothing was built badly. The deliverable was simply defined as a data asset when it needed to be defined as a product.
Four things that change adoption
- Ship a segment builder, not a schema. A brand manager needs to answer "cardiologists in tier-two Maharashtra cities who attended an event in the last six months and have WhatsApp consent" without writing a query. If that takes SQL, they will email an analyst instead, and the 360 becomes an analyst tool rather than a commercial one.
- Show provenance in the interface, not just in the schema. When a brand manager sees a doctor's institution, they should be able to see where it came from and when it was verified. Visible provenance is what converts scepticism into trust — and scepticism is the default, because everyone has been burned by a bad list before.
- Make the feedback loop one click. A rep who knows a doctor has moved should be able to flag it from wherever they are looking, and that flag should enter the resolution process as a source rather than an email to the data team. The field force is your highest-frequency verification signal and it is almost always wasted.
- Publish a coverage and freshness page and keep it current. Which specialties, which cities, how many records, verified when. Teams trust a system that tells them where it is weak far more than one that presents everything with equal confidence — and they stop asking the data team the same three questions every month.
The adoption metric worth reporting to leadership - Not the match rate. The share of campaigns and target lists in the last quarter that were built from the 360 rather than from a spreadsheet.
- That number is uncomfortable at first and it is the only one that measures whether the investment changed anything. A 360 with an 88% match rate and 20% campaign adoption is a worse outcome than one at 78% and 90% adoption — the first is a data asset, the second is an operating change.
- Report it monthly from launch. What gets measured here is what gets used.
What good looks like at 90, 180 and 365 days
A realistic delivery shape for a team starting from scattered sources. It is deliberately unambitious in the first quarter, because the alternative is a broad shallow build that nobody trusts.
| Window | Scope | What exists at the end | What to report |
|---|---|---|---|
| Days 1–90 | Identity, professional and governance layers. Two or three highest-quality sources only. Resist the pressure to ingest everything | A resolved universe for one therapy area or one region, with provenance and consent state on every record. Review queue running | Match rate and false-merge rate on the initial sources. Coverage for the chosen segment, stated honestly |
| Days 91–180 | Affiliation and behavioural layers. Remaining internal sources. First version of the serving layer | Brand teams in one business unit building their own segments. Field feedback loop live | Campaign adoption share. Freshness by layer. Review-queue throughput and backlog |
| Days 181–365 | Commercial and influence layers. Full national coverage. Refresh cadence operating as routine rather than as a project | The 360 as the default source for targeting across business units | Adoption share across all units. Cost per verified record. Decay rate by field |
The most common planning error is inverting the first two rows — ingesting every source in the first quarter to demonstrate breadth, then spending the second fixing resolution problems created by sources that should have waited. Depth on a narrow universe first, breadth second. A brand team that trusts one therapy area completely will advocate for the programme far more effectively than one given national coverage it does not believe.
Frequently Asked Questions For How to Build a Doctor 360 (HCP 360) View
Seven layers. Identity — internal ID, name and name variants, contact points, and every source system identifier held as an alias. Professional — qualification, state medical council, registration number, specialty and sub-specialty. Affiliation — current institution, department, designation, addresses, and the dates those changed. Behavioural — calls, emails, WhatsApp, event attendance and content engagement across every channel. Commercial — potential, segment, territory and assigned representative. Influence — publications, trials, society memberships and speaking history. Governance — consent state per channel and purpose, source lineage per field, last-verified date and merge history. Most teams build the first four and skip the seventh, which is the one that makes the others defensible.
Nine steps: define the entity precisely, design an identity model with a surrogate key and an alias table, inventory every source, ingest raw before standardising, standardise the matching fields aggressively, resolve identity in two passes with deterministic matching first and probabilistic second, apply field-level survivorship rules, publish with provenance attached, then instrument and set a refresh cadence. The two steps that decide the outcome are the identity model and the resolution pass. And build the serving layer for brand teams rather than for analysts — the most common reason these projects fail is not bad data engineering but low adoption.
Because the United States has the NPI — a public, unique, universal identifier that makes matching largely deterministic. India has no equivalent in general use; registration sits with individual state medical councils, historically without centralised electronic synchronisation. Without a join key, matching must be built on names, registration numbers, qualifications, locations and contact data, all of which are noisy. Indian names also arrive in multiple legitimate romanisations across sources. Match rates quoted in US-authored guidance are not achievable benchmarks here, and a team measuring itself against them will wrongly conclude its data is uniquely poor.
Possibly, and sooner than most data teams have planned for. The National Medical Commission published a draft amendment in August 2026 — the Registration of Medical Practitioners and Licence to Practice Medicine (Amendment) Regulations, 2026 — proposing a Unique Identification Number issued under the National Medical Register, structured to embed the state or union territory code alongside the state register number, with practitioners then eligible to practise anywhere in India without fresh registration. This is a draft proposal and not implemented law, so it must not be treated as an available identifier. It should, however, change your design today: use a stable surrogate primary key, hold every source identifier as an alias, and leave a slot for the UID. That decision costs nothing now and avoids a rebuild later.
Identity resolution is deciding which records refer to the same person. It takes time because the decision is probabilistic rather than lookup-based when no shared identifier exists, and because the cost of the two error types is asymmetric. Failing to merge two records for one doctor creates a visible duplicate. Wrongly merging two different doctors creates invisible corruption that propagates into targeting, territory alignment and incentive compensation. Roughly 80% of the effort in an HCP 360 build sits here, and the correct approach is two passes — deterministic matching on strong keys, then scored probabilistic matching with two thresholds and a human review queue between them.
There is no honest universal number, and any vendor quoting one without seeing your data is guessing. What matters more is that you measure match rate and false-merge rate separately and track both against your own baseline rather than against external benchmarks built on markets with a national identifier. A 95% match rate with an 8% false-merge rate is worse than 80% with none. Sample merged clusters manually every month — nothing in the system will report a false merge to you.
Both, split by component. Buy the doctor universe, because maintaining a verified national universe is a continuous operation rather than a project, and the name-variant and registration handling is specialised work. Buy or platform the identity resolution engine, since deterministic plus probabilistic matching with blocking and review queues is a solved problem and building it in-house is roughly six months to reach parity. Build your behavioural and commercial layers, because nobody else has your CRM history or your sales data. Build the survivorship rules, because they encode your judgement about your own sources. And build the serving layer, because adoption is a product problem.
On the record, as a field, per channel and per purpose — not in a separate system that campaign tools are expected to query. Holding it on the record makes the 360 the enforcement point: a campaign selecting an audience cannot select doctors whose consent state does not permit that purpose on that channel, because those records are not returned. One consequence to plan for: an identity resolution failure then becomes a consent failure, since an unmerged duplicate of a doctor who withdrew remains contactable. That raises the stakes on the review queue and the monthly false-merge sampling. Note also that consent must survive to the most restrictive value, never the most recent.
Let's Discuss Your Requirements