Micro-Segmentation of Doctors: A Practical Guide
Micro-segmentation replaces one ranked list with several small, behaviourally distinct groups that each get a different action. The method is: decide the decision the segmentation must drive, inventory the signals you actually hold, build 4–8 segments across potential, behaviour, adoption stage and channel preference, name them so a rep understands them without training, and assign one differentiated action per segment. In markets without prescriber-level prescription data — India included — potential must be modelled from proxies rather than measured from scripts, which changes the technique but not the logic. A segmentation nobody acts on differently is not a segmentation; it is a report.
Search for how to segment doctors and you will find consulting content of a very high standard. ZS and Indegene have been shaping this discipline for decades and their published thinking is good. It is also, by design, abstract — the specific method is what you engage them for, and their writing is a demonstration of judgement rather than a set of instructions.
There is a second and larger problem with almost all of it, and it is not a criticism of the authors. Every published segmentation methodology assumes prescriber-level prescription data. Deciles are built from TRx. Behavioural personas are built from NBRx trends. Patient-volume segmentation is built from claims. One thorough 2026 guide to HCP targeting lists IQVIA Xponent, Symphony Health, Komodo, Veeva Link, the AMA Masterfile, Doximity and Medscape as the input stack — and contains no discussion of data availability outside the United States at all.
For an Indian brand manager, that renders most of the literature unusable in a specific and under-discussed way. This article is written for that situation: what micro-segmentation looks like when the inputs everyone assumes are simply not available, and how to build something better than the rep-rated A/B/C list that most Indian brands are still running.
Why decile targeting breaks
Worth being precise about, because deciles are not stupid — they are a reasonable answer to a question that has since changed.
Decile targeting ranks prescribers one to ten by total prescriptions in a therapy area over a rolling period, and concentrates effort on the top bands. It is the oldest framework in commercial pharma and still the most common. It works because past volume is genuinely predictive of future volume, and because it is simple enough for a field force to act on.
The three failures, in the US where the data exists
| Failure | What goes wrong | Evidence |
|---|---|---|
| Volume is not fit | A high-volume prescriber in a category may have none of the specific patients your brand serves | A documented launch case: an SGLT2 brand targeted high-decile prescribers who lacked treatment-resistant diabetic patients and achieved a 4% NBRx rate against an expected 12% |
| It is backward-looking | Deciles describe last year. A doctor who wrote their first three new-to-brand scripts last month looks identical to one who has written none | Separating NBRx and NRx from TRx surfaces movement that TRx alone conceals entirely |
| Specialty is too coarse | Targeting all endocrinologists for a diabetes therapy includes reproductive endocrinologists who will never prescribe it | Sub-specialty precision routinely matters more than specialty |
| It ignores reachability | A decile-9 doctor who refuses rep visits and never opens email is a lower practical opportunity than a decile-6 who engages on WhatsApp | Not a data problem — a modelling omission. Deciles have no channel dimension at all |
The consequence is measurable. Poorly targeted programmes have been shown to waste a meaningful share of spend on misidentified prescribers, while well-executed targeting produces new-to-brand lift in the range of 5 to 15 percentage points over unexposed controls. That range is a US benchmark on US data, and it is quoted here as a direction of travel rather than as something an Indian team should expect to reproduce.
Why it breaks completely in India
In India, the inputs that make decile targeting possible largely do not exist at prescriber level. The dominant commercial data infrastructure — secondary sales audit, stockist and retail panel data from providers such as AIOCD AWACS — operates at distributor, stockist and chemist level, not at doctor level. Sales are visible by territory and by outlet, not by prescriber. So a decile built in India is usually not built from that doctor's prescriptions at all; it is built from a representative's estimate of that doctor's prescriptions, which is a different instrument with different failure modes.
This is the single most important thing to understand before designing an Indian segmentation, and it is absent from essentially all published guidance. The practical consequences are worth stating plainly.
- Rep-rated potential is subjective and gameable. A rating that influences territory targets and call plans is a rating the rater has an interest in. Ratings tend to correlate with relationship strength and with effort already invested, not with untapped opportunity.
- It is systematically conservative about doctors you do not yet know. A rep cannot rate a doctor they have never met, so unknown doctors default to low ratings, and low ratings mean no calls. The segmentation quietly encodes and perpetuates the existing coverage pattern.
- It cannot detect movement. The NBRx signal that makes US behavioural segmentation powerful — a doctor who just started prescribing — has no direct Indian equivalent in most therapy areas.
- It is not auditable. When a target list is questioned, "the rep said so" is not a defensible basis, and under UCPMP scrutiny that matters more than it used to.
None of this means Indian segmentation is hopeless. It means the potential dimension must be modelled from proxies rather than measured from scripts — which is a solvable problem, and the subject of the next section.
What you actually have — the Indian signal inventory
Before choosing a technique, inventory the signals honestly. Most Indian brand teams have more than they use and less than the literature assumes.
| Signal | Where it comes from | Strength | What it can tell you |
|---|---|---|---|
| Secondary sales by territory and outlet | Stockist and retail audit panels | Strong at territory level, silent at doctor level | Where the volume is, not who wrote it. Useful as a territory-level ceiling on any doctor-level estimate |
| Chemist-proximity signal | Sales at outlets near a doctor's clinic | Moderate, and underused | A real proxy for prescribing. Sales lift at the three chemists around a clinic is genuine evidence about that doctor |
| Rep call and feedback data | CRM, field app | Strong on activity, weak on outcome | What has been invested in this doctor, and what the rep observed. Use as an input, never as the potential score itself |
| Digital engagement | Email, WhatsApp, content, webinars | Strong, first-party, objective | Genuine interest, channel preference and topic affinity. The most under-used reliable signal in Indian pharma |
| Sampling and requests | Sample management, medical information requests | Strong | Intent. A doctor requesting information is displaying something a rating cannot capture |
| Patient support programme enrolment | PSP records where they exist | Very strong where available | Actual patients on therapy attributable to that prescriber — the closest thing to script-level evidence available in India |
| Practice characteristics | Doctor 360 — institution type, bed count, catchment, footfall | Moderate, stable | Structural potential. A consultant at a 400-bed hospital has different capacity from a solo clinic |
| Qualification and sub-specialty | Registration and verification data | Strong, stable | Clinical fit. The cheapest available accuracy gain and the most commonly ignored |
| Influence signals | Publications, speaking, society roles, peer networks | Moderate | Whose prescribing changes other people's |
The two signals worth building on first
Digital engagement and patient support programme data.
Both are first-party, objective, timestamped and already in your systems. Neither depends on a rep's judgement. Neither can be gamed by the person being measured. And in most Indian brand teams both are collected diligently and used for almost nothing beyond campaign reporting.
A segmentation that combines modelled potential with observed digital behaviour is dramatically better than a rep-rated list, and it can be built from data you already own without licensing anything new. That is the fastest available improvement in Indian HCP targeting, and it is sitting unused in the marketing platform.
How to segment doctors for pharma marketing — the six steps
- Start from the decision, not the data. Write down what will change as a result. "Which doctors get a rep visit versus digital only", "which get the switch message versus the initiation message", "where does the launch effort go first". A segmentation without a named decision behind it will not be used, because nobody knows what to do differently on Monday.
- Inventory the signals you hold against the table above, and mark each as reliable, indicative or absent. This determines technique. A team with strong digital engagement and no PSP data should build a different model from one with the reverse.
- Model potential from proxies. In the absence of prescriber-level scripts, build a potential estimate from practice characteristics, sub-specialty fit, catchment, chemist-proximity signal and PSP evidence. Validate it against something — territory sales, or a sample of doctors where you have better information.
- Layer behaviour on top. Current engagement, direction of travel, channel responsiveness. This is where micro-segmentation separates from ranking: two doctors with identical potential and opposite engagement behaviour need different treatment.
- Cut four to eight segments and name them in plain language. "High potential, digitally engaged, not yet prescribing" is a segment a brand manager and a rep both understand. "Cluster 4" is not.
- Assign one differentiated action per segment, and write it down. Different message, different channel, different frequency, different owner. If two segments get the same action, they are one segment.
Step six is the test that most segmentation projects fail. It is worth applying it early — if you cannot articulate a distinct action for a proposed segment, do not build it.
The five dimensions of micro-segmentation
Deciles use one dimension. Micro-segmentation uses the intersection of several — which is what makes segments small, specific and actionable rather than merely numerous.
| Dimension | Question it answers | How to build it in India | Typical bands |
|---|---|---|---|
| 1. Potential | How much opportunity exists with this doctor? | Modelled from proxies — sub-specialty fit, practice setting and size, catchment, chemist-proximity sales, PSP enrolment. Not rep-rated | High / Medium / Low — three bands, not ten |
| 2. Current behaviour | What are they doing with us today? | First-party: calls, digital engagement, sample requests, information requests, event attendance | Loyal / Occasional / Lapsed / Never engaged |
| 3. Adoption stage | Where are they on the journey with this brand? | Combination of engagement depth and observed or inferred usage | Unaware / Aware / Trialling / Adopting / Advocating |
| 4. Channel preference | How can we actually reach them? | Observed, never assumed — response rates by channel over the last two quarters | Rep-led / Digital-led / Hybrid / Unreachable |
| 5. Influence | Does their prescribing change others'? | Publications, speaking, society roles, referral position, peer network | Influencer / Connected / Independent |
The intersection is the segment. A doctor who is high potential, never engaged, unaware, digital-led and connected is a completely different commercial problem from one who is high potential, loyal, advocating, rep-led and independent — and a decile model puts both in band nine and sends both the same detail aid.
The dimension almost everyone omits
Channel preference — dimension four.
It is the only dimension that determines whether any of the others can be acted on. A high-potential, high-influence, early-adopting doctor who does not accept rep visits and does not open email is, operationally, a low-yield target until you find a channel that works.
It is also the easiest dimension to build honestly, because it is pure first-party observation — response rate by channel over two quarters, no modelling required. Add it before you add anything sophisticated, and expect it to reorder your target list more than any other single change.
How do I segment physicians beyond decile targeting?
Replace the single ranked list with a matrix of potential against current behaviour, then qualify each cell by channel preference. That gives roughly six to nine cells, of which four to eight will be commercially meaningful. Each cell gets its own action, not its own call frequency. The shift from deciles is not that you have more granularity — it is that you stop treating engagement as a volume problem and start treating it as several different problems that happen to share a field force.
A worked version of the core matrix, with the action stated for each cell. This is the artefact to hand a brand team.
| **Never engaged** | **Lapsed** | **Occasional** | **Loyal** | |
|---|---|---|---|---|
| High potential | Acquire. Highest-value cell. Clinical fit message, senior rep, whatever channel they accept | Win back. Find out why. Usually a service failure or a competitor switch — diagnose before messaging | Grow. Deepen indication or line of therapy. This is where most incremental volume sits | Protect. Low frequency, high value. Advocacy and peer programmes, not detailing |
| Medium potential | Test cheaply. Digital only until a signal appears. Do not spend rep time here | Selective. Re-engage only where the lapse was recent and the cause is fixable | Nurture digitally. Automated, consistent, low cost per contact | Maintain. Low-touch, mostly digital. Do not over-invest |
| Low potential | Ignore, deliberately. Record the decision so it is not revisited every cycle | Ignore. | Digital only. Zero rep time | Service, do not sell. Keep them supplied and satisfied |
Two things this matrix does that a decile list cannot. It makes the "ignore, deliberately" decision explicit and recorded — most target lists never formally exclude anyone, so low-potential doctors quietly absorb rep time every cycle. And it separates acquire from grow from protect, which are three different messages that a frequency-based model collapses into one.
The adoption ladder
Dimension three deserves its own treatment, because it is the most useful concept in this article for launch brands and the most frequently misapplied.
The adoption ladder places each doctor on a journey — unaware, aware, trialling, adopting, advocating — and the operating principle is that you can only move a doctor one rung at a time. Messaging a doctor who has never heard of the brand as though they are deciding whether to expand usage does not accelerate them; it produces no response, which then reads as low potential.
| Rung | What is true | The job | The message that fails |
|---|---|---|---|
| Unaware | Does not know the brand or the mechanism | Create awareness of the clinical problem, not the brand | Comparative efficacy data. They have no frame to place it in |
| Aware | Knows it exists, no experience, no conviction | Build clinical rationale. Peer evidence works here | Dosing and access detail. Too early |
| Trialling | Has used it in one or two patients | Remove friction and confirm the decision. Support, access, follow-up | New efficacy claims. They are evaluating, not reconsidering |
| Adopting | Uses it routinely in the core indication | Expand indication, line of therapy or patient type | Basic efficacy. They are past it and it reads as not knowing them |
| Advocating | Recommends it to peers | Enable them. Speaking, peer programmes, advisory roles | Detailing. Actively counterproductive at this rung |
For launches in specialty markets with a small defined universe, the adopter lifecycle framing — innovators, early adopters, early majority, late majority, laggards — is a useful overlay, because seeding evangelists early matters disproportionately when the total prescriber base is a few thousand rather than a few hundred thousand.
The common misapplication is treating the ladder as a segmentation on its own. It is one dimension. A doctor at the trialling rung with low potential and a doctor at the trialling rung with high potential and strong peer influence require very different investment, and the ladder alone cannot tell them apart.
AI-driven micro-segmentation examples in pharma
Concrete rather than aspirational, and honest about which of these need data an Indian team may not hold.
| Example | What the model does | Data needed | Available in India? |
|---|---|---|---|
| Rising-star detection | Identifies prescribers whose new-to-brand trend is inflecting upward, reportedly 60 to 90 days before conventional reporting surfaces them | Prescriber-level NBRx time series | No, not directly. The Indian analogue is inflection in digital engagement and sample or information requests — weaker, but real and first-party |
| Explainable propensity scoring | Scores each doctor with a driver breakdown — for example around 30% prescribing behaviour, 25% engagement history, 20% access, 15% peer influence, 10% specialty and site of care | Mixed. Degrades gracefully when a driver is missing | Partly. Drop the Rx driver and reweight. The explainability is the point — a score a brand manager cannot interrogate will not be trusted |
| Threshold and movement alerts | Triggers when a doctor crosses a defined behavioural boundary — moving several deciles in 90 days, or a loyal prescriber going quiet against their own baseline | Any consistent behavioural time series | Yes. Works on engagement data alone. The single most transferable technique in this table |
| Look-alike modelling | Finds doctors resembling your best adopters across practice, behaviour and profile features | A defined seed set and a rich doctor 360 | Yes, and it is particularly valuable in India precisely because it does not require script data. Strong answer to the cold-start problem |
| Behavioural clustering | Groups doctors by engagement pattern rather than by volume, surfacing segments nobody hypothesised | Engagement history across channels | Yes. First-party data is sufficient |
| Uplift modelling | Predicts who will change behaviour because of contact, rather than who will prescribe anyway | Contact history plus outcome, ideally with holdout groups | Partly. Requires disciplined holdouts, which few Indian brands run. Worth building the holdout habit for this reason alone |
| Natural-language segment building | A brand manager describes the segment in plain English and the system builds it, removing the analyst bottleneck | A clean, well-modelled doctor 360 | Yes — and adoption improves sharply when brand teams can self-serve |
The honest framing on AI here
Four of those seven work on first-party data an Indian brand already owns. Movement alerts, look-alike modelling, behavioural clustering and natural-language segment building need no prescription panel at all.
That is worth stating clearly, because the usual conclusion drawn from the data gap is that sophisticated targeting is not possible in India until script data arrives. It is not true. What is not possible is copying the US method. The techniques that depend least on Rx data happen to be among the most useful, and they are unbuilt in most Indian brands not because the data is missing but because nobody has assembled the doctor 360 underneath them.
Which is the real prerequisite, and the subject of a separate article
Choosing the model — and why you should start with rules
A recurring failure is jumping to clustering because it sounds like the sophisticated answer. It usually is not the right first step.
| Approach | When it is right | When it is wrong | Time to first output |
|---|---|---|---|
| Business rules | Almost always first. When the decision is clear and the dimensions are known | When you genuinely do not know what distinguishes your best doctors | Two to three weeks. Explainable, arguable, shippable |
| Unsupervised clustering | Exploratory — when you suspect patterns you have not hypothesised | As the first attempt. Produces mathematically valid segments nobody can act on or name | Six to ten weeks, plus interpretation time that is routinely underestimated |
| Supervised propensity | When you have a defined outcome and enough labelled history | When the outcome variable is itself unreliable — which in India it often is | Eight to twelve weeks |
| Uplift modelling | The most commercially correct approach, when you have holdouts and contact history | Without holdout discipline. It cannot be built retrospectively from data that has none | Twelve weeks-plus, and worth it |
| Hybrid — rules plus model | The realistic destination. Rules define the frame, models score within it | As a starting point before either component is understood | Iterative |
The argument for starting with rules is not that models are worse. It is that a rules-based segmentation can be argued with. A brand manager who disagrees that sub-specialty should outrank practice size can say so, and the discussion improves the segmentation. A clustering output cannot be argued with in the same way, so it is either accepted uncritically or rejected entirely — and in our experience it is usually rejected, quietly, by a field force that carries on using its own list.
Ship rules in three weeks, learn what the segmentation is actually for, then earn the right to model.
How many segments, and how big?
Sizing rules that are boring and load-bearing.
- Four to eight segments. Below four you have grouped rather than segmented. Above eight, reps cannot hold them in mind and the segmentation is used by head office and ignored in the field.
- No segment below about 5% of the addressable universe unless it is a deliberate high-value micro-segment with its own dedicated programme. Small segments fragment execution and cannot be measured — you will never detect a lift on 40 doctors.
- No segment above about 40%. A segment containing nearly half the universe has not distinguished anything and will be treated as the default.
- Every segment needs a distinct action. The test stated earlier, repeated because it is the one that matters: if two segments receive the same action, they are one segment.
- Stability matters more than precision. A segmentation where 30% of doctors move segment every quarter destroys field trust and makes measurement impossible. Target under 10% quarterly churn, and if you exceed it, the boundaries are drawn on noise.
That last point deserves emphasis because it is rarely discussed. A more precise segmentation that reassigns doctors constantly is operationally worse than a slightly cruder one that holds still. Reps build relationships over quarters, not weeks, and a doctor who moves from grow to protect and back again three times a year teaches the field that the model is noise.
Where segmentation projects actually fail
Rarely in the analysis. Almost always in what happens next.
- No differentiated action. Every segment receives the same detail aid at a different frequency. This is the dominant failure, and it means the entire project produced a call-planning input rather than a strategy.
- Segments the field cannot recognise. If a rep cannot tell which segment a doctor is in without opening a system, the segmentation does not exist in the field. Names must describe observable characteristics.
- Built on unverified doctor data. A segmentation on a database with duplicates and stale affiliations assigns the same doctor to two segments and mis-assigns others entirely. Clean first — this is why deduplication precedes segmentation rather than following it.
- No holdout, so no proof. Without a control group, nobody can demonstrate the segmentation worked, so it gets replaced by the next brand manager's approach within eighteen months.
- Refreshed too often or never. Quarterly is usually right. Monthly creates churn; annual means the segmentation is describing a market that has moved.
- Owned by analytics, not by the brand. A segmentation the brand team did not help build is one they will not defend when it constrains something they want to do.
The test to apply before you build anything
Ask the brand manager: "What will a rep do differently on Monday for a doctor in segment three versus segment five?"
If the answer is a frequency — more calls, fewer calls — the segmentation is a call plan and does not need five dimensions and a model to produce it.
If the answer is a different conversation, a different channel, a different piece of evidence or a different owner, you have a real segmentation and it is worth the effort.
Ask it at the start, not at the review.
How to measure whether it worked
Four measures, and one design decision that determines whether any of them mean anything.
The design decision is the holdout. A randomly selected group of doctors within each segment who receive the previous approach. Without it, every result is confounded by market movement, competitor activity and seasonality, and you will spend the review arguing about attribution rather than about the segmentation.
| Measure | What it tells you | Watch for |
|---|---|---|
| Differential response by segment | Whether segments behave differently under differentiated treatment. If they respond identically, the segmentation is not real | The most important measure and the most commonly skipped |
| Lift against holdout | Whether the new approach beat the old one. In the US, well-executed targeting produces new-to-brand lift of roughly 5 to 15 percentage points over unexposed controls | That is a US benchmark on US data. Use your own baseline; do not import the number as a target |
| Segment stability | What share of doctors change segment each quarter | Above 10% suggests the boundaries are drawn on noise rather than on structure |
| Field adoption | What share of calls and campaigns actually followed the segment action | The leading indicator of everything else. Low adoption invalidates the other three measures before you can interpret them |
Report field adoption first and separately. A segmentation with excellent differential response and 30% field adoption has not been tested — it has been tested on a third of the field, and the two-thirds who ignored it are the more interesting finding.
Frequently Asked Questions For Micro-Segmentation of Doctors
Six steps. Start from the decision the segmentation must drive, not from the data. Inventory the signals you actually hold. Model potential from proxies if you do not have prescriber-level prescription data. Layer current behaviour on top of potential. Cut four to eight segments and name them in plain language a rep understands. Then assign one differentiated action per segment. The test that matters: if two segments receive the same action at a different frequency, they are one segment and you have built a call plan rather than a segmentation.
Replace the single ranked list with a matrix of potential against current behaviour, then qualify each cell by channel preference. That produces six to nine cells of which four to eight will be commercially meaningful, and each gets its own action rather than its own call frequency. The shift is conceptual rather than technical: deciles treat engagement as a volume problem, while micro-segmentation treats it as several different problems — acquire, win back, grow, protect — that happen to share a field force. Add a fifth dimension for influence if peer effects matter in your therapy area.
Because the input it requires largely does not exist at prescriber level. Deciles rank doctors by their total prescriptions, and India's dominant commercial data infrastructure — secondary sales audit and retail panel data — operates at distributor, stockist and chemist level rather than at doctor level. So an Indian decile is usually built from a representative's estimate of a doctor's prescribing, which is subjective, correlated with relationship strength rather than untapped opportunity, systematically conservative about doctors the rep has never met, and not auditable when a target list is questioned. Potential has to be modelled from proxies instead of measured from scripts.
More than most teams use. Practice characteristics and sub-specialty from a verified doctor database give you structural potential and clinical fit. Chemist-proximity sales — volume at the outlets nearest a clinic — is a genuine and underused prescribing proxy. Digital engagement across email, WhatsApp, content and webinars is first-party, objective and timestamped. Sample and medical information requests indicate intent. Patient support programme enrolment, where it exists, is the closest thing to script-level evidence available in India. Digital engagement and PSP data are the two strongest signals and the two most commonly left unused.
A model placing each doctor on a journey — unaware, aware, trialling, adopting, advocating — with the operating principle that you can only move a doctor one rung at a time. Its practical value is in message selection: a doctor at the trialling rung needs friction removed and their decision confirmed, not new efficacy claims, while a doctor at the advocating rung needs enablement rather than detailing. The common error is treating the ladder as a segmentation on its own. It is one dimension of five — two doctors on the same rung with different potential and different peer influence warrant very different investment.
Four to eight. Below four you have grouped rather than segmented; above eight, reps cannot hold them in mind and the segmentation is used by head office and ignored in the field. No segment should be below roughly 5% of the addressable universe unless it is a deliberate high-value micro-segment with its own programme, because smaller than that cannot be measured. None should exceed about 40%, because a segment holding nearly half the universe has not distinguished anything. And aim for under 10% quarterly movement between segments — stability matters more than precision, since a model that reassigns doctors constantly destroys field trust.
Eventually, but not first. Start with business rules, because a rules-based segmentation can be argued with — a brand manager who disagrees that sub-specialty should outrank practice size can say so, and the discussion improves the model. Clustering output cannot be interrogated the same way, so it tends to be either accepted uncritically or quietly ignored by a field force that carries on using its own list. Rules ship in two to three weeks and teach you what the segmentation is actually for. The realistic destination is hybrid — rules define the frame, models score within it. Uplift modelling is the most commercially correct approach and needs holdout discipline you should start building now.
Four of the most useful. Movement and threshold alerts work on any consistent behavioural time series, including engagement data alone. Look-alike modelling finds doctors resembling your best adopters from practice and behaviour features, and is especially valuable in India because it directly addresses the cold-start problem for doctors you have never contacted. Behavioural clustering groups doctors by engagement pattern rather than volume. And natural-language segment building lets brand teams self-serve, which improves adoption more than any modelling refinement. The real prerequisite for all four is not prescription data — it is a properly assembled doctor 360 underneath them.
Let's Discuss Your Requirements