Programme at a glance

A company selling speech recognition for medical documentation was entering a new market. Its product was already localized, its workflows were proven, and its recognition accuracy in its home languages was strong enough to have made it a standard in the hospitals that used it. None of that carried across, because a speech model trained on one language does not transfer to another, and the vocabulary the product depends on is not the vocabulary in a dictionary.

Person names46,000, with phonetic transcription
Place names19,000
Clinical and pharmaceutical terms7,800
Total entries72,800
Languages4
Contributors90 native speakers across 14 regions
Pronunciation variants2.4 per entry on average
Validation100 percent dual transcription with adjudication on disagreement
Timeline12 weeks, brief to delivery

The gap that privacy compliance creates

Clinical dictation is dense with proper nouns. A physician dictating a note says patient names, referring clinician names, hospital and district names, and drug names, continuously. Those words carry most of the meaning in the record and almost none of them appear in a general language corpus.

They also cannot be harvested from real clinical data, and the reason is the point of the whole project. Health data has to be anonymized before it can be used for development, under data protection regimes across every market worth entering. Anonymization works by removing personally identifying information, which means it removes person names and place names specifically and deliberately.

Anonymization strips out exactly the words a clinical speech model most needs to learn. The more rigorous your privacy pipeline, the larger your proper noun gap.

So a company can hold years of compliant, well governed clinical audio and still have no way to teach its model how a name in the target market is spelled or spoken. The data it is permitted to keep is the data with the hard part cut out. Closing that gap requires building the vocabulary separately, from sources that never contained patient information in the first place.

Why proper nouns are not a rounding error

In most speech applications a misrecognized word is an inconvenience. In clinical documentation, the words most likely to be got wrong are also the ones where being wrong matters most: which patient, which clinician, which facility, which drug.

A system with excellent general accuracy and weak named entity coverage produces notes that read fluently and are wrong in the specific fields a clinician relies on. Buyers in this category evaluate on exactly that.

Why one pronunciation per name is wrong

The instinct on a lexicon build is one entry, one transcription. That produces a dataset that fails in the field.

A name spelled one way is pronounced differently depending on the region the speaker is from, the language they carry their accent from, and in many cases the community the name originates in. Place names are worse, because local pronunciation frequently diverges from the official romanized spelling, and the person dictating uses the local form.

The lexicon therefore carried 2.4 pronunciation variants per entry on average, each transcribed and tagged with the regional context it belongs to rather than flattened into a single canonical form. That tagging is what allows the client to weight variants by deployment region later, instead of shipping one national model that is slightly wrong everywhere.

Getting there required contributors distributed across 14 regions rather than concentrated in one, which is a sourcing requirement more than a linguistic one. A transcription team drawn from a single city will produce a confident, internally consistent lexicon that encodes that city's pronunciation as the standard.

Building and validating the lexicon

Entries were assembled from openly available public sources rather than from any clinical record, which keeps the build clean of patient data by construction rather than by remediation. Name lists were checked for distributional realism, so that the coverage reflects names actually common in the market rather than an alphabetical sample weighted toward the unusual.

Every entry was phonetically transcribed by a native speaker and independently transcribed a second time by a different contributor, with disagreements routed to a linguist for adjudication rather than resolved by majority. On phonetic work the disagreements are informative: two native speakers transcribing the same name differently usually indicates a genuine regional variant, which means the item belongs in the lexicon twice rather than once.

Clinical and pharmaceutical terms were handled separately, reviewed by contributors with healthcare backgrounds, because drug names carry pronunciations that follow neither ordinary phonetic rules nor the spelling, and a general language transcriber will confidently regularise them.

72,800 Phonetically transcribed entries across names, places and clinical terms.
2.4 Regional pronunciation variants per entry, tagged rather than flattened.
12 Weeks from brief to delivered lexicon across four languages.

Outcome and what it unlocked

The delivered lexicon gave the client a vocabulary base it could not have assembled from its own data at any budget, because the constraint was regulatory rather than commercial. That turned a market entry that was blocked into one that was scheduled.

The structure mattered as much as the volume. Because variants were tagged by region rather than merged, the same lexicon supports tuning per deployment, and because the clinical terms were separated from general named entities, the client can extend either set independently as its product moves into new specialties.

The broader lesson generalises past speech. Privacy compliant pipelines are now standard in every regulated industry, and they are working correctly when they remove identifying information. The consequence is that the data a regulated company legally holds is systematically missing the categories its models most need. That gap does not close by collecting more of the same data. It closes by commissioning the missing vocabulary separately, from sources that never held personal information to begin with.

Entering a new language market?

We build phonetic lexicons, named entity sets and speech corpora with native speakers across 100+ countries, structured for regional variation rather than flattened to a single standard.

Talk to our AI data team