Expanding an AI data project into a new language is widely understood to be challenging. What is less understood is that language is not actually the hardest boundary to cross. The harder boundary is script. Collecting and annotating data in a writing system that the tools, the quality processes, and the contributor pool were never built for introduces a set of difficulties that show up in places most teams do not think to anticipate, and they show up after the project is already underway.

A project moving from English to French stays within the same Latin script, and most of the operational machinery carries over. A project moving from English into Devanagari, Arabic, Bengali, Tamil, Thai, or any of the dozens of writing systems that do not share Latin's structure crosses into territory where assumptions that were invisible suddenly stop holding. The script, not the language, is where the real operational work lives.

Why script, not just language, is the real boundary

Language differences are largely about vocabulary, grammar, and meaning, and a competent annotation operation handles those through native-speaker contributors and well-localised guidelines. Script differences are about the fundamental representation of the text, and they affect the tooling, the display, the input methods, the quality checks, and the verification processes in ways that language differences alone do not.

Two languages can share a script and present few script-level challenges to each other. Two languages in entirely different scripts present challenges that have nothing to do with how well the contributors speak the language and everything to do with whether the surrounding infrastructure can handle the writing system at all. This is why script is the boundary that matters operationally, and why projects underestimate it when they think only in terms of language.

The operational challenges of unfamiliar scripts

Several specific challenges appear when a project crosses into a script the operation was not built around.

Tooling that assumes Latin characters. Many annotation platforms and data processing pipelines were built with Latin script as the implicit default. They may technically support other scripts but break in subtle ways: character counts that are wrong, text that displays incorrectly, search and filter functions that do not work as expected, export formats that corrupt non-Latin characters. These problems are not always obvious until the data is being processed and something does not line up.

Right-to-left and complex scripts. Scripts like Arabic and Hebrew run right to left, which affects how text is displayed, edited, and annotated. Scripts like the Indic family have complex character combinations, conjuncts, and diacritics that do not map cleanly onto the character-by-character assumptions built into many tools. Annotation tasks that are straightforward in Latin script become genuinely difficult when the writing system itself does not behave the way the tool expects.

Quality review across unfamiliar writing systems. Quality review depends on reviewers who can actually read and evaluate the content. For an unfamiliar script, the reviewer pool is narrower, and the central coordination team may not be able to spot-check the work themselves the way they could for a script they read. The quality process has to be restructured so that script-level review is handled by people who can genuinely do it, with coordination layers that connect them to the central team.

Contributor verification. Verifying that a contributor genuinely reads and writes a particular script fluently is harder than verifying spoken language ability, and it matters just as much. A contributor who speaks a language fluently but is not fully literate in its script will produce work with errors that a script-literate reviewer would catch but that automated checks and non-literate coordinators would miss.

Two languages in the same script share most of the operational machinery. Two languages in different scripts share almost none of it. The writing system, not the language, is where projects underestimate the work.

Why this is consistently underestimated

The underestimation comes from thinking about expansion in terms of language count rather than script count. A team planning to add ten languages mentally models ten times the language work. If those ten languages span five different scripts, several of which the operation has not worked in before, the actual challenge is shaped by the scripts far more than by the language count, and the planning that focused on language misses it.

The problems also tend to surface late. The contributors are sourced, the guidelines are localised, the work begins, and then the script-level issues appear: a tool that mishandles a conjunct character, a quality check that cannot evaluate the script, an export that corrupts the text. By the time these surface, the project is committed, and fixing them mid-stream is more expensive than designing for them would have been.

What handling unfamiliar scripts well requires

A project crossing into unfamiliar scripts needs the script-level requirements identified during design, not discovered during execution. The tooling has to be verified to handle the specific scripts, including the complex cases like conjuncts and bidirectional text, before collection begins. The contributor pool has to be verified for genuine script literacy, not just spoken fluency. The quality review structure has to include people who can read and evaluate each script, with coordination that connects them to the central team. And the processing and export pipeline has to be tested against the actual scripts to confirm it preserves the characters correctly.

What to check before crossing into an unfamiliar script

Does the tooling genuinely handle the script, including complex characters and bidirectional text, not just nominally support it? Is the contributor pool verified for script literacy, not just spoken fluency? Can the quality review be done by people who actually read the script? Does the processing and export pipeline preserve the script's characters without corruption? Have these been confirmed before collection starts, rather than assumed?

How ConsultBae approaches this

ConsultBae's contributor network covers scripts that many annotation operations are not built for, including the range of Indic scripts and other non-Latin writing systems across the geographies we operate in. The network is part of it, but the operational discipline around the network is what makes script-level work actually deliverable: verifying script literacy, structuring quality review per script, and confirming that the tooling and pipeline handle each writing system correctly before the work begins.

Crossing into an unfamiliar script is one of the points where the gap between an operation that has genuinely done the work and one that is improvising shows most clearly. The difficulties are specific, they surface late if they are not designed for, and they are exactly the kind of thing that a network and an operation built across many writing systems handles as a matter of course rather than as a surprise.

Vanshika Jain works in AI Data Collection and Annotation at ConsultBae, focused on annotation operations and data quality across projects in multiple modalities and domains.

Collecting data across multiple scripts?

ConsultBae's network covers writing systems most operations are not built for. Let us talk about what your project needs.

Talk to us