Programme at a glance
A graphic design software company had built a multimodal model that generates original images from text prompts, and was extending it well beyond English. Generating an image in another language is the easy part. Generating one that a designer in that market would actually put in front of a client is a different problem, and it is not one the model builder can assess internally.
| Image evaluations delivered | 9,400 |
| Prompts localized | 2,600 |
| Languages | 18 |
| Locales | 22, separating regional variants of the same language |
| Reviewer pool | 140, native fluency plus design background |
| Scoring | 6 dimensions, 2 independent reviewers per image, adjudication on splits |
| Control | Every prompt also generated and scored in English |
Translated prompts are not localized prompts
A prompt is an instruction, and translating an instruction accurately can still change what it asks for. A phrase describing a seasonal scene carries a different set of visual expectations depending on which hemisphere the reader is in. A reference to a family gathering implies a different table, a different room and a different number of people depending on where it lands. The translation can be word perfect and the resulting image still wrong.
So the first phase was localization by native speakers, with transcreation wherever direct translation would have produced a technically correct prompt describing the wrong picture. Around 11 percent of prompts needed that treatment. The rest translated cleanly, and identifying which was which is itself expert work, because the prompts that need transcreation do not look problematic in English.
Separating language from locale mattered more than expected. Two markets sharing a language do not share stylistic conventions, colour associations or what reads as contemporary rather than dated. Treating them as one bucket produces evaluation data that averages two audiences into an aesthetic that satisfies neither, which is why the programme ran 22 locales against 18 languages rather than collapsing them.
Text rendered inside generated images degrades sharply outside Latin scripts. Character forms break, joining behaviour in connected scripts fails, and diacritics land in the wrong place.
An English speaking reviewer sees plausible looking lettering and scores the image as fine. A native reader sees nonsense text in the middle of a design. This failure is close to invisible without native reviewers, and it disproportionately affects exactly the markets a multilingual launch is meant to serve.
The reviewer nobody can hire quickly
This programme needed people at an unusual intersection: native fluency in the target locale, plus enough graphic design literacy to judge composition, typography, colour and whether an output meets a professional standard.
Either qualification alone is straightforward to source. A native speaker without design training rates whether the image looks nice. A designer working outside their own culture rates craft accurately and misses cultural fit entirely. Both produce clean looking scores and neither answers the question the client is asking.
Assembling 140 people who hold both, distributed across 22 locales, is a specialist recruitment problem before it is a data problem. It runs on the same sourcing and verification infrastructure our staffing business uses, which is what makes a pool like this assemblable in weeks rather than quarters.
Why every prompt was also run in English
The design decision that did the most work in this programme was the least visible one. Every prompt was generated and scored in English as well as in its target locale.
Without that control, a low score is uninterpretable. If a Japanese prompt produces a poor image, there are two entirely different explanations. The model may be weaker at Japanese conditioning, which is a model problem. Or the localized prompt may have drifted from the original intent, which is a localization problem. Those two findings lead to opposite remediation, and a scorecard without a control cannot distinguish them.
Without an English control, you cannot tell whether the model failed or the prompt did. Those look identical on a scorecard and require opposite fixes.
Scoring ran across six dimensions rather than an overall verdict: prompt adherence, cultural appropriateness, design quality, stylistic fit for the locale, text rendering accuracy, and format compliance. Two independent reviewers scored each image, with splits going to adjudication. As with any subjective task, reviewer disagreement was treated as signal, because two qualified native designers disagreeing usually means the prompt itself is ambiguous in that locale.
Outcome and what it changed
The client received a comparative, locale segmented view of model performance with the English baseline sitting alongside every score, which let their team route findings correctly. Conditioning weaknesses went to the model. Prompt drift went back into localization. Script rendering problems went to a specific, known defect rather than into a general sense that quality was lower outside English.
The wider point is that multilingual capability in a generative model is not a translation layer bolted onto the front. The model produces different quality in different languages, prompts change meaning as they move, and both effects appear in the same output. Separating them requires reviewers inside the culture and a control that holds the prompt constant.
Cultural relevance is often treated as a soft criterion, the part of the evaluation that gets trimmed when timelines compress. In a design product it is the entire proposition. An image that is technically well composed and culturally off is not a partial success. It is unusable, and the user who receives it does not file a bug, they stop using the feature.
Launching a generative model outside English?
We build native reviewer panels across 100+ countries with the domain literacy your evaluation actually needs, including prompt localization, transcreation and controlled multilingual benchmarking.
Talk to our AI data team