Programme at a glance

An artificial intelligence company working in computer vision and natural language processing analyses web page content so that advertising can be placed against relevant material and kept away from material that would damage a brand. The approach avoids tracking individual users, which means the intelligence has to come from the page itself, which means the models have to be very good at reading both images and text in context.

Image throughput240,000 per month at peak
Text throughput65,000 rows per month
Prior in-house ceiling40,000 images or 18,000 text rows per month
Languages6, annotated by native speakers
Turnaround48 hours standard, same day on priority batches
Accuracy97.4 percent against an adjudicated gold set
TeamElastic pool, 55 to 130 annotators by demand
Pipeline2 stage: safety screening, then classification

The ceiling two annotators hit

The company had been running annotation internally with two full time annotators. That team could produce roughly 40,000 images or 18,000 rows of text in a month, working well and consistently. The number was not the problem. The problem was what the number did to everything upstream of it.

When labelling capacity is fixed and small, it stops being a queue and starts being a filter. Research proposals get evaluated on whether the data can be produced rather than on whether the idea is good. Experiments that would need a fresh labelled set do not get proposed, because everyone knows what the answer will be. Scientists start labelling small batches themselves to unblock their own work, which is the most expensive labelling in the organisation and the least consistent, because it is done by people optimising for their own immediate experiment rather than to a shared guideline.

The cost of in house annotation is not the annotators' salaries. It is the experiments the research team stopped proposing.

That is why throughput and turnaround matter more than they appear to on a scoping call. A team that can get 10,000 labelled rows back in two days runs a fundamentally different research process from a team that waits three weeks, even if both eventually receive the same volume. The wait does not just delay the work, it changes which work gets attempted.

Screening before classifying

The pipeline runs in two stages, and keeping them separate is a deliberate design decision rather than a workflow convenience.

Every image passes a safety screen first, checked against the client's policy for the categories that would make an advertising placement unacceptable. Only material that clears the screen moves to the second stage, where it is classified for what it actually contains, at whatever granularity the model needs, from broad object categories through to specific recognisable subjects.

Why the two stages are not merged

Combining them looks more efficient and produces worse data. Screening is a policy judgement with a defensible right answer and a low tolerance for misses. Classification is a descriptive task where the cost of an error is a slightly weaker model.

Merged into one pass, the stricter standard either drags down classification throughput or, more often, the looser standard quietly erodes screening rigour. Separating them lets each stage carry its own guidelines, its own quality bar and its own reviewers.

The screening stage also carries the same duty of care as any content review work. Annotators on that queue encounter material that is genuinely unpleasant, so exposure is capped and rotated and support is available as part of the engagement, for the same reason set out in our trust and safety work: accuracy on a judgement task degrades under sustained exposure well before anyone says so.

Why multilingual classification breaks first

Text classification was the part of the programme that could not be solved by adding capacity, because it was never a volume problem.

Judging whether a passage of text is relevant to a topic, or safe for a brand to sit beside, requires reading it the way a reader in that language would. Tone, idiom, sarcasm and implication carry the judgement, and all four are precisely what survives translation worst. An annotator working through a translation layer will produce confident, consistent, wrong labels, and the confidence is what makes it dangerous, since nothing in the output signals that anything is missing.

So all six languages were annotated by native or fully fluent speakers working in the original text, each with the guideline set localized rather than translated, because a brand safety threshold written for one market does not describe the same line in another. Quality was measured per language against a gold set built in that language, not against an aggregate that lets a strong language mask a weak one.

6x Increase in monthly image annotation capacity over the in-house baseline.
48hr Standard turnaround, with same day delivery on priority batches.
97.4% Accuracy against an adjudicated gold set, measured per language.

Outcome and what changed upstream

Throughput rose roughly sixfold on images and comparably on text, and the elastic pool meant a spike in demand no longer required a hiring decision. Priority batches turn around the same day, which is what allows an experiment to be designed on a Monday and evaluated within the week.

The more consequential change was not in the numbers. Once labelling stopped being scarce, the research team stopped rationing its own ideas, and the scientists stopped annotating. Annotation is a discipline with its own guideline design, calibration and quality measurement, and a researcher doing it between meetings is not doing it well. Giving it to people whose job it is improved consistency and gave back the time it had been consuming.

The general lesson for teams weighing this: measure your in house annotation against the experiments you are not running rather than against a vendor's per unit rate. The salary comparison usually favours keeping it in house. The research velocity comparison rarely does, and that is the one that determines how fast the models actually improve.

Is labelling capacity setting your research agenda?

We run managed image and text annotation with elastic capacity, native language coverage across 100+ countries, and quality measured per language rather than in aggregate.

Talk to our AI data team