Programme at a glance

A global platform working on the personalization of its content feed needed human ratings at scale. The ranking model had to learn which items users found important, useful and worth seeing, and that signal can only come from people. The constraint was that the raters had to represent the platform's actual audience, not a convenient subset of it.

Contributors onboarded600
Markets covered12
Languages9
Time to full ramp12 days
Initial engagement5 week pilot
Data flow7 days a week, including holidays
Quality standard95 percent inter-rater agreement maintained

Why relevance is a representation problem

Every personalization system is trained on somebody's judgement of what is worth seeing. That judgement is not neutral. What counts as important news, useful content or an item worth keeping in the feed varies by country, by language, by age and by the everyday context a person is reading in.

So a rating pool concentrated in a small number of markets does not simply produce less data. It produces a confidently wrong signal. The model learns the preferences of the people who happened to be available and applies them to an audience that does not share them. The failure is invisible in the metrics, because agreement scores can look excellent inside a homogeneous pool. Everyone agrees because everyone is similar.

This is a structural limitation rather than a vendor failing. Rating pools tend to form where contributor recruitment is easiest, and they are difficult to widen after the fact, because expanding into a new market means sourcing, vetting, training and paying people in that market before a single rating arrives.

Inter-rater agreement inside a narrow pool measures how alike your raters are, not how well they represent your users.

Building a pool that mirrors the user base

The work began with a profile rather than a headcount. Before any recruitment, we mapped the demographic and geographic shape the pool needed to have in order to reflect the platform's audience, then treated that profile as the specification to hire against, market by market and language by language.

Sourcing ran through the same infrastructure our staffing business uses, which is the part most data vendors have to build from scratch. Finding, screening and contracting several hundred people across a dozen countries inside two weeks is a recruitment problem before it is a data problem, and it is the reason the ramp took 12 days rather than a quarter.

Onboarding was built as a structured module with visual and interactive components rather than a document to read. For a subjective task this matters more than it does for a mechanical one. Contributors are not learning a rule, they are calibrating a judgement, and calibration comes from worked examples and practice items with feedback, not from a definition.

Why continuous coverage changes the model, not just the schedule

Content consumption does not pause on weekends or public holidays, and the content people engage with on a Sunday is not the content they engage with on a Tuesday.

A pool that only rates during one region's working week trains the model on a partial week. Running coverage across time zones seven days a week closes that gap, and the benefit is representational rather than simply throughput.

Holding quality on a subjective task

Ratings of importance and impact have no answer key. There is no ground truth file to check against, which rules out the validation method most annotation pipelines depend on.

What replaces it is agreement measured against calibrated reference raters, gold items seeded invisibly into live queues, and drift monitoring on individual contributors over time. The pattern worth watching is not the person who is wrong, it is the person whose ratings slowly stop matching their own earlier ratings, which is the signature of fatigue or of a guideline being interpreted differently as edge cases accumulate.

Guidelines themselves needed a revision cycle. On subjective work, ambiguity is discovered rather than anticipated, and each new market surfaced cases the original brief had not considered. We ran a two week revision loop with the client's team, and treated every guideline change as a recalibration event rather than a memo, since a rule that is published but not retrained against simply widens the spread.

600 Contributors sourced, vetted and trained across 12 markets and 9 languages.
12 Days from programme start to full rating volume.
95% Inter-rater agreement held across a task with no ground truth.

From pilot to standing programme

The engagement was scoped as a five week pilot intended to fill a coverage gap. It became a standing programme, and expanded into four further markets on the same operating model.

The reason was less about the ratings themselves than about what a representative pool made possible. Once the pool mapped closely to the real audience, the client could run experiments against segments that had previously been invisible in the training data, and could extend the same rating infrastructure to adjacent problems such as filtering low quality and spam content from the feed.

The transferable lesson is that personalization quality is bounded by who you are able to ask. Model architecture, ranking logic and evaluation frameworks all sit downstream of that constraint. If the pool cannot reach the audience, no amount of tuning further along the pipeline recovers the signal that was never collected.

Building a rating programme that reflects your users?

We source, vet and train contributor pools to a demographic and geographic specification across 100+ countries, with quality frameworks designed for subjective tasks.

Talk to our AI data team