Programme at a glance

A children's video platform needed two connected things. Search had to return results that were relevant and appropriate for the age of the person searching, and newly published content had to be checked against the platform's own safety policy before it reached that audience. Both problems scale with publishing volume, which meant the review capacity had to scale with it too.

Review volume60,000 video items per month, ongoing
Age bands4, each with its own threshold
Reviewer pool210 trained reviewers
Coverage11 markets, 7 languages
Quality standard96 percent agreement against an adjudicated gold set
Escalation3 tiers, highest severity to specialist review in under 2 hours
Exposure control4 hour daily cap on sensitive queues, with rotation

Age appropriate is not a single standard

The most common design error in child safety review is treating children as one audience. A single safe for kids label collapses an enormous developmental range into one threshold, and that threshold is wrong at both ends. Set it for the youngest users and older children get a search experience so restricted they leave the platform for one with no protections at all. Set it for the oldest and the youngest are exposed to material they are not ready for.

So the programme ran four defined age bands, each with its own written threshold covering not just prohibited material but themes, intensity and complexity that are fine for one band and unsuitable for the one below it. A reviewer is never asked whether a video is appropriate. They are asked whether it is appropriate for a specified band, which is an answerable question.

Why banding changes the search problem too

Relevance and appropriateness are separate judgements and they can point in opposite directions. A result can be the most relevant answer to a query and still be wrong for the age band that submitted it.

Scoring the two dimensions separately lets the platform tune ranking and filtering independently. Collapsing them into one score hides which of the two is failing.

What a thumbnail check misses

The reason this work resists automation is that unsuitable material on a children's platform does not usually announce itself. Content can look entirely appropriate in its thumbnail and opening seconds while carrying something quite different elsewhere in the file, including in audio that never appears in the visual track at all.

That shapes the review method more than any other factor. Review covers the full item rather than a sample, treats the audio track as a first class signal rather than an afterthought, and is carried out by someone fluent in the language actually spoken in the video. A reviewer working outside their language can assess the visuals and will miss the rest entirely, and the gap is invisible in their output because their scores look confident and complete.

Guidelines carry the weight here, and they were written by a dedicated quality team rather than inherited from a general policy document. Each guideline carries worked examples, including deliberately near threshold ones, and every reviewer clears a calibration set before touching live queues. Guidelines are revised on a fixed cycle, because new patterns appear continuously and a rule written six months ago will not describe them.

Anything above the lowest severity tier leaves the general queue immediately and goes to a smaller specialist group, with the highest tier reaching them in under two hours. Volume reviewers are not asked to make the hardest calls, and they are not asked to sit with the worst material while deciding.

Protecting the people doing the reviewing

This is the part of content moderation that most case studies leave out, and it belongs in the operating model rather than in a policy appendix.

Reviewers on a child safety queue encounter material that is genuinely distressing, and the effect is cumulative rather than incidental. Time on sensitive queues is capped at four hours a day with rotation onto other work, exposure is tracked per person rather than per shift, and access to professional psychological support is provided as part of the engagement and not on request.

A moderation programme that exhausts its reviewers does not only have an ethics problem. It has a quality problem, and the quality problem shows up first.

The two concerns are the same concern. Accuracy on a judgement task degrades with fatigue and distress before anyone reports feeling unable to continue, and it degrades in a specific direction, toward faster decisions and toward the default option. On a safety queue the default is approval. A programme that pushes reviewers past their limit is quietly increasing the rate at which unsuitable material passes through, and it will not see that in its throughput numbers.

60,000 Video items reviewed each month against age banded policy.
4 Age bands scored separately, each with its own written threshold.
96% Reviewer agreement against an adjudicated gold set.

Outcome and ongoing operation

The programme is continuous rather than a project with an end date, which is the correct shape for this work. Publishing does not stop, patterns change, and a filter validated once is a filter validated against last quarter's content.

Alongside the review volume, the work gives the platform two things it did not have. Age banded scoring shows where filtering is actually failing rather than that it is failing somewhere, and the labelled output feeds back into automated classification, so the systems handling the first pass improve on exactly the cases they currently get wrong.

The honest framing for any programme like this is that the goal is not a perfect filter, because there is not one. The goal is a review layer that catches what automation cannot, escalates the hardest cases to people equipped to judge them, learns fast enough to keep pace with what is being published, and does all of that without using up the people doing it.

Running a trust and safety review programme?

We build trained, native language review teams across 100+ countries, with calibrated guidelines, tiered escalation and reviewer wellbeing designed into the roster from day one.

Talk to our AI data team