Programme at a glance

A company building a consumer product around meal planning and home cooking needed its models to make the kind of judgements a trained cook or a qualified dietitian makes without thinking. What is in this dish. How was it prepared. What can be substituted without ruining it. Those are inference tasks, and they only work if the model has been shown enough correctly interpreted examples.

Food images labelled310,000
Recipes structured92,000
Annotation layers3: ingredient identification, preparation method, attribute tagging
Expert pool62 chefs, registered dietitians and food scientists
Cuisine coverage9 regional cuisines, annotated within their own tradition
Quality signal96 percent expert to expert agreement, against 71 percent for generalists
CadenceWeekly delivery on a 5 day cycle, continuous

Data rich and label poor

The client's own description of its position was that it was data rich and label poor, and that phrase describes most companies that have been operating long enough to accumulate anything. Years of images, recipes, user submissions and scraped text, stored carefully, growing steadily, and almost entirely unusable for supervised training.

Unlabelled data feels like an asset because it was expensive to gather and it sits on a balance sheet of sorts. It behaves like an obligation. It has no training value until someone has interpreted it, and the interpretation is the part that was never budgeted.

Models can be built without it, and the client had done exactly that, reaching respectable accuracy on the easy majority of cases. The ceiling appears where the product actually lives. A meal planning tool that is right about obvious dishes and confidently wrong about the rest is not a shippable product, because users notice errors far more than they notice correct answers.

What a generalist misses in a photograph of food

The instinct is to treat food labelling as easy work. Everyone eats, so everyone is qualified. That instinct is the reason so much culinary and nutritional training data is quietly poor.

A capable generalist annotator will identify the visible components of a dish accurately. What they cannot reliably determine is what was done to those components, and that is usually the label that matters. Whether something was grilled, roasted, pan seared or braised changes the nutritional profile, the preparation time and the substitutions that make sense. Much of that information is present in the image, in browning, in surface texture, in how the fat has rendered, but it is only legible to someone who has cooked.

A generalist can tell you what is in the photograph. Only someone who cooks can reliably tell you what was done to it.

We measured the gap rather than asserting it. On preparation method labels, generalist annotators agreed with expert adjudication 71 percent of the time. Expert to expert agreement on the same items was 96 percent. The generalist labels were not random, which is what makes them dangerous. They were plausible, internally consistent and wrong in a patterned way, which is precisely the kind of error a model learns most efficiently.

Why cuisines were annotated from inside

A dish annotated by someone outside its culinary tradition typically gets the name right and the method wrong, because the visual result is familiar while the technique that produced it is not.

Each of the nine cuisines was therefore annotated by experts who cook within that tradition. Grouping them under one general food expert pool would have produced a dataset most accurate on the cuisines already best represented online, which is the opposite of what a global product needs.

A pipeline, not a batch

The engagement was scoped initially as a labelling project with an end date, and that framing was wrong for the problem.

A product like this generates new unlabelled data continuously, from user submissions, new recipes and new content. Treating annotation as a one time exercise means the labelled set is most current on the day it is delivered and decays from then on, and the client is back in the same position within a year.

So the work was restructured as a standing weekly cadence on a five day cycle, sized to absorb incoming volume rather than to clear a backlog. That changes what the client's engineering team can build. Retraining becomes a scheduled event rather than a project, model performance on new material stops drifting quietly, and the annotation step sits inside the normal engineering pipeline instead of being a procurement exercise every time it is needed.

Guidelines are revised on the same cadence, because expert annotators surface genuine ambiguity continuously, and disagreements between two qualified chefs are treated as items for adjudication and guideline refinement rather than as errors to be scored against someone.

310K Food images labelled across three annotation layers.
62 Credentialed chefs, dietitians and food scientists annotating in their own field.
71% vs 96% Generalist agreement with expert adjudication, against expert to expert agreement.

Outcome and where the approach generalises

The client moved from holding a large archive it could not train on to running a continuous labelled data supply feeding scheduled retraining. The accuracy gains concentrated where they were needed, on the harder half of cases that generalist labelling had been getting confidently wrong.

The pattern here is not specific to food. Any product that promises the judgement of a professional runs into the same wall: the model can only be as good as the expertise encoded in its labels, and expertise cannot be approximated by careful non experts. It applies wherever the label requires a practitioner's reading of the evidence rather than a description of what is visible in it.

The practical test is simple enough to run before committing to a large annotation programme. Take a sample, have it labelled by generalists and independently by qualified practitioners, and measure the disagreement. If the two groups agree closely, generalist annotation is fine and cheaper. If they diverge the way they did here, you have just found the ceiling your model will hit, and you have found it before paying to build a dataset that sits underneath it.

Does your model need expert judgement in its labels?

We build credentialed annotator panels across 40+ domains and 50+ countries, with continuous delivery cadences designed to sit inside your engineering pipeline.

Talk to our AI data team