A client asks for images of pets. A hundred images arrive. Every file is sharp, correctly sized, properly sourced. Then the client comes back and asks what happened, because the model they are building needs dogs and cats, and most of what they received is fish.

Nothing in that delivery was wrong. It was simply not what they wanted. That gap, between a delivery that is technically correct and one that is actually right, is where most AI data projects quietly go wrong. Not at the labelling stage, where everyone expects trouble, but at the very beginning, in a conversation that took fifteen minutes and should have taken an hour.

What clients mean when they say high quality data

When a company says it needs high quality data, what it usually expects is data that is one hundred percent correct. What it usually means is whatever the engineer told the person who is now talking to you, passed along second hand and stripped of the detail that mattered.

So the brief arrives as a sentence. I need images of plants, I am building a model that identifies plant species. That sentence is not wrong, it is just not a specification. It could describe forty different projects with wildly different costs and timelines.

Underneath the phrase high quality data, the real ask is simpler and harder than perfection: understand what we need, then deliver what we need. No dataset at scale is flawless. An error rate of one or two percent is normal, because nothing built by people at volume is perfect. But that tolerance applies to the execution, not to the understanding. If your understanding of the requirement is wrong, the error rate is not one percent. It is the entire batch.

The plants example

A client asks for images of plants. Work starts. Then they realise the common species they listed are already freely downloadable from the internet, which means the dataset teaches their model nothing it could not have learned for nothing.

So the requirement changes. Not plants. These species, photographed in this forest, in these conditions. The opening sentence was identical. The project underneath it is completely different, and so is the price, the timeline and the collection method.

A correct delivery can still be the wrong delivery

The pets example sounds like a joke until it happens to you. If the brief says pets and the dataset that lands is ninety percent fish, the vendor has a defensible position and the client has a useless dataset. Both things are true at once, and that is what makes it expensive.

It matters because a model learns the world it is shown. A dataset skewed heavily toward one class does not simply underperform on the others, it teaches the system that the dominant class is the safe answer whenever an image is ambiguous. You rarely catch this at delivery. You catch it at evaluation, weeks later, when the model behaves strangely and someone works backwards to find out why. By then the cost is not recollection alone. It is the collection window already spent, the annotation hours layered on the wrong files, and a training cycle that has to be rerun.

The data does not have to be one hundred percent correct. Your understanding of what to deliver does.

The question almost nobody asks first

There is one clarification question that resolves more risk than any other, and it is unglamorous. What is the split?

If I am delivering a hundred images to you, do you want forty dogs, forty cats, ten fish and ten birds? Is that the bifurcation you have in mind, or something else entirely? Once that is on the table, the useful questions follow naturally. Is there a split by breed within each class. Is there a split by age. Is there a split by colour, keeping in mind that colour variation is meaningful for some breeds and irrelevant for others. What resolution. What minimum and maximum file size. Which environments and lighting conditions.

None of that is pedantry. Each answer removes a decision that someone would otherwise make silently, on your behalf, at collection time. And silent decisions are the ones that surface as rejected deliveries.

Sometimes the buyer does not know the answers yet, and that is a good outcome rather than a bad one, because the right move is to go back to the engineer who owns the model and get the requirement clarified. Increasingly, teams do this before they contact a vendor at all. Those projects run visibly better than the ones that start with a sentence.

1 to 2% Realistic error tolerance on a large delivery. Understanding the brief carries no such tolerance.
4 Layers worth agreeing before collection: class split, breed, age and colour.
2 Missed quality checks out of twelve is enough to fail an entire delivery.

The most expensive assumption in this work

Most of the time, the difficulty in this work is not in understanding the requirement. It is in not asking about it.

There is a quiet belief in delivery teams that asking a lot of questions makes you look uncertain, and that the smart vendor reads the brief, nods, and gets started. It reads as confidence. It behaves as risk. The team that assumes it does not need to ask is the team that finds out at handover that the client wanted dogs and received cats.

If you are on the buying side, invert this into a signal. The vendor who spends the first week asking uncomfortable, granular questions about your data is not slow or unsure. That is the vendor who will not hand you the wrong dataset in week six. A partner who accepts the first version of your brief without pushing back on any of it is not saving your time. They are deferring the conversation to a point where it costs both of you far more.

Agree the checks before you build them

The same failure repeats at the other end of the project, at the quality gate, and it is worth naming because it is the most common reason a finished delivery gets rejected.

Every data delivery needs a checking mechanism. Yours might have ten steps in it, carefully built and properly tested. If the project actually needed twelve, those two missing steps are what come back to bite you. The data was collected well and checked well against the wrong list.

The fix is the same discipline as the brief, applied later. Before you build the pipeline, tell the client exactly what you intend to check. These are the ten things I am going to verify. Then let them come back and say no, check these twelve. Now you add the two. The gap was never technical. It was a conversation nobody had.

Quality checks also need checking. At volume, an automated workflow can reject the right assets just as easily as it can accept the wrong ones, and an over-tight pipeline is harder to spot than a loose one because rejection looks like rigour. Watch for patterns where rejections spike, pull a small sample of those files, and inspect them by hand. More often than you would like, the rule is wrong rather than the data. Manual intervention is still required at some point in every pipeline, and pretending otherwise is how good data gets thrown away.

If you are commissioning training data, the highest leverage hour of the project is the one before anything is collected. Write down the split, the bifurcations inside it, the file requirements, and the checks you expect to see run. Then hand that over and watch what the vendor does with it. Good data is not a grade. It is a description of what you asked for, made specific enough that two different people reading it would build the same dataset.

About the author

Mohit Singh Katewa leads the AI data vertical at ConsultBae, where he runs collection, annotation and quality validation projects across image, video, audio and text, including physical AI and multilingual datasets.

Planning a training data project?

We help teams turn a vague requirement into a specification that can actually be collected against, then build and validate the dataset end to end.

Talk to our AI data team