A few weeks ago a client told me he had two quotes in front of him for the same image labeling project. One was nearly five times the other. Same images, same instructions, two numbers that did not look like they belonged to the same job. He wanted to know which vendor was overcharging him.
Neither was. That is the uncomfortable part of data annotation pricing. The same task description can produce two honest quotes that sit a long way apart, because the description is not the work. What you write in a brief and what actually happens to each file are two different things, and the gap between them is where the price lives.
If you have ever tried to compare annotation vendors and walked away more confused than when you started, this is why. So let me break down what actually moves the number, without anyone needing to hand you a rate card.
Why two quotes for the same task can differ by 5x
Start with a request that sounds simple. Match an image to a product and confirm they are the same item. I got exactly this last month: an image of a phone, a link to that phone on a retailer's site, and one instruction, verify they are an exact match.
It sounds like a job for an off-the-shelf vision tool. It is not. If it were that easy, the client would have done it in-house and never called us. The hard cases are the ones that look almost identical but are not. A different storage variant. A different region's model. Last year's edition photographed at a flattering angle. Catching those takes a trained person who knows what to look for, not a quick script.
That is the whole point. A task that reads as one line in a brief can be thirty seconds of work or three minutes of work per item, and the price follows the minutes, not the sentence. Two vendors quoting the same brief are often quoting two different mental pictures of the actual work.
If a labeling task were simple enough to price in a single number, the client would have just done it themselves.
The five things that actually set the price
When I scope a project, five things decide the number. Most of the quotes you receive are really just different assumptions about these five.
The first is modality. Text, image, audio and video do not cost the same. A line of text is quick. A minute of video can carry hundreds of frames, and many of them need attention one by one.
The second is label complexity. A single image can carry one label or forty. Drawing one box around a car is cheap. Marking every pedestrian, lane line and sign across a busy street scene is a different job at a different price, even though both are still called image annotation.
The third is the quality bar. Some projects can ship at good enough. Others, the ones training models that make real decisions, need every file checked, rechecked and validated before it leaves. Each layer of review is real human time, and it is the layer a cheap quote quietly removes.
The fourth is language and region. Collecting and labeling data in one widely spoken language is straightforward. Doing the same across many languages and locales, with reviewers who actually speak them, is harder to staff and gets priced accordingly.
The fifth is volume and turnaround. A steady project run over two months and a rush of the same size compressed into five days are not the same cost, because the fast one means standing up more trained people, sooner.
Why "price per label" hides more than it tells you
The most common way buyers compare vendors is price per label, or price per file. It feels objective. It is also the easiest number to make look good.
A low per-label price means very little on its own, because labels are not equal units. One vendor's label might be a single tag with no review. Another's might include two passes of quality control and a final validation step. Same word, very different product.
When you compare only the unit price, you reward whoever defined the smallest, lightest version of a label. That is not the vendor doing the most careful work. It is usually the one who removed the parts you cannot see in a spreadsheet.
Where a cheap quote gets expensive later
Here is the cost that never shows up in the original quote. Data that arrives cheap and wrong does not stay cheap.
If a batch comes back with errors, someone has to find them, send them back, wait for the fixes and check again. If the errors slip through into training, the model learns them, and you find out much later when its behaviour is off and you are tracing the problem backwards through everything you built on top. The cheapest label in the world is expensive if you cannot trust it.
This is why I am wary of the lowest quote in the stack by reflex. Not because cheap is always bad, but because a price that far below the others usually means a corner was cut somewhere, and the corner is almost always the checking.
How to read a quote so you can compare vendors fairly
You do not need anyone's rate card to read a quote well. You need to make every vendor quote the same job.
Pin down the five factors before you ask for a number. State the modality, the exact label set, the quality bar you need, the languages, and the volume with the deadline. Then ask each vendor to quote against that, and ask them plainly what their price includes in terms of review and rework. The quotes will move much closer together, and the ones that do not will tell you who was quietly assuming less work.
A fair quote is not the lowest one. It is the one where you understand exactly what you are paying for, and the vendor understood exactly what you asked.
Ask what a single label actually includes, and whether quality review is part of the price or billed separately.
Ask how many review passes the data goes through before it reaches you.
Ask who handles errors and rework, and whether that is already inside the number.
Give every vendor the same brief, then compare. A quote is only comparable when the job behind it is the same.
About the author
Mohit Singh Katewa leads the AI data vertical at ConsultBae, where he builds the collection and quality pipelines behind the company's annotation work. He writes about the operational side of training data, the part that happens long before a model ever sees it.
Stop comparing guesses. Scope it once, properly.
Tell us the modality, the label set and the quality bar you need, and we will quote the actual job, not a lighter version of it.
Talk to our data team


