A vendor proposal lands on the desk. The headline is a single number: cost per record, cost per hour, cost per image. The buyer compares that number against two or three other proposals and picks. Sometimes that comparison is straightforward. Often it is not, and the difference between a well-priced project and a badly-priced one is not visible from the headline figure alone.
Understanding what actually goes into the cost of AI training data work makes the comparison meaningful. It also makes it possible to evaluate whether a low number is genuine efficiency or hidden cuts that will surface later as quality problems, and whether a higher number reflects real investment in operations or simply margin that the vendor has decided to charge.
Why the per-record number is misleading
Cost-per-record is a useful summary metric for a finished project. It is a misleading basis for vendor comparison upfront because two vendors can quote the same headline number while running very different operations underneath it. One has invested in contributor quality, multi-stage review, and edge case management. The other has cut corners on all of these because they are not visible in the headline price.
Both will deliver a dataset that looks similar at first glance. The differences show up later, in production performance, in rework cycles, in the dataset's ability to handle the cases the buyer did not specifically test for in acceptance. By that point, the cost comparison made on per-record price has produced an outcome where the cheaper option turned out not to be cheaper.
The six cost components underneath any data project
A fairly-priced AI data project includes six cost components, and the proportions between them vary depending on the work but the components themselves are consistent.
Contributor compensation. The actual payment to the people producing the data. This is the most visible component and the one most likely to be optimised when a vendor competes on price. Compressing contributor compensation typically degrades contributor retention and quality, both of which surface as project problems later.
Sourcing and onboarding. Finding the right contributors, verifying their identity and demographic profile, onboarding them onto the project, and calibrating them against the guidelines. This is significant work that does not appear on the per-record line but shapes whether the project starts well or starts with quality problems already built in.
Quality review labour. The cost of reviewers, sample-based quality checks, expert review of flagged cases, and the iteration between collection and quality teams. A vendor with minimal quality investment will price aggressively. The dataset will reflect that choice.
Infrastructure and tooling. The annotation platform, the data storage, the security controls, the metadata management. Building and maintaining this infrastructure is a real cost that gets absorbed across projects. A vendor without it is improvising, which usually shows up in delivery format issues and chain-of-custody gaps.
Project management overhead. The dedicated project lead, the client communication, the cross-team coordination, the documentation. Underfunded project management is one of the most common reasons data projects miss timelines or deliver inconsistent quality. The cost is real and the lack of it is felt by the client.
Margin. The vendor's actual profit on the work. A reasonable margin is healthy for both sides because it is what funds the operational improvement that makes the vendor better over time. A vendor working at near-zero margin is either subsidising the project to win it or cutting elsewhere to keep their numbers viable.
The headline price is the easiest number to compare and the least informative. The six components underneath it are what determines whether the project actually delivers what was promised.
How the mix shifts depending on data type and complexity
The proportions across these components shift meaningfully based on what is being collected.
Generic high-volume tasks have contributor compensation as the largest component. Quality review can be sample-based and lighter touch. Project management is moderate. Margin is competitive because the work is commodified.
Specialist annotation work has contributor compensation higher per record but quality review and project management consume a larger share of the total cost. Edge case handling alone can be a significant line. Margins tend to be higher because the work is operationally harder and fewer vendors can do it well.
Physical AI and on-site collection projects have a different cost structure entirely. Infrastructure, location, and coordination costs are substantial. Contributor compensation is still important but no longer dominant. Project management is the single largest line item in many cases.
A buyer who expects the same cost structure across all of these will misread proposals for the projects whose economics work differently.
Why cheaper vendors are not always cheaper
The hidden costs of a cheap vendor land after the contract is signed. Rework cycles when batches fail quality review. Communication overhead when the vendor's project management is thin. Acceptance disputes when the brief was interpreted differently than the buyer expected. Delays when contributors drop off mid-project because compensation was too low to retain them.
The total cost of a project is the contracted price plus all of the unbudgeted time the buyer's team spent absorbing problems the vendor's pricing structure did not allow them to handle. By that measure, the cheaper headline number frequently produces a higher total spend than a fairly-priced project with a higher upfront figure.
Contributor compensation is at market rate for the geography and skill required, not aggressively compressed.
Quality review investment is visible in the proposal. The vendor can describe the specific review process and the cost it represents.
Project management is a named person with capacity for the engagement, not a generic account function.
Infrastructure costs are absorbed by the vendor through scale, not hidden by being shifted onto the client.
Margin is reasonable. Both sides understand they need the vendor's business to be healthy for the work to be sustainable.
How ConsultBae approaches this
We price projects transparently and we walk clients through the cost structure when they ask. The work is easier to do well when both sides understand what is being paid for. It also helps clients evaluate other proposals more accurately, which sometimes confirms our pricing and sometimes surfaces options that look attractive on the headline number but have gaps in the underlying components.
The best data projects we have run are ones where the buyer understood the economics of the work and the vendor relationship was structured around fair value on both sides. The projects that go badly are usually the ones priced aggressively at the start by someone trying to win a contract they could not deliver on. The market sorts this out eventually, but the cost of the learning lands on the buyer.
Mohit Singh Katewa leads the AI Data vertical at ConsultBae, overseeing data collection, annotation, and quality operations across 100+ countries.
Trying to compare AI data proposals?
ConsultBae prices transparently because the work is easier when both sides understand the economics. Let us walk you through ours.
Talk to us


