The question of which type of annotator a project needs is usually treated as a budget question. Generalist annotators cost less per hour. Domain experts cost significantly more. When the budget is tight, the decision defaults to generalists. When quality is paramount, the decision defaults to domain experts. Neither heuristic is actually useful, because the cost of getting the decision wrong in either direction consistently exceeds the cost of applying a proper framework to get it right.
Using domain experts on tasks that generalists can handle equally well wastes budget and slows delivery without improving the dataset. Using generalists on tasks that require domain judgment produces a dataset with errors that look correct on the surface, pass quality review, and degrade model performance in production in ways that are difficult to trace back to their source. Both mistakes are common. Both are avoidable with a clearer decision process.
What generalist annotators are genuinely good at
Generalist annotators perform well on tasks where the correct answer is visible to any careful, attentive reader who has been properly trained on the annotation guidelines. The defining characteristic of these tasks is that domain knowledge does not change the output. A careful generalist and a domain expert, given the same guidelines and the same data point, will produce the same label.
High-volume image classification with clear visual categories. Sentiment labelling on unambiguous text samples. Bounding box annotation on objects with well-defined boundaries. Transcription of clearly recorded audio. Named entity recognition on text where the entities are unambiguous. These are tasks where the quality ceiling is determined by attention, consistency, and guideline adherence, not by subject matter knowledge. Generalist annotators who are well-selected and properly trained can reach that ceiling without any domain background.
For these tasks, using domain experts adds cost without adding quality. The domain expert brings judgment they have no opportunity to apply because the task does not require it.
What domain experts provide that generalists cannot
Domain experts become necessary when the correct annotation requires knowledge that cannot be transferred through a guideline document. The distinction is not complexity in the sense of difficulty. It is complexity in the sense of judgment: situations where two annotators following the same guidelines will produce different outputs because the right output depends on understanding that the guidelines can describe but not fully convey.
A radiologist annotating a medical scan can recognise that a feature flagged for annotation is a known imaging artefact rather than a pathological finding. A generalist following the same guidelines will annotate it as a pathological finding because the guideline says to annotate features with those visual characteristics, and the guideline cannot enumerate every artefact the annotator might encounter. The error looks correct. It passes quality review. It trains the model to recognise artefacts as pathology.
A contracts lawyer reviewing AI-generated legal text can identify that a clause summary is technically accurate but omits a qualification that changes its legal meaning. A generalist assessing the same output against a quality rubric will pass it because the rubric cannot capture every legal nuance the evaluator might need to apply.
Domain expertise matters when the gap between what the guidelines say and what the task requires can only be bridged by someone who knows the subject well enough to operate in that gap correctly.
The question is not whether the task is complex. It is whether the correct answer requires knowledge that guidelines alone cannot transfer. If it does, a generalist will produce errors that look like correct annotations.
The decision framework: four questions
Before assigning annotator type to a task, four questions determine the right answer.
Can the correct annotation be determined from the data and the guidelines alone? If a careful, intelligent person with no background in the domain could read the guideline and produce the correct annotation reliably, a generalist annotator is sufficient. If producing the correct annotation requires knowledge the guideline cannot fully capture, a domain expert is required.
What does an error look like and how visible is it? If errors in this task are visible on review, producing incorrect labels that a quality reviewer will catch, a generalist with strong quality review processes can work. If errors are invisible on review, producing annotations that look correct but are substantively wrong in ways only a domain expert would recognise, generalist annotation will pass quality checks and embed errors into the dataset that surface only in model behaviour.
How frequent are edge cases and how consequential are they? If the data is largely straightforward with infrequent edge cases, generalists can handle the volume with a domain expert review layer for escalated cases. If edge cases are frequent and consequential, the annotation pool needs domain expertise throughout rather than only at escalation.
What is the downstream consequence of annotation error? A dataset used for a consumer entertainment model has a different error tolerance than a dataset used to train a medical diagnostic model. The higher the consequence of model error in production, the more the annotation pipeline needs to account for the categories of error that only domain expertise can prevent.
Generalist-appropriate: Image classification with clear visual categories, sentiment labelling on unambiguous samples, bounding box annotation on well-defined objects, audio transcription of clean recordings, basic named entity recognition.
Domain expert required: Medical image annotation, legal or financial document evaluation, scientific literature classification, code quality assessment, clinical conversation annotation, any task where the error type is invisible to a non-specialist reviewer.
Hybrid appropriate: Large-volume tasks with a specialist edge case layer, projects where generalists handle clear-cut cases and domain experts handle flagged ambiguous ones, multi-stage pipelines where domain review follows generalist first-pass.
Where teams get this wrong
The most common mistake is using generalists on tasks that sit just outside their competence boundary. The task looks straightforward on the surface: the guidelines are clear, the categories are defined, the volume is manageable. The problem is that the data contains a non-trivial proportion of edge cases where the guidelines do not clearly apply and the correct answer requires judgment the annotator does not have.
Generalists resolve those cases with their best guess, which is not random but is also not reliable. The errors are distributed unevenly across the dataset, concentrated in exactly the cases where correct annotation matters most, and they are invisible in quality metrics because they passed review. The model learns from them anyway.
The second mistake is the reverse: treating domain expert annotation as the default for all quality-sensitive work, regardless of whether the specific task requires domain knowledge. This is an expensive form of caution that does not improve the dataset on the tasks where generalists would have performed identically. The budget spent on domain expert rates for clear-cut annotation is budget not available for the specialist work that actually requires it.
The hybrid model
Most well-designed annotation projects do not use one annotator type exclusively. They use a structure where the type of annotator is matched to the type of decision being made at each stage of the pipeline.
A common and effective structure is generalist first-pass annotation on the full dataset, with a flagging mechanism for cases that fall outside the clear-cut guidelines, followed by domain expert review of the flagged cases and a sample of the unflagged cases to validate that the generalist judgments are holding up. This structure captures most of the cost efficiency of generalist annotation while applying domain expertise precisely where the task requires it.
The design of the flagging mechanism matters significantly. If the guidelines do not clearly specify what counts as an edge case worth escalating, generalists will either over-escalate, routing work to domain experts that they could have handled themselves, or under-escalate, resolving cases silently that should have gone to review. Getting this right requires careful attention to guideline design before the project begins.
How ConsultBae approaches this
At ConsultBae, annotator type is a project design decision made at the scoping stage, not a default applied uniformly across the work. We map the task requirements against the four questions, identify where the generalist-to-domain-expert boundary falls for the specific dataset and use case, and design the pipeline accordingly.
For projects that span both types of work, we build hybrid pipelines that route annotation decisions to the right level. We have a generalist annotator pool for high-volume work and a subject matter expert network across 40-plus domains for the tasks that require it. The two pools are not interchangeable and they are not treated as such.
The right annotator type is not the cheaper one or the more cautious one. It is the one whose capability matches what the task actually requires. Getting that decision right at the start is one of the highest-leverage choices in annotation project design.
Vanshika Jain works in AI Data Collection and Annotation at ConsultBae, focused on annotation operations and data quality across projects in multiple modalities and domains.
Scoping an annotation project?
ConsultBae designs annotation pipelines where the annotator type is matched to what each task actually requires. Let us talk about what your project needs.
Talk to us


