AI Data
A self-driving model does not fail on the ordinary miles. It fails on the rare moment nobody captured, and that moment is exactly the data that is hardest to find and label.
MK
Mohit Singh Katewa AI Data Lead · 6 min read
Table of Contents
There is a number quoted about self-driving cars that should worry you more than it reassures you. A system that handles the road correctly ninety-nine percent of the time sounds excellent, until you remember what the other one percent contains. It is not one percent of easy driving. It is the child stepping out from behind a parked van, the truck carrying a load that confuses the sensors, the flooded intersection at night with the glare of oncoming headlights.
The rare moments are the whole problem. They are also the moments that are hardest to capture, and hardest to label correctly. That is the real challenge behind autonomous vehicle training data, and it is why this corner of the AI data world is far harder than it looks from the outside.
The 99% problem with self-driving data
Most driving is repetitive. Clear lanes, predictable traffic, good light. A model learns that routine quickly, which is why collecting more of it stops helping almost immediately. Once a system has seen ten thousand hours of ordinary highway driving, the ten thousand and first hour teaches it close to nothing.
The danger lives in the long tail, the unusual events a car might meet once in a million miles but absolutely must handle when it does. You cannot solve that by simply driving more and collecting more. You solve it by finding, capturing and carefully labeling the rare situations, which is slow, deliberate work and the opposite of bulk collection.
Why one camera feed is never enough
A self-driving system does not see the world through a single camera. It fuses several sensors at once, because each one is strong exactly where the others are weak.
Cameras read colour, signs and traffic lights, but struggle at night and in glare. Lidar measures precise distance and shape in three dimensions, but degrades in heavy rain. Radar tracks speed and works in bad weather, but lacks fine detail. The model needs all of them together, and it needs them synchronised in time, so that an object seen by the camera lines up with the same object in the point cloud at the exact same instant.
That synchronisation is where a lot of the difficulty sits. Producing training data is not just labeling three separate streams. It is making sure those streams agree about what is where, and when. An error in alignment quietly teaches the model that two different things are the same object.
The annotation work most people underestimate
When people picture annotation, they picture a box drawn around a car in a photo. Autonomous vehicle data is several steps beyond that.
It involves three-dimensional cuboids placed inside lidar point clouds, semantic labels applied to every region of a scene, and the same object tracked consistently across frames so the model learns how things move, not just where they sit in a single image. A single minute of driving can be thousands of frames across multiple sensors, and every one needs labels that stay consistent with the frames before and after it.
This is why small errors are so costly here. A mislabeled object in one photo is a small mistake. The same object mislabeled the same way across hundreds of consecutive frames teaches the model a wrong pattern, over and over, until it believes it.
A self-driving model does not fail on the easy miles. It fails on the one strange moment nobody thought to label.
Where the real value hides: the long tail
Put those two facts together and the economics of this data look unusual. The routine footage is cheap and, past a point, nearly worthless. The valuable work is in the rare events, and rare events are expensive precisely because they are rare.
Good autonomous vehicle data work is therefore as much about selection as it is about labeling. Someone has to sift through enormous volumes of ordinary driving to surface the handful of genuinely instructive moments, the strange weather, the unusual road users, the ambiguous scene that even a careful human has to think about. A partner who only labels what they are handed misses the harder half of the job, which is deciding what is worth labeling in the first place.
What to look for in a data partner for physical AI
The demand in AI data is shifting noticeably toward physical AI, the models that operate in the real world rather than purely on a screen, and autonomous vehicles are the most demanding example of it. The data needs are a clear step up from flat image labeling, so the questions you ask a partner should rise to meet that.
Ask whether they can work with multiple sensor streams and keep them aligned, not just label one feed. Ask whether they handle three-dimensional annotation and object tracking, not only two-dimensional boxes. Ask how they deal with the ambiguous edge cases, because that is where the real value is. A partner who understands that the goal is the long tail, and not raw volume, is a partner who understands the actual problem.
What a realistic AV data brief includes
Which sensors are involved, and whether their streams need to be time-synchronised against each other.
The annotation types required, including three-dimensional cuboids, segmentation and cross-frame tracking, not just boxes.
The specific edge cases and rare events that matter most for your system's safety.
A clear definition of consistency across frames, since that is where small errors quietly compound.
3Sensor streams commonly fused: camera, lidar, radar
1,000sFrames in a single minute of multi-sensor driving
1%The rare events that hold most of the real value
About the author
Mohit Singh Katewa leads the AI data vertical at ConsultBae, where he runs the collection and quality pipelines behind the company's annotation work. He has watched demand move steadily toward physical AI, and the harder, multi-sensor data it depends on.
Building for the real world needs more than boxes on images.
If your models operate in physical space, tell us the sensors and the edge cases that matter, and we will scope the data work around them.



