There are cars on the road in parts of the United States carrying four passengers with no steering wheel and no brake pedal in front of any of them. The system drives. The people ride. Most of the coverage treats this as a milestone, and it is. It is also worth sitting with for a second, because something quietly changed in that car that has nothing to do with the software.
For the entire history of driving, the risk was yours and you knew it. With a combustion engine, a gearbox and a set of brake pads, you had a rough sense of how reliable your vehicle was on any given morning. If the car was not stopping the way it used to, you could get the pads changed, or you could decide not to and carry that risk knowingly. Either way the decision sat with you.
The trade nobody explicitly agreed to
A driverless taxi takes that decision away. Not maliciously, and not without giving something back, but completely. The passenger has no mechanism to intervene, no feel for the condition of the system, and no way to judge whether today is a good day for the model or a strange one.
This is not an argument against autonomy. It is a description of what actually happened, and it did not happen suddenly. Automatic transmissions removed a decision. Electric drivetrains removed the maintenance signals you used to feel through the pedals. Lane keeping and emergency braking removed a few more. Each of those was sold as convenience, and each of them moved a small amount of responsibility from the human to the machine. The driverless vehicle simply completes a transfer that has been running for decades.
What matters, if you build these systems, is that the burden of proof moved with it. When the person inside cannot intervene, mostly right stops being an acceptable standard. The system has to be right in the situations nobody thought to describe in advance, because there is no longer a human in the loop to cover for the gap.
What it took for the elevator to become boring
We have done this before with a different machine. When elevators were new, stepping into one was itself the experiment. How much load will it hold. What happens if it holds less than that today. Getting in was a decision a person had to weigh.
Nobody weighs it now. An elevator is the most boring machine in any building, and that boredom is an achievement. It was earned over a long stretch of time through testing, iteration, regulation and hard lessons learned the expensive way. The technology became trustworthy the slow route, in public, with the public forming part of the experiment.
Every machine we now consider safe passed through a period where it was not. The elevator, the passenger aircraft, the seatbelt-era car. In each case the maturity curve was real and it had to be paid for by someone.
The open question for physical AI is not whether it will get there. It is whether the curve has to be paid for the same way it was last time.
Version ten is a data problem, not a time problem
I would be happy to get into version ten of the robotaxi. It is the first nine versions I would like to avoid.
That instinct is close to universal, and it contains the whole problem. Nobody volunteers for version one of a safety critical machine, yet every safety critical machine has a version one, and versions two through nine have to happen somewhere.
The only real question is where they happen. A version can be burned down in simulation and against collected, annotated real world data, quietly, before anything is deployed. Or it can be burned down on a live road with passengers inside, publicly, one incident at a time. The engineering work is similar. The cost of being wrong is not remotely similar.
This is the actual case for serious investment in physical AI training data, and it is not really an accuracy argument. It is a sequencing argument. Every failure mode you can capture, label and train against in advance is a version you never have to ship. Data is how you skip the versions you do not want the public to experience.
The data that moves you along the curve
Here is where most physical AI data programmes quietly underperform. The instinct when a system is not ready is to collect more data, and more data usually means more of what the system already handles well. Clear conditions. Standard tasks. Cooperative humans behaving predictably.
That material is cheap and abundant precisely because it is common, and the model is already competent at it. It moves you almost nowhere along the curve. The versions you are trying to skip do not fail on the routine. They fail on the three seconds that look nothing like the rest of the recording.
For embodied and egocentric collection, capturing that tail is harder than it sounds. You need people performing real tasks naturally rather than robotically, with both hands staying in frame, across genuinely different environments rather than one room dressed three ways. And you need the awkward variants on purpose. The dropped object, the obstructed reach, the ambiguous instruction, the unusual lighting, the person who does the task in the wrong order because that is how people actually behave.
A dataset that is ninety five percent clean routine, with edge cases arriving only by accident, teaches a system that the world is routine. That is exactly the belief you do not want inside a machine that people cannot override.
What this should change about what you buy
If you are commissioning data for anything embodied or safety critical, the useful evaluation question is not how many hours a vendor can deliver. It is how much of the tail those hours contain, and whether the tail got there deliberately.
Ask how edge cases were sourced, by design or by luck. Ask how rare event coverage is measured, and what the target balance actually is, because a number nobody can state is a number nobody is managing. Ask how many genuinely distinct environments were used. Ask what annotator agreement looked like on the ambiguous frames specifically, since that is where labelling quality either holds or quietly collapses, and it is the part of the dataset your model most needs to be right about.
It is worth applying the same thinking one level up. Your first collection round is also a version one. Treat it like a pilot, look hard at what it failed to capture, and fix the specification before you scale it rather than after. The elevator became boring eventually. Physical AI will too. The only decision still open is how much of that gets learned in public.
Mohit Singh Katewa leads the AI data vertical at ConsultBae, where he runs collection, annotation and quality validation across image, video, audio and text, including physical AI and egocentric datasets.
Building something people cannot override?
We design and run physical AI data collection built around edge case coverage, not raw volume, with quality validation at every stage.
Talk to our AI data team


