Almost every enquiry we take on physical AI data now arrives with the same shape. The volume is large, the timeline is compressed, and the buyer wants to know what we can start this month rather than next quarter. That urgency is usually read as competitive pressure, and it partly is. But there is a second reason underneath it that gets discussed less, and it is the more important one.
Data of humans doing physical things is not a durable asset in the way most training data is. It has a window, and the window is narrower than the people commissioning it tend to assume.
Why demand looks the way it does right now
Humanoid and general purpose robotic platforms have moved from research demonstrations into active testing, and every group working on them needs the same input: recordings of people carrying out ordinary tasks, seen from roughly where the machine's own sensors will sit. Picking things up. Putting them down. Moving through a kitchen. Handling a phone. Doing the small, unremarkable sequences that make up a day.
Nobody has this at the scale required, because until recently nobody needed it. There is no accumulated archive of first person recordings of household routines the way there is an accumulated archive of text. So the entire field is collecting simultaneously, from a standing start, which is why the market feels the way it does and why timelines have compressed to the point where a programme collecting 500 hours in ten days is a normal request rather than an exceptional one.
Embodied data behaves differently from text
Here is the part that changes the buying decision. A corpus of text collected five years ago is still useful. Language has not moved much, the material is still representative, and it can be reused across model generations. The same is true of most image data. These are durable assets.
Task demonstration data is not durable in the same way, and the reason is not that the recordings degrade. It is that the demand for them is tied to a specific and finite problem: teaching machines a defined set of physical competencies. Once a competency is learned reliably, the marginal value of the ten thousandth hour of humans demonstrating it drops sharply. The data has not become worse. The question it answers has been answered.
Text data is a stock you can draw on for years. Embodied task data is closer to a flow, and it is priced by how much of the problem is still unsolved.
My own read, and I want to be clear that this is a judgement rather than a forecast anyone can evidence, is that the bulk of training for common human tasks gets done within roughly three to five years. Not that robotics is finished at that point, but that the broad, high volume collection of people doing ordinary things will have largely served its purpose, and what remains will be narrower and more specialised.
Which parts of the window close first
The window does not shut all at once, and knowing the order is what makes this actionable.
Common, high frequency tasks close first. They are the easiest to collect, everyone is collecting them, and they are the first competencies any platform will get right. If your programme is built around generic household and office activity, you are working in the most crowded and fastest closing part of the space, and the data you gather in eighteen months will be worth considerably less than the same data gathered now.
What stays valuable much longer is everything specific. Tasks inside particular working environments, industrial or clinical or agricultural. Unusual conditions and awkward variants. Regional differences in how the same task is performed, which are real and almost entirely uncollected. Edge cases generally, since those are the last thing any system learns and the first thing that causes trouble in deployment.
If you are going to collect generic activity data at all, collect it now rather than later, because that segment is depreciating fastest.
If you are choosing where to invest for the longer term, the specific and the difficult hold their value. They are harder to source, which is precisely why they stay scarce.
The supply side is time limited too
There is a second clock running, on capacity rather than demand, and it is easy to miss.
Collection at this scale depends on physical infrastructure and on people. Headsets and capture devices are in limited supply and are allocated to programmes already running. Trained contributors who understand how to record naturally while keeping their hands in frame are not instantly replaceable, and the ones who are good at it are already working. Households willing to have their daily routines recorded are a finite pool in any given market, and once a household has completed a programme it is usually done.
This work is also, by its nature, temporary for the people doing it. It is a genuinely useful earning opportunity while it lasts, and it will not last permanently, which everyone involved understands. The practical consequence for a buyer is that contributor pools are not a resource you can assume will be sitting there when you get around to needing them. They are assembled for programmes and they disperse afterwards.
What to do if you are buying
None of this argues for rushing into a badly specified programme. A collection run built on a vague brief wastes the window rather than using it, and that failure is entirely self inflicted.
What it argues for is sequencing. Decide which parts of your data requirement are competitive with everyone else's and which are genuinely yours, then front load the first category and take your time over the second. Generic activity data is a race you are running against every other group in the field, so treat it as a now decision. Specialised environments, unusual conditions and edge cases are yours to build carefully, because nobody else is fighting you for them and they will still be valuable when the general collection wave has passed.
And be honest with yourself about which one you are actually commissioning. A programme described as physical AI data collection that turns out to be a few hundred hours of generic household activity is not a differentiated asset. It is the same thing several other teams are collecting this quarter, and its value is falling while you scope it.
Mohit Singh Katewa leads the AI data vertical at ConsultBae, where he runs collection, annotation and quality validation across image, video, audio and text, including physical AI and egocentric datasets.
Scoping a physical AI collection programme?
We run egocentric and embodied collection across home, workplace and field environments, with quality validation inside the capture window rather than after it.
Talk to our AI data team


