Project at a glance
A leading global AI company partnered with ConsultBae to support development of next generation physical AI and embodied AI systems. The objective was to collect large scale egocentric, first person recordings using wearable devices across real world indoor environments.
| Total data collected | 500+ hours |
| Project duration | 10 days, onboarding to final delivery |
| Device | PICO wearable headsets |
| Environment scope | Home and office |
| Recording type | Egocentric, first person point of view |
| Collection model | Distributed global workforce |
| Quality outcome | 98 percent approved accuracy |
The challenge: speed without scripted behaviour
The requirement carried a tension inside it. The client needed a heavily accelerated execution model, and at the same time needed the recordings to capture natural human interaction patterns, not performances. Those two goals usually pull against each other. Speed pushes teams toward instructing contributors tightly, because tight instruction is easier to manage at pace. Tight instruction is exactly what produces the stiff, unnatural footage that embodied models learn nothing useful from.
Four constraints defined the work. Large numbers of distributed contributors had to be managed simultaneously inside a strict ten day limit. Recording consistency had to hold across home and office setups that had never been vetted, each with its own lighting. Motion heavy egocentric streams had to stay clear, without accumulating the motion artefacts that first person capture generates so easily. And the accuracy had to be high enough to feed a machine learning pipeline directly.
Speed is not the hard part of a ten day collection window. Holding quality steady while moving that fast is the hard part.
How the programme was run
Contributors were onboarded and trained on days one and two, drawn from an existing distributed workforce rather than recruited from scratch, which is what makes a two day onboarding phase realistic at all.
Device calibration and a pilot ran on day two, before full scale recording opened. This is the single most consequential decision in a compressed timeline. A pilot on day two costs a few hours. Discovering a systematic framing or lighting problem on day eight costs the project. The pilot surfaces the mismatch between how a brief reads and what contributors actually record while there is still time to correct it.
Full scale recording ran from day three to day eight, covering typing and laptop interaction, object picking and placement, mobile phone handling, walking, workspace movement, reading, writing, and unstructured daily household routines. The unstructured material matters more than it looks. A model trained only on named tasks learns tasks. A model that also sees ordinary, undirected activity learns what a home and an office actually look like from behind a person's eyes.
Nobody wears a headset productively for a full working day. Eye strain and fatigue set in, and tired contributors produce exactly the shaky, inattentive footage that fails validation.
The answer is rotation rather than endurance. Contributors work in shifts with others on standby, so a device is handed over rather than left idle, and power backup keeps sessions from being cut short. Throughput comes from device utilisation, not from asking individuals to push through.
The quality layer behind 98 percent
Quality assurance ran from day four to day nine, overlapping recording rather than following it. In a ten day programme this is not a scheduling preference, it is the whole design. Sequential quality assurance means problems are found after the collection window has closed, when the only remedy left is recollection nobody has time for. Running validation alongside capture means a contributor whose lighting is consistently failing gets corrected on day five, not judged on day ten.
A multi layer framework covered four validation areas: video stability, environmental illumination consistency, behaviour authenticity, and structural metadata log validation. Behaviour authenticity is the one that separates a usable embodied dataset from a large one. It is a judgement on whether the person in the recording is doing a task or demonstrating a task, and it is difficult to automate, which is why it needs human reviewers who understand the intent behind the brief.
Outcome and what the dataset supports
The full dataset was delivered on day ten at 98 percent approved accuracy, ready to enter the client's training pipeline without a remediation cycle in between.
The material supports human motion understanding models, physical AI learning pipelines, context aware indoor navigation, and human object interaction recognition. Those capabilities all depend on the same underlying property: enough first person footage of people behaving normally in real interior spaces that a system can form a reliable sense of how humans move through and handle the world.
The transferable lesson is about sequencing rather than scale. Compressed timelines do not fail because contributors work too slowly. They fail because onboarding, calibration and validation get treated as phases that must finish before the next one begins. Pilot early, validate while you are still capturing, and rotate devices instead of exhausting people, and ten days is a workable window rather than a gamble.
Planning a physical AI collection programme?
We run egocentric and embodied data collection across home, office and field environments, with quality validation built into the capture window rather than bolted on after it.
Talk to our AI data team