100K hours of first-person video won't save your model
Passive video vs deliberate demonstration, and what actually improves robot performance.
You can buy 100K hours of first-person video and barely improve a Physical AI model's performance on the robot.
Human egocentric data collection has grown quickly in the last few years. From where I stand, the excitement is grounded in its potential for quick task adaptation, generalization to novel objects, and robotics tasks that only unlock with hand manipulation.
I'm still putting the full picture together, but a few patterns are becoming clear.
Promising results show that downstream robot performance improves with more human data, especially in novel environments and with objects unseen during training. Both are exactly the conditions robot deployment companies care about.
There are two types of egocentric data collected today: passive video and deliberate demonstration.
Passive video is simpler to collect, but rarely action-labelled. It is captured with less intent, which typically results in unplanned head motion and long stretches of idle, irrelevant activity.
Deliberate egocentric demonstrations are collected with specialized hardware that captures video alongside hand positions, which can be mapped onto the robot's actual hand. These datasets are closer to what a buyer would pay for if the target use case is low-shot manipulation adaptation.
Both types of egocentric data have a role, because each teaches a model something different.
Ultimately, large volumes of POV footage are available, but what matters is how much they improve Physical AI model performance on robot tasks.
Running multimodal analysis on petabyte-scale egocentric data to determine its quality means ingesting and processing it frame by frame.
We heard others were running into challenges doing so, which motivated us to build Pareto as open source.
We're excited to help advance dexterous manipulation by providing quality indexing for robot data.
If you're an egocentric data company, send us your gnarliest dataset and we'll show you which parts of it are robot-ready.