HoMMI and the case for egocentric data

HoMMI suggests we may not need teleop data to bridge egocentric video to robot performance.


If you'd asked me a year ago whether egocentric data could be the most consequential data source in robotics, I wouldn't have believed you.

That could change this year.

HoMMI, out of Stanford and TRI, makes a strong case that egocentric video integrated with wrist-camera video (UMI) could be the thing that unlocks the next frontier of physical AI for long-horizon mobile manipulation.

The data collection hardware is surprisingly simple: take UMI (with iPhones) and add a third iPhone on a cap. The rig uses ARKit's multi-device collaboration to lock all three phones, two on the grippers and one on the head, into a single shared world frame at 60 Hz.

The output of this rig is synchronized wrist views, an egocentric view, 6-DoF poses, depth, and gripper widths.

The catch is that naively bolting a head camera onto UMI with wrist cameras will likely fail. The paper's RGB-only baseline shows this, posting 0% success on two of the three evaluation tasks.

The reason is a gap between the collected training data and what the robot will see or do at deployment. Humans and robots differ in height, in neck degrees of freedom, and in the fact that a human's own arms sit in the frame where the robot's won't.

HoMMI closes these gaps with three design choices, each detailed in the paper.

The results across three tasks, folding a cloth into a bin, carrying a box across a room to a trolley, and unfolding a mat on a table: 90%, 85%, and 80% success, respectively.

Each task was trained on 100 to 200 human demonstrations, with no robot teleoperation, no co-training, and no fine-tuning on robot data.

One caveat: the closed control loop is not run by the AI model alone. Whole-body IK solvers sit in it too. HoMMI is not end-to-end from pixels to torques with nothing hand-designed in between.

The implication for me is that prior egocentric systems assumed we needed teleop data to bridge the gap to downstream robot performance. This work suggests we may not need it at all.

ActiveUMI, EgoMI, and EgoMimic are similar works showing promising results in the same direction.

The industry is moving fast, and we're excited to help push the frontier of long-horizon mobile manipulation.

If you're collecting egocentric data for robot policies, reach out!