Physical AI data collection: Where are today's biggest bottlenecks?

I agree with Yuuki that hardware is probably the biggest bottleneck, and I’d add that it’s not just about having the sensors, but having a reliable, repeatable data collection pipeline.

One reason is that today’s VLA models are surprisingly sensitive to camera setup. We recently ran a small experiment showing how even relatively minor viewpoint changes can significantly affect model performance. ( Why VLAs are so viewpoint brittle? ) That means camera placement, calibration, and synchronization aren’t just implementation details, they’re effectively part of the data distribution the model is trained on.

Because of that, I don’t think a truly standardized setup has emerged yet. Most teams still end up building custom rigs around their embodiment, sensor suite, and tasks. Even if two teams are both collecting teleoperation data, differences in cameras, calibration, control interfaces, and robot kinematics can make the resulting datasets much less interchangeable than they appear.

1 Like