HiPHI: A Large-Scale Benchmark for High-Precision Human Motion and Object Interaction

IEEE Spectrum and Wiley have released a sponsored white paper from Noitom Robotics describing HiPHI, a large-scale dataset for humanoid robot learning. The paper is presented as an overview for robotics researchers and engineers and is available as a free download.
According to the description, HiPHI is a 617.5-hour whole-body human motion dataset captured with optical motion capture at sub-millimeter accuracy. It includes 245.7 hours of human-object interaction data with synchronized object trajectories and meshes.
The dataset's coverage is organized using FrameNet, a linguistic framework for human action. The paper also introduces a benchmark suite for measuring motion diversity and interaction grounding, and reports results from policies trained on the dataset and deployed on a physical Unitree G1 humanoid robot. The described aims include supporting reinforcement learning and sim-to-real transfer for tasks such as carrying, pushing, and pulling.
The paper addresses what it characterizes as a data gap: internet video shows diverse behavior but cannot capture precise physical states, while laboratory motion capture systems record accurate movement but usually cover only a narrow set of actions, according to the description.
Noitom Robotics is identified as the sponsor and builds ModalityNet, described as a human-centric data substrate for embodied AI. The white paper is published in partnership with IEEE Spectrum Magazine. Registration is required to access content on the hub.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
This White Paper gives robotics researchers and engineers an overview of a new large-scale motion capture dataset built to close the data gap limiting humanoid robot learning. It also shows how policies trained on the dataset transfer to a real humanoid robot. What you will learn about: Why humanoid robot learning, a central problem in embodied AI and Physical AI, needs data that internet video and existing motion capture datasets cannot provide. How FrameNet , a linguistic framework for human action, can guide mot