Researchers at LimX Dynamics in Shenzhen, with a co-author at Southern University of Science and Technology, have posted HWAM, a humanoid world action model that predicts the robot’s own body state after a movement as well as the movement it means to make.
The problem it targets is specific to layered humanoid systems. A vision-language-action policy issues a reference action; whole-body control, dynamics, balance and contact then decide what the robot really does. The paper calls the difference the action-execution gap, and argues that a model asked to predict future images has to explain both how the scene changed and how far the robot missed, which muddles the link between an action and its physical result. HWAM makes the post-execution proprioceptive state an explicit prediction target, trained through three paths: a policy path conditioned only on current observations, forward dynamics that predicts future images from actions and body states, and inverse dynamics that reconstructs the joint trajectory from visual change.
On three real-robot tasks using the company’s own LimX OLI humanoid, HWAM reports the highest success rate of the baselines tested, including 70.6 per cent on candy picking against 43.3 per cent for Fast-WAM. Worth keeping in view: these are the authors’ figures, on their own hardware, in a paper marked as under review, with no independent replication.
Source: arXiv:2610.12026, submitted 8 October 2026.
