Abstract
Scaling Generalist Policy Models with heterogeneous data is limited by the lack of unified, low-noise action supervision. Human egocentric videos are abundant, but only a small fraction comes with high-quality hand-action labels. Observed world transitions offer a common source of action-related supervision across data sources. We introduce WLA3 (World Latent Action Modeling for Semantics, Dynamics, and Kinematics), a unified Generalist Policy Model framework built around representations learned by a World Latent Action Model (WLAM). WLAM encodes synchronized camera views and available embodiment-state changes into a compact local latent action and a richer transition feature. WLA3 reuses these representations across semantics, dynamics, and kinematics through latent action-conditioned world modeling, Semantic Latent Aggregate supervision, and joint latent–robot action generation. The final 32D latent action reaches 67.89% average classification accuracy on LARYBench, while WLA3 achieves 81.9% average success across six real-robot tasks versus 66.2% for π0.5.