simmediummanipulation-datametric · varies

Joint-Aligned Latent Action: Towards Scalable VLA Pretraining in the Wild

Description

Despite progress, Vision-Language-Action models (VLAs) are limited by a scarcity of large-scale, diverse robot data. While human manipulation videos offer a rich alternative, existing methods are forced to choose between small, precisely-labeled datasets and vast in-the-wild footage with unreliable hand tracking labels. We present JALA, a pretraining framework that learns Jointly-Aligned Latent Actions. JALA bypasses full visual dynamic reconstruction, instead learns a predictive action embeddin

Source

http://arxiv.org/abs/2602.21736v1