Awesome Embodied VLA / VA / VLN is a curated research collection that organizes state-of-the-art work in embodied AI, focusing on vision-language-action models, vision-language navigation, and multimodal learning approaches for robotics.
The collection addresses the need to navigate a rapidly expanding landscape of embodied AI research by organizing papers and resources across multiple specialized domains. It structures content around key paradigms including vision-language-action models that enable robots to understand and execute tasks from language instructions, world-action models for predictive planning, vision-language navigation for spatial reasoning, vision-action models with diffusion policies, and multimodal large language model approaches to embodied reasoning. The repository also covers physics-aware policy learning, sim-to-real transfer techniques, evaluation benchmarks, and simulation platforms.
Developers and researchers working on robot learning systems should use this collection as a reference for understanding the current state of embodied AI research. It suits anyone building or evaluating vision-language-action systems, navigation pipelines, or multimodal robotic agents who needs to quickly locate relevant papers and approaches. The repository is particularly valuable for those exploring how large language models and vision-language models can drive robotic reasoning and planning, as well as those investigating the transfer of learned policies from simulation to physical robots.
The project maintains active curation with papers organized by recency within each year, with particularly influential works highlighted regardless of publication date. The maintainers explicitly welcome community contributions through pull requests and issues, indicating an open approach to expanding the collection. The repository includes structured sections for surveys, benchmarks, simulators, and related resources, suggesting ongoing effort to provide comprehensive coverage of the embodied AI landscape.