|
3DVLA: Enhancing Vision-Language-Action Models via 3D Spatial and Instance Understanding
Zhongyu Xia, Yousen Tang, Bingqing Wei, Yongtao Wang
vlarobotic-manipulation3d-perceptionmulti-viewself-supervised
|
A plug-and-play framework that injects multi-view-consistent 3D spatial fusion, object-centric instance tokens, and a masked self-supervised occlusion predictor into pretrained VLAs, achieving 86.0% average on LIBERO-Plus and lifting π₀ to 54.5%/23.2% on RoboTwin 2.0 Easy/Hard. |
2026-05-28 |
2026-06-03 |
Explanation
Original
|
|
AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment
Weijie Kong, Zhian Su, Wei Yu, Huixu Dong
vlaaffordancerepresentation-learningrobotic-manipulationfeature-alignment
|
AffordVLA distills a zero-shot affordance teacher into the intermediate visual features of a π0.5-style VLA via a cosine-similarity feature-alignment loss, steering attention to functional interaction regions without any added inference cost. |
2026-05-17 |
2026-06-03 |
Explanation
Original
|
|
Behavior Cloning of MPC for 3-DOF Robotic Manipulators
Theo Guegan, Wen Jie Dexter Teo
behavior-cloningmodel-predictive-controlrobotic-manipulatorsneural-surrogatesreal-time-control
|
Trains feedforward and recurrent neural networks to imitate an IK+MPC expert on a 3-DOF MuJoCo manipulator, achieving ~3× faster inference (1.1 ms/step) and 84.98% success while showing static MLPs beat temporal variants. |
2026-05-08 |
2026-06-03 |
Explanation
Original
|
|
VLA-REPLICA: A Low-Cost, Reproducible Benchmark for Real-World Evaluation of Vision-Language-Action Models
Alex S. Huang, Jiahui Zhang, Shiqing Tang, Yu Xiang
vlabenchmarkreal-world-roboticsreproducibilitymanipulation
|
A ~$1,050 off-the-shelf real-world benchmark for VLA models: SO-101 arm + RGB-D + light box, 10 tasks, 500 demos, ID/OOD evaluation, and cross-lab reproducibility. |
2026-05-20 |
2026-06-03 |
Explanation
Original
|
|
X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model
Jinliang Zheng, Jianxiong Li, Zhihao Wang, Dongxiu Liu et al.
vlacross-embodimentsoft-promptflow-matchingtransformerrobotics
|
Cross-embodiment VLA that handles heterogeneous robot data via per-source learnable soft prompts, a streamlined Florence-2 + Transformer + flow-matching architecture, and a two-step adapt-to-new-robot recipe. The 0.9B model sets SOTA on 5/6 simulation benchmarks and three real robots. |
2026-01-26 |
2026-06-03 |
Explanation
Original
|
|
Attention Is All You Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit et al.
transformersattentionself-attentionnlpmachine-translationdeep-learning
|
Introduces the Transformer — a sequence-transduction model built entirely on attention, dropping recurrence and convolution to enable massive parallelism and new state-of-the-art translation quality. |
2017-06-12 |
2026-06-03 |
Explanation
Original
|