AdaRoboVLG: A Task-Adaptive Visual-Language Grasping Framework for Different Manipulators
5 hour ago / Read about 0 minute
Author:小编   

Researchers from Huazhong University of Science and Technology, Peking University, Keenon Robotics, and other institutions have jointly proposed the AdaRoboVLG, a task-adaptive visual-language grasping framework. This framework cleverly decouples task requirements from the physical grasping capabilities of manipulators, achieving the synthesis of task understanding and physical grasping through a unified grasping interface. The upper layer transforms task contexts into specific grasping constraints, while the lower layer generates stable grasping postures for different manipulators. The AdaRoboVLG framework integrates three types of foundational model priors—spatial, cognitive, and temporal—to generate flexible, combinable grasping constraints, paired with a foundational grasping strategy that can be reused across different manipulators. New manipulators can be easily integrated by configuring the kinematic mapping and compatible grasping types corresponding to the manipulator. Verified through multiple sets of simulation and real-world experiments, the framework demonstrates exceptional performance in tasks such as grasping in cluttered scenes, language-guided functional grasping, and dynamic target tracking grasping. Without retraining the foundational grasping strategy, new capabilities such as dual-arm sorting and transparent object handling can be extended simply by integrating new modules, providing a practical path for scalable robotic grasping. However, the framework currently has limitations such as upper-layer reasoning errors and perceptual noise, which could be further optimized through directions such as tactile closed-loop control and pre-operation in the future.