The Showlab team from the National University of Singapore has proposed the Show-Harness solution, which designs a set of semantic action interfaces between vision-language models (VLMs) and robot control. Instead of requiring the model to directly output low-level continuous control signals, this interface transforms actions into semantic units that the model can understand and reason about. Subsequently, these semantic units are converted into locally compliant control commands with safety constraints by dedicated interpreters for different robots, forming a continuous perception-reasoning-execution loop. The team also developed the GUMI graphical control interface, which unifies the action semantics for human teleoperation, web-based agent control, and direct VLM control, facilitating cross-platform demonstration data collection. The Show-Harness solution enables closed-source cutting-edge VLMs to adapt to real robot closed loop (closed-loop) control in a zero-shot manner without fine-tuning, while small open-source VLMs require only minimal parameter adjustments for adaptation. Moreover, the same interface can be used across different robot bodies, tasks, and environments, achieving significantly higher success rates than traditional baselines in cross-task, cross-environment, and cross-body generalization tests on real robots. Additionally, the solution excels in simulation-to-reality transfer and physical-semantic adaptability tests. Research indicates that clear and stable semantic-physical effect conventions are key to enabling action implementation through the interface, providing a new interface paradigm for extending foundational intelligence from the digital world to physical embodied systems.
