Researchers from 15 institutions have released TouchScale, a 500-hour human visual-tactile dataset constructed under a unified sensing and collection process. It aims to explore whether scaling up visuo-tactile data can yield similar benefits to large-scale video data, and translate these into improved robotic manipulation capabilities. The dataset, collected through standardized wearable devices, includes rich interaction records, tasks, objects, and diverse scenarios, ensuring data quality from both hardware and algorithmic perspectives. Experiments demonstrate that, under equivalent training volumes, the TouchScale dataset exhibits superior cross-sensor generalization capabilities, with tactile supervision enhancing the accuracy of visual representations in related tasks. Using only this dataset for pre-training—without human motion labels or hand motion mapping—the average success rate of four real-world robotic contact tasks increased from 22.5% to 57.5%. As data scale expands, tactile prediction performance and robotic manipulation capabilities improve simultaneously, proving the potential for scalable expansion in the visuo-tactile domain and offering a new pathway for robots to acquire transferable physical interaction experience.
