Embodied AI Track: Data Cleaning Emerges as the Key Bottleneck in Production Scalability
2 day ago / Read about 0 minute
Author:小编   

This year, the funding scale in the embodied AI sector has surpassed 10 billion yuan. However, the primary constraint on industry growth lies in the data cleaning stage, akin to quality assurance in conventional manufacturing. For example, training a robot to execute a cup-grasping task typically demands around six months of manual data refinement by a single individual, with only 30% of the initial data deemed usable. Gathering effective data through real-world teleoperation is expensive, and the data cleaning process is indispensable. Unlike the highly refined and automated data annotation process for large-scale models, data cleaning in embodied AI involves multimodal temporal stream data generated by robots in the physical environment. This requires a comprehensive understanding of the causal chain throughout the entire operational trajectory. Consequently, refining a single data point is labor-intensive and, due to the necessity for specialized domain knowledge and the severe repercussions of mistakes, it cannot be outsourced or crowdsourced in the manner of large model annotation. At present, this task is primarily undertaken by in-house engineers and graduate students, leading to inefficient data production and legal ambiguities, such as unclear definitions of labor relations. Some companies have started to explore automation in embodied data cleaning, but implementation remains difficult, and a fully automated end-to-end cleaning process has yet to be realized. Furthermore, the industry is grappling with a talent gap, with a significant shortage of full-stack data infrastructure engineers who possess expertise in robotics, signal processing, distributed data engineering, and related fields. Many companies are now elevating their data infrastructure teams to the same level as their algorithm teams, and the bargaining power of data layer companies is on the rise. Data cleaning in embodied AI is still in a manual, artisanal phase, similar to the early days of large model annotation, characterized by higher costs per data hour and significant growth potential. If the complexities of the physical world can be encoded into automated systems, data cleaning has the potential to become the foundational infrastructure of the embodied AI era.