Training data collection is the binding constraint on humanoid robot performance. Unlike language models, which train on text drawn from the internet, robots learning physical tasks need labeled, embodied demonstrations gathered in the real world, and assembling those demonstrations at scale has proven slow. Several Chinese AI startups are now attacking that bottleneck directly, each testing a distinct method for collecting and applying robot training data more efficiently.
The mechanism behind that slowness is structural. Humanoid robots learning to fold laundry or handle components on a production line require physical demonstration data: records of hand position, grip force, body movement, and task sequencing, each captured during actual physical activity. Every training example must be physically produced, logged, and labeled, which makes the data pipeline a function of human time and physical throughput rather than compute. Scale in language model training is primarily a compute problem. For humanoid robots, it is a collection problem, and that distinction changes what kind of company can solve it.
That is the constraint CNBC's The China Connection newsletter places at the center of the current race. Several startups in China are experimenting with different approaches to gathering and applying that training data more efficiently, the newsletter reports. The variety of strategies in play is itself informative: the field has not converged on a single dominant pipeline, and the companies now forming around the problem are competing on an open engineering question with direct commercial stakes for every humanoid robot manufacturer waiting on data to activate hardware that already exists on the factory floor.
The concentration of this startup activity in China tracks where humanoid robot hardware development is densest. Where that hardware is being built and deployed at scale, demand for training data infrastructure follows, and that demand is what the newsletter says is now pulling AI companies into the space.
Completing human tasks in roughly human time is the benchmark the hardware has not yet reached. The number of different collection approaches being tested simultaneously signals where the field actually stands: the data problem for humanoid robots is genuinely unsolved, and the startups now organizing around it are working without a confirmed answer.