Skip to content
Technology

Who is earning the “tuition” of robots?

After decades of doing housework, Gao Bo, a stay-at-home mother in her 50s, discovered for the first time that her actions could be sold for money.

She lives in Shandong and has to take care of her teenage son, making it difficult for her to work outside the home. Now, Gao Bo uses his mobile phone to film his movements every day while cooking, washing and cleaning. She told the media that she can earn 120 yuan for six hours a day.

After processing, these first-person videos will become training materials for robots to learn housework.

Gao Bo is not an exception. There are more and more similar scenes on social media: in McDonald’s, a girl wears a collection device to clean the table; in a food stall, the chef records the angle and movement trajectory of the wrist while stirring the pot…

According to comprehensive information from Tianyancha Media, embodied data is entering the era of “crowdsourcing”. In the past, robot data were mainly produced by collectors remotely controlled by real machines in training centers. Nowadays, with the popularization of bodyless collection equipment such as first-view cameras, data mining workstations are no longer limited to training centers. The place where everyone works may become a classroom for robots.

Behind this wave of crowdsourcing craze is the huge data gap in the robotics industry. If a robot wants to work in the physical world, it must first “replenish its brain.” However, the current situation is that there is a serious lack of data to understand the physical world.

An industrial chain quickly emerged around the data collection and training of robots, and enthusiasm in the primary market followed.

But behind the excitement, the score is not easy to settle. Who makes money first in this industry chain? How much of the ever-expanding data production capacity can truly turn into robot capabilities and customer orders?

The entire industry is scrambling to teach robots “lessons”

According to Tianyancha Media Comprehensive Information and Interact Analysis data, as of the end of April 2026, there were 64 data collection and training centers in operation across the country, and at least 90 projects under construction and planning, of which at least 13 centers have deployed hundreds of robots.

In an area of ​​several thousand square meters, homes, supermarkets, warehouses and factories have been restored one to one. Dozens or hundreds of robots line up to “take classes”, and collectors wear VR headsets and control robotic arms to repeatedly carry boxes, organize goods and operate tools.

On recruitment websites, embodied intelligent data collectors have also begun to appear in batches, and some positions do not require work experience.

A data collection service provider, which previously served Internet customers such as Baidu Maps, added a team of more than 20 people at the beginning of this year to specialize in binocular collection of general data business. According to the relevant person in charge, two head robot manufacturers and a senior care real estate company have taken the initiative to cooperate. “The demand is still relatively large. It should be said that we have caught up with the trend.”

This boom comes from the shift in the development focus of the robotics industry. With the continuous advancement of the body, joints and whole-body motion control, robots can already run, dance and box. The industry has begun to shift more energy to models and actual tasks.

New bottlenecks have also emerged: the robot can repeatedly complete an action in the training ground, change objects, adjust the placement angle, or enter a room with different light and environment. The success rate may drop rapidly, that is, the robot’s generalization ability is weak.

The robotics industry therefore attempts to copy the Scaling Law of large AI models: when the model architecture and training methods are relatively stable, increased data, parameters and computing power are exchanged for increased capabilities.

This rule is still in the verification stage on robots, but it has changed the industry’s expectations for data scale.

According to comprehensive information from Tianyancha Media and industry estimates, the entire industry has currently accumulated approximately 500,000 hours of high-quality training data, and to achieve the emergence of intelligence, 100 million hours of training data may not be enough. Calculated at this magnitude, the difference between the two is at least about 200 times.

As demand increases, the industry also needs to solve the problem of where the data comes from. It is difficult for traditional real-machine collection to fill this gap. Practitioners revealed that a trainer can usually only control one robot, and a robot often requires the cooperation of two people, and can only complete dozens of real-machine collections a day, because the robot’s “working” speed is much slower than that of humans.

Ontology-less collection, which will become increasingly popular in 2026, has lowered the threshold for expanding data production capacity.

This route will appear as early as 2024, and this year it will begin to move from papers and prototypes to complete sets of products. For example, the EGO device (Note: First-person perspective data device) worn on the person’s head records body movements, and the UMI (Note: Universal Operation Interface) held in the hand records hand movement, rotation, opening and closing; wrist cameras and tactile devices supplement close-range images and contact information.

Technology opens up supply, and policies further amplify construction demand. In June 2026, the Ministry of Industry and Information Technology and the State-owned Assets Supervision and Administration Commission of the State Council launched a special action for real-life training, requiring ten provinces and cities to select no less than 20 key scenarios respectively, and plan to form more than 100 high-value application scenarios by the end of the year to promote the implementation of 10,000-unit scale implementation capabilities.

With multiple forces stacked together, the data mining industry is quickly crowded with companies from different backgrounds.

Hardware companies such as Aobi Zhongguang and Tijian Technology are most like “shovel sellers”. Obi Zhongguang started from 3D vision and extended its 3D camera, calibration and large-scale manufacturing capabilities to EGO, UMI and wrist cameras; Tuzhen started from flexible electronic skin and used tactile gloves to record contact force and force changes that video cannot provide.

Ontology and model companies such as UBTECH and Independent Variables also enter the data acquisition process for equipment orders and training data. When local data mining centers are built, robots, teleoperation equipment and training systems are often packaged and purchased. The two projects in Huizhou and Hohhot that Youbi won the bid for have a total amount of more than 130 million yuan. Independent variables simultaneously develop models, ontology, and data collection tools, which can determine what data to collect in the next round based on model performance, shortening the cycle of collection, training, and testing.

Meifeng Technology, incubated by Zhiyuan, attempts to turn the data capabilities that originally served its own robots into an independent platform to undertake external collection, management, training and evaluation, and act as the organizer of the entire data production chain.

The advantage of JD.com and Xiagong Intelligence is the ready-made scenarios. JD.com has logistics, retail scenarios and personnel organization capabilities, and can extend collection into warehouses, shopping malls, factories and ordinary households; Xiangong Intelligence already has a large number of robots running in warehouses and factories, and hopes to have these devices generate data while working. The company disclosed that its robot brain has more than 50,000 installed units and has accumulated more than 500,000 hours of multi-form real machine data.

The entire industry is accelerating the printing of teaching materials for robots, but having too many teaching materials does not mean that students can learn them.

What kind of “teaching materials” do robots need?

In June this year, robotics company XDOF shared an experiment.

The team is preparing to teach the robot to fold T-shirts. After training using ordinary imitation learning methods, the robot succeeded 20 times in 20 tests. Later, the team continued to add more demonstrations that had completed the task. They thought that the increase in data would make the model perform better, but the success rate dropped to 2 times at first, and finally failed to 20 times.

The problem is that in those videos that seem to be qualified, the clothes are indeed folded in the end, but the process is mixed with pauses, hesitations, re-grabs and ineffective adjustments. Humans look at it once and think that this is just a normal operation; but the robot will take the entire process as it is and learn even the hesitation as a standard answer.

This is close to the experience of domestic practitioners. An industry insider told “NoNoise” that since the beginning of this year, the demand for data collection has expanded rapidly, but after a lot of data was recovered, it was found to be of no use.

What is really scarce in the industry is gradually becoming the ability to judge what robots should learn from it.

If you want to turn a human action into a robot capability, the first step is to determine what to use. Moving boxes requires stability and rhythm, tightening screws requires precision and force control, and folding clothes requires dealing with deformation, occlusion and multi-step operations. Depending on the task definition, the acquisition requirements for cameras, motion trajectories, tactile sensations, and joint states will also change accordingly.

These contents can be converted into hours on the “capacity table”, but they play a completely different role after entering the model.

After the collection is completed, the data first passes the first acceptance check: whether the task is completed, whether the picture is blocked, whether the sensor is disconnected, whether the image and control instructions can be synchronized, whether any faces, mobile phone screens or business secrets are captured.

Some data service providers said that currently they can make at least one of the three or four pieces of collected data valid, which is relatively efficient in the industry.

After cleaning and annotation, the data has to face a more troublesome level: when the same batch of data is in the hands of different companies, the acceptance results may be completely different.

Embodied data is deeply tied to the robot body and model architecture. The degree of freedom of a six-axis robotic arm is different from that of a seven-axis robotic arm. The action spaces of a two-finger gripper and a five-finger dexterous hand are different. The camera position, sensor configuration, coordinate system and control frequency are also different. The same trajectory can be trained directly by one company, but if it is changed to a different robot or model, it may need to be remapped or even completely unusable.

Therefore, there are at least three levels of valid standards for embodied data.

The first layer is that the collection is valid, confirming that the task is completed and the files are available; the second layer is that the data set is valid, confirming that the data has been cleaned, annotated and aligned, and can enter the customer’s training pipeline; the third layer is that the model is valid, after adding this batch of data, whether the success rate, generalization ability and failure recovery of the robot are really improved.

At present, these three levels of “validity” are easily confused. When a supplier delivers 1,000 hours of valid data, it may only have completed the first two levels; what customers expect is an increase in robot capabilities.

This also determines that embodied data is temporarily difficult to trade like ordinary commodities. It is difficult for customers to just place an order for 1,000 hours of data. A more common requirement is: the success rate of a certain robot in performing a specific task is only 60%. Suppliers need to determine which scenes and actions are missing, and then design collection, cleaning, training and supplementary acquisition plans.

Validity is therefore hardly a fixed property of the data itself; it depends on whether the data can match a model, a body, and a task.

Wang Xingxing, founder of Yushu, once pointed out: For robots, every time input and output are performed, deviations and losses may occur. This is an important reason why the generalization ability and task success rate of current robot models are still insufficient.

This means that there is a chain of data loss between the time the action is captured and the robot actually learns it.

As a result, a dislocation has occurred in the industry chain: the previous costs have been paid, but there is still uncertainty about the final training effect. There is no formula for a robot to automatically become smart after being fed enough hours.

When data output and capacity gain cannot be directly converted, problems in the data acquisition business also arise: suppliers invest costs based on equipment, labor and collection time, but customers are only willing to pay for model effects. Who will bear the intermediate losses has become an account that the data mining industry must settle.

Survival issues in the data mining industry

Looking along the industrial chain, the first thing to be “received” is the money for laying infrastructure. Cameras, gloves, grippers, remote operating systems and robot bodies can be settled on a per-unit basis and per set. The training ground can also be accepted based on area, equipment quantity and construction period.

Once it reaches the data transaction stage, uncertainty increases sharply.

Many embodied intelligence practitioners revealed to NoNoise that some places are willing to pay to purchase robots and build training grounds, but will require robot companies to buy back data during project negotiations.

The “Research Report on Embodied Intelligent Training Fields (2026)” jointly released by the Institute of Artificial Intelligence of the China Academy of Information and Communications Technology pointed out that training fields are heavy asset investments. Although the sale of data products generates revenue the fastest, it is difficult to cover heavy asset investments by selling data alone, and the return period is long.

According to public information estimates, a professional remote operator can only produce 2 to 3 hours of effective data on average after working for 8 hours. The price quoted for domestic real-machine data is about 500 to 1,000 yuan per hour, and the monthly salary of some data collectors who require technical background or on-site presence reaches 8,000 to 15,000 yuan.

Requiring data repurchase is equivalent to adding an “insurance policy” to high investment costs.

But robot body manufacturers and model companies also have difficulties——

The cost of Ego and binocular equipment has been declining, and general tasks such as cleaning, storage and sorting are the easiest to expand and the most prone to duplication. Different collectors cleaning tables in different rooms does increase the time; if the changes in objects, actions, and environments are limited, the new capabilities for the model may not be as much as the new data set seems.

Figure, the American humanoid robot unicorn, once tried to purchase data from external suppliers, but later found that it was difficult for the data to meet the requirements of its own model for data scale, diversity and quality. So it turned to its own data collection system, the Index platform, to collect daily housework or work videos from ordinary users around the world through crowdsourcing. Currently, Figure has paid $15 million to creators.

Real machine data is not easy to sell either. It is more closely integrated with the target task and the robot body, is slower and more expensive to produce, and has fewer customers to buy it from. A practitioner told NoNoise that model companies often choose lower-cost simulation data first when it comes to the procurement process. Only when tasks such as moving boxes and assembly are close to actual delivery will customers add targeted real machine data.

These problems are forcing the industry to change the way it expands production. What the industry needs to reduce is the extension from the cost of acquisition per hour to the total cost of a robot learning a capability.

One approach is to save the expensive real machine for where it is needed most. For example, using bodyless acquisition equipment to expand the scale of human demonstrations, using simulation to batch create environmental changes and failure situations, and then using real machine data to complete task adaptation and final calibration. NVIDIA once used a small number of human demonstrations to generate 780,000 synthetic trajectories, which is equivalent to 6,500 hours of manual demonstration data. The entire generation process only took 11 hours.

Another approach is to reduce the cost of repeated adaptation between different customers. The industry is trying to unify data formats, annotation specifications and interface standards to reduce the costs of collecting, converting and using data from different sources. By the end of 2025, the Zhejiang Testing Center will standardize data formats and interfaces to integrate thousands of pieces of data collected at 11 points across the country into the same platform to achieve compatibility of real, simulated and test data.

These attempts cannot eliminate the physical differences of robots, but they can first unify how data is recorded and read, reducing the work of service providers in defining formats from scratch, and making it easier for customers to judge whether a batch of data can enter their own training pipelines.

Domestic standard construction has also been launched. Since March 2026, a number of embodied intelligence data and training specifications have been released or launched.

From an industry perspective, the data collection industry, as the data infrastructure part for robots to “replenish their brains”, is on the eve of an explosion. In the future, once the technical direction of basic models converges and the industry reaches a consensus on data standards, the value of high-quality data may be repriced.

About Us · 關於我們