Who is paying the robot's "tuition fees"?

Who is paying the robot's "tuition fees"?

2026-09-02 12:00

Lead-in: Crowdsourced data lowers entry barriers, but validity remains the bottleneck. After decades of household chores, 50-something full-time mother Gao Bo discovered for the first time that her movements could be monetized.

After decades of domestic labor, 50-year-old full-time homemaker Gao Bo first realized her daily actions could generate income.

She lives in Shandong Province and typically cares for her teenage son, making external employment difficult. Now, Gao Bo films her actions—cooking, laundry, cleaning—using her smartphone during routine tasks. According to media reports, she earns around ¥120 per day by working six hours.

These first-person videos, once processed, become training datasets for robots learning household tasks.

Gao Bo is not an isolated case. On social media, similar scenes are increasingly common: a young woman wearing data collection gear cleaning tables at McDonald’s; a chef recording wrist angles and motion trajectories while stir-frying at a street stall...

Embodied data is entering the era of crowdsourcing. Previously, robot data was primarily generated by operators remotely controlling real robots in training centers. Now, with the proliferation of first-person cameras and other no-body data acquisition devices, data collection workstations are no longer limited to centralized facilities—anywhere people perform tasks can become a robot's classroom.

This surge in crowdsourcing stems from a massive data gap in the robotics industry. To operate effectively in the physical world, robots must first "train their brains"—yet current data on physical-world understanding remains severely insufficient.

Surrounding robot data collection and training, a complete industrial chain has rapidly emerged, drawing strong interest from venture capital markets.

Yet beneath the surface excitement, profitability remains elusive. Who will be the first to profit along this chain? And how much of the expanding data production capacity will truly translate into robot capabilities and customer orders?

01 The Entire Industry Is Racing to "Teach" Robots

According to Interact Analysis, as of the end of April 2026, China has already deployed 64 data collection and training centers, with at least 90 additional projects under construction or in planning, including at least 13 centers deploying over a hundred robots each.

▲ Interact Analysis《Inventory of Over 90 Humanoid Robot Data Collection Centers in China》

In thousands of square meters of space, homes, commercial stores, warehouses, and factories are recreated at full scale. Dozens or even hundreds of robots queue up to “attend classes,” with data collectors wearing VR headsets remotely operating robotic arms to repeatedly move boxes, organize goods, and manipulate tools.

On job platforms, embodied intelligence data collectors are now appearing in batches, with some positions requiring no prior experience.

A data collection service provider, previously serving clients like Baidu Maps, launched a team of over 20 people this year dedicated exclusively to dual-camera general-purpose data collection. According to a company representative, two top-tier robot manufacturers and one elderly care real estate developer have approached them for collaboration—“the demand is substantial; we’ve caught the wave.”

This wave is driven by a strategic shift in the robotics industry. As advancements in robot bodies, joints, and whole-body motion control enable running, dancing, and even boxing, the sector is increasingly shifting focus toward models and real-world task execution.

New bottlenecks have emerged: while robots can repeat a single action flawlessly in training environments, success rates drop sharply when objects change, placements shift, or lighting and environments vary—indicating weak generalization capability.

To address this, the robotics industry is attempting to replicate AI large model Scaling Laws: when model architecture and training methods stabilize, performance gains are pursued through increased data volume, parameters, and computational power.

Though this principle remains under validation in robotics, it has already transformed industry expectations regarding data scale.

Industry estimates suggest the total accumulated high-quality training data across sectors currently stands at around 500,000 hours. To achieve intelligent emergence, even 1 million hours may fall short. At this scale, the gap exceeds at least 200-fold.

With rising demand, the critical question becomes: where does the data come from? Traditional real-robot data collection struggles to close this gap. An industry insider revealed that a single trainer typically controls only one robot, often requiring two people to coordinate, yielding just dozens of real-robot data captures per day—because robot operation speed lags far behind human efficiency.

The widespread adoption of no-body data collection in 2026 has lowered the barrier to scaling data production capacity.

▲ Self-variable showcased no-body data collection solutions at the World Robot Conference

This approach emerged as early as 2024, and this year it has transitioned from academic papers and prototypes to integrated products. For example, EGO devices worn on the head (first-person perspective data capture) record body movements; UMI devices (universal manipulation interface) held in hand capture hand motion, rotation, and grasping; wrist-mounted cameras and haptic sensors provide close-up visuals and contact information.

Technological advances unlock supply, while policy amplifies demand. In June 2026, MIIT and SASAC launched the Real-World Training Initiative, mandating ten provinces to select no fewer than 20 key application scenarios each, aiming to establish over 100 high-value use cases by year-end and drive deployment at the ten-thousand-unit scale.

Multiple forces converging have rapidly filled the data collection industry with players from diverse backgrounds.

Hardware companies like Orbbec and Tujian Technology resemble “miners selling shovels.” Orbbec, starting from 3D vision, expanded its 3D camera, calibration, and mass manufacturing capabilities into EGO, UMI, and wrist-mounted cameras; Tujian entered via flexible e-skin technology, using haptic gloves to record contact forces and force variations invisible to video alone.

▲ Video from Orbbec’s official website demonstrates personnel wearing EGO devices for data collection

Companies like Ubtech and Self-variable, which produce robot bodies and models, also enter data collection to secure equipment orders and train datasets. When local governments build data centers, they often bundle purchases of robots, remote operation systems, and training platforms. Ubtech secured two major projects in Huizhou and Hohhot, totaling over ¥130 million. Self-variable simultaneously develops models, hardware, and data collection tools, enabling decisions on next-phase data collection based on model performance—shortening the cycle between data collection, training, and testing.

MiFeng Technology, incubated by Agibot, aims to transform its internal robot data capabilities into an independent platform, offering end-to-end services in data collection, governance, training, and evaluation—positioning itself as the orchestrator of the entire data production pipeline.

Jingdong and Xian Gong Intelligence leverage existing operational scenarios. Jingdong possesses logistics, retail settings, and organizational capabilities, allowing data collection to extend into warehouses, malls, factories, and ordinary households. Xian Gong Intelligence already operates numerous robots in warehouses and factories, hoping these devices generate data while performing tasks. The company disclosed that its robot brain installations exceed 50,000 units, accumulating over 500,000 hours of multi-modal real-robot data.

The entire industry is accelerating the process of “printing textbooks” for robots—but more textbooks do not guarantee better student performance.

02 What Kind of "Textbook" Does a Robot Need?

In June this year, robotics company XDOF shared an experiment.

The team trained a robot to fold T-shirts using standard imitation learning. After training, the robot succeeded 20 out of 20 test attempts. Then, the team added more successful demonstration clips, expecting improved model performance—but success rate dropped to just 2 out of 20, eventually falling to zero.

▲ XDOF’s robot folding T-shirt experiment

The issue lay in seemingly valid videos: although shirts were correctly folded at the end, the process included pauses, hesitation, re-grasping, and ineffective adjustments. Humans watching once perceive it as normal operation; robots, however, absorb the entire sequence—including hesitation—as the standard behavior to emulate.

This mirrors sentiments among domestic practitioners. An industry insider told Denoise NoNoise: “Since this year, data collection demand has surged, but many collected datasets prove useless upon receipt.”

What the industry truly lacks is the ability to judge what robots should learn from data.

For a human action to become a robot capability, the first step is determining what to collect. Lifting boxes demands stability and rhythm; screwing requires precision and force control; folding clothes involves handling deformation, occlusion, and multi-step procedures. Different tasks require distinct sensor configurations—camera positioning, motion trajectory capture, haptics, joint state monitoring—all tailored to the specific objective.

While these elements can be quantified in terms of hours on a “production ledger,” their roles in models differ fundamentally.

After collection, data undergoes initial verification: Was the task completed? Are images obstructed? Did sensors disconnect? Are image and control signals synchronized? Were sensitive content like faces, phone screens, or trade secrets captured?

Some data service providers claim they can ensure at least one valid dataset per three to four collections—a relatively efficient benchmark in the industry.

After cleaning and annotation, data faces a more complex challenge: identical datasets yield vastly different acceptance results across different companies.

Embodied data is deeply tied to robot bodies and model architectures. Six-axis and seven-axis robotic arms differ in degrees of freedom; two-finger grippers and five-finger dexterous hands operate in distinct action spaces; camera placement, sensor configuration, coordinate systems, and control frequencies vary significantly. A single trajectory may train directly in one company’s system but require remapping—or remain unusable—on another robot or model.

Thus, embodied data must meet at least three layers of validity criteria.

The first layer: data collection validity—confirming task completion and file usability.

The second: dataset validity—ensuring cleaned, annotated, and aligned data can enter client training pipelines.

The third: model validity—verifying whether adding this data actually improves robot success rate, generalization, and failure recovery.

Currently, these three validity layers are easily conflated. A supplier delivering 1,000 hours of “valid” data may only satisfy the first two layers; clients, however, expect tangible robot capability improvements.

This reality prevents embodied data from being traded like ordinary commodities. Clients rarely order 1,000 hours of raw data; instead, they request: “Our robot achieves only 60% success rate on this task—what specific scenarios or motions are missing?” Suppliers must then design collection, cleaning, training, and supplementary collection plans.

Validity, therefore, cannot be an inherent attribute of data—it depends on alignment with a specific model, body, and task.

Wang Xingxing, founder of Unitree Robotics, noted: “Every input-output cycle in robotics introduces potential bias and loss. This is a primary reason why current robot models still lack sufficient generalization and task success rates.”

This implies: a direct data decay chain exists between capturing an action and a robot truly mastering it.

Consequently, a misalignment arises in the industrial chain: upfront costs are incurred, yet final training outcomes remain uncertain. There is no formula dictating that robots automatically become smarter after consuming a certain number of hours of data.

When data volume and capability gain cannot be directly correlated, the core dilemma of data collection business emerges: suppliers base pricing on equipment, labor, and collection duration; clients pay only for model outcomes. Who bears the cost of the intermediate losses? This is a financial equation the data collection industry must solve.

03 Survival Challenge for the Data Collection Industry

Tracing the industrial chain, infrastructure investment is the first to see returns. Cameras, gloves, grippers, remote operation systems, and robot bodies can be billed per unit or set; training facilities can be assessed by area, equipment count, and project timeline.

But once data trading enters the picture, uncertainty spikes dramatically.

Multiple practitioners in embodied intelligence told Denoise NoNoise that while some local governments are willing to fund robot purchases and training facility construction, they often demand that robot companies repurchase the resulting data during project negotiations.

The jointly released Research Report on Embodied Intelligence Training Facilities (2026) by the AI Institute of China Academy of Information and Communications Technology notes that training facilities represent heavy asset investment. While data product sales generate revenue fastest, relying solely on data sales cannot cover such heavy investments, leading to extended payback periods.

Based on public data, a professional remote operator working eight hours produces only 2–3 hours of effective data on average. Domestic real-robot data pricing ranges from ¥500 to ¥1,000 per hour; data collectors with technical expertise or on-site requirements earn monthly salaries between ¥8,000 and ¥15,000.

▲ Data collection job listings on recruitment platforms

Requiring data repurchase effectively adds an insurance policy against high capital expenditure.

Yet robot body manufacturers and model companies also face challenges:

EGO and dual-camera device costs continue to decline, making tasks like cleaning, storage, and sorting easy to scale—and thus prone to redundancy. Multiple collectors wiping tables in different rooms increases total duration, but if object types, actions, and environments remain largely unchanged, the model gains little new capability beyond what the data volume suggests.

U.S.-based humanoid unicorn Figure initially tried purchasing data from external suppliers but found it difficult to meet internal model requirements for scale, diversity, and quality. They pivoted to building their own data collection system—Index Platform—using crowdsourcing to gather daily household and work videos from global users. To date, Figure has paid creators $15 million.

▲ Figure 03 robot demonstrating household chores

Real-robot data is also hard to sell. It is tightly coupled with target tasks and robot bodies, produced slowly and at high cost, limiting its buyer pool. An industry insider told Denoise NoNoise: “At procurement stage, model companies often opt for cheaper simulation data first. Only when tasks like box lifting or assembly approach actual delivery do clients supplement with targeted real-robot data.”

These issues are pushing the industry to rethink expansion strategies. The goal is to reduce not just per-hour collection cost, but the total cost required for robots to master a given skill.

One approach is reserving expensive real-robot data for critical stages. Use no-body devices to scale human demonstrations, simulate environmental variations and failure cases in bulk, then rely on real-robot data for task adaptation and final calibration. NVIDIA once generated 780,000 synthetic trajectories from just a few human demonstrations—equivalent to 6,500 hours of human input—using only 11 hours of computation.

Another approach is reducing repeated integration costs across customers. The industry is experimenting with unified data formats, annotation standards, and interface protocols to lower the cost of integrating, converting, and utilizing data from diverse sources. By end-2025, Zhejiang Testing Center standardized data formats and interfaces, merging tens of thousands of data records from 11 national sites onto a single platform, achieving compatibility among real, simulated, and test data.

While these efforts cannot eliminate robot body differences, they can standardize how data is recorded and read—reducing the need for vendors to redefine formats from scratch and helping clients quickly assess whether a dataset fits their training pipeline.

Domestic standardization initiatives have also begun. Since March 2026, multiple embodied intelligence data and training specifications have been published or initiated.

From an industry perspective, the data collection sector—acting as the foundational data infrastructure for robot “brain augmentation”—is poised for explosive growth. Once fundamental model directions converge and consensus forms around data standards, the value of high-quality data may be reassessed and re-priced.

Source: Denoise NoNoise

Disclaimer: Contains third-party opinions, does not constitute financial advice

Share To
X
Telegram
WeChat
QQ
Link
Recommended Reading

HuggingFace's "AI Duck" Goes Viral, Boosting Chinese Chip Supplier Rockchip

18 days ago
HuggingFace's "AI Duck" Goes Viral, Boosting Chinese Chip Supplier Rockchip

Yushu Surges 629% on First Day of Trading, Humanoid Robots Sell 5,500 Units

08-19
Yushu Surges 629% on First Day of Trading, Humanoid Robots Sell 5,500 Units

14 Executives Depart, Horizon Unleashes a Robot Legion

08-11
14 Executives Depart, Horizon Unleashes a Robot Legion

Yushu's 61 Billion IPO: Who Will Make Money?

08-07
Yushu's 61 Billion IPO: Who Will Make Money?

Former Huawei Genius Youth: Building 5,000 Reliable Robots This Year Won't Exceed 3 Companies

08-05
Former Huawei Genius Youth: Building 5,000 Reliable Robots This Year Won't Exceed 3 Companies

China's climbing robotic vacuum cleaner captures 70% of the global market

08-05
China's climbing robotic vacuum cleaner captures 70% of the global market

Musk spent nearly half of the Tesla earnings call discussing AI and robotics, with automotive business accounting for less than one-third of total revenue

08-05
Musk spent nearly half of the Tesla earnings call discussing AI and robotics, with automotive business accounting for less than one-third of total revenue