Mifeng Turns Data Collection into Crowdsourcing as Embodied Intelligence Welcomes "Data Riders"

Deep News
Sep 24

On September 23, Mifeng Technology officially launched its data crowdsourcing platform "Mifeng Pai," attempting to open up embodied intelligence data collection to the general public on a large scale.

Mifeng Technology was founded in February 2026, with business operations covering the collection and processing of real-machine data, bodyless data, and simulation data, positioning itself as a third-party data platform serving the entire industry.

Through the "Mifeng Pai" platform, ordinary people can apply for MEgo devices, accept collection tasks via the App, and then receive compensation from the platform based on the verified duration of valid data.

Mifeng has simultaneously opened recruitment for "robot trainers" and, together with more than 50 enterprises and institutions in sectors such as hotels, supermarkets, and manufacturing, launched a scenario data alliance. The company hopes to break apart data production that was originally concentrated in data collection factories and distribute it across more real work and life scenarios.

Over the past year, the answer to "where does data come from" in the embodied intelligence industry has been changing rapidly. The most direct early approach was to concentrate robots in data collection centers, where humans teleoperated real machines to repeatedly complete tasks such as grasping, carrying, and organizing.

But real machines are expensive, collection efficiency is limited, and it is very difficult to rely on linearly increasing robots and collectors to push data scale from millions of hours to tens of millions or even hundreds of millions of hours.

As a result, the industry began to separate more data production from the robot body itself.

Mifeng chose to further open up the production side: allowing more people to enter real scenarios with lightweight devices, and then having the platform uniformly process these scattered data into training data that models can use.

As embodied data begins to attempt crowdsourced production, is it a rapidly expandable data network, or a more complex labor business?

A Meituan for data, how do the numbers work?

For ordinary people, participating in Mifeng Pai roughly requires several steps: download the App, rent a MEgo device, claim tasks, complete specified operations in life or work scenarios, then upload the data, and receive commissions based on valid data duration and withdraw them.

Users can select tasks in the task hall, and ongoing tasks adopt a duration quota system. The MEgo device rents for 39 yuan per day, with a promotional discount price of 19 yuan per day, and the App shows a base task return of approximately 20 yuan per valid hour.

This is not the uniform hourly wage that collectors ultimately receive.

According to the company, compensation consists of a city base rate and dynamic subsidies: the former references income levels in different regions, while the latter depends on task completion efficiency and data quality.

Tasks range from a few minutes to over an hour, and final settlement is based on the verified valid duration.

Once a certain type of data has been collected in sufficient quantity, the platform can stop distributing it; for scarce or harder-to-access scenarios, additional subsidies are used to increase appeal.

Currently open collection tasks cover more than 20 fields including production maintenance, catering services, logistics handling, elderly care, and household organization. Even everyday operations such as counting shuttlecock tubes at badminton halls or restocking drinking cups at convenience service stations have been included in the collection scope.

"Crowdsourcing" easily evokes a variant of the food delivery rider model: the platform connects demand with dispersed workers, then schedules them through task rules and pricing.

But the delivery difficulty of the two is not the same.

Food delivery only requires delivering goods to a specified location, while embodied data must satisfy the model's specific requirements for scenarios, actions, and quality, and dispersed collectors differ greatly in operating habits, professional skills, and work environments.

From a business model perspective, Mifeng believes crowdsourcing can first reduce labor costs on the collection side.

Traditional centralized data collection requires specially hired collectors to repeatedly complete tasks in fixed locations.

The crowdsourcing model instead hopes to embed collection into participants' existing lives and work. What the platform pays for is the incremental data contribution, not the cost of fully employing a full-time collector.

Correspondingly, the complexity of back-end operations also increases.

Model companies' needs must first be broken down into tasks that ordinary people can understand, clarifying action steps, device positions, completion standards, and acceptance rules.

The platform also needs to complete basic training, device pickup, and circulation through offline service stations to minimize operational deviations among different collectors.

Unified devices solve part of the input variation. Data that passes initial screening still needs to go through action segmentation, annotation, trajectory extraction, and manual spot checks before it can become data usable by models.

Mifeng does not simply divide data into "usable" and "unusable," but grades it according to quality and usage value.

Data with obvious capture failures is removed, while the rest may be used separately for pre-training, specific task training, model evaluation, or long-tail scenario learning.

Yao Maoqing, Chairman and CEO of Mifeng Technology, said that more than 95% of data that passes initial screening settlement can ultimately still be utilized after post-processing.

According to figures disclosed by Mifeng, within one month of internal testing, Mifeng Pai reached 20,000 registered users, generated a cumulative 13,000 task submissions, and the highest-earning participant made more than 5,000 yuan in a single month.

The more dispersed the scenarios, the more complex the operations

Mifeng's launch of a crowdsourcing platform at this point is backed by two simultaneous changes in downstream data demand: customers need ever-larger volumes of data, and the scenarios demanded are becoming increasingly granular.

Yao Maoqing said that one million hours is changing from an industry supply target into the training demand of a single customer. Some customers have already raised demand at the tens of millions of hours level, and some leading customers even require suppliers to deliver 50,000 to 100,000 hours per week.

But if the same kind of data is simply produced ten times over, the value of crowdsourcing remains limited.

In the past, embodied data collection was concentrated in home scenarios, with many teams repeatedly collecting tasks such as folding clothes, and this type of data has become saturated. What is becoming scarcer are new skills, new environments, and new operational processes that models have never seen.

Mifeng is in fact expanding scenarios on both the demand side and the supply side.

Mifeng Technology stated that the company regularly holds production plan review meetings to consolidate market orders, first determining which needs can be met by existing inventory, and then allocating the missing portion to different collection paths based on order value and resource conditions.

Yao Maoqing also mentioned that leading customers have already formed some common requirements for real scenarios, spatial precision, and action semantic labels, but specific needs will continue to sink down into different processes and actions.

Centralized data collection centers can replicate standard environments such as kitchens, living rooms, and tabletops, but it is very difficult to reproduce auto repair, hotel services, agricultural production, or specific factory production lines in fixed locations.

Relying on a full-time team to find these scenarios one by one is costly and slow. Crowdsourcing disperses scenario discovery and data collection to people who already work in different industries, making it easier to reach long-tail occupational scenarios and their operational skills.

At the same time, Mifeng Technology hopes to use the "Scenario Data Alliance" to connect scenario parties such as factories, communities, hotels, stores, and logistics centers in advance, resolving issues of safety, authorization, and partnership before customer demand arrives.

Zhang Zhifu, Vice President of the Mifeng Pai business, also mentioned that in the future the platform may match collectors based on their occupation and task history, and open self-declaration of scenarios, with the platform reviewing them before determining whether they match downstream demand.

In his explanation, it is very difficult to list tens of thousands of tasks worth collecting in advance using only the platform's internal team. People truly on the front lines of agriculture, maintenance, and production know better what scenarios and skills they possess.

What Mifeng ultimately wants is not a larger data collection factory, but a scenario network that can be scheduled according to orders.

This also makes operations one of the core capabilities of embodied data companies.

Yao Maoqing judged that once the industry truly enters the stage of scale and profitability, it will "place great importance on operational capability":

It must not only translate model companies' data gaps into tasks and find the corresponding people and scenarios, but also process highly heterogeneous data into unified specifications, while maintaining cost advantages and delivering on schedule.

More importantly, as procurement volumes expand, customers' requirements for supplier scale, quality, and delivery capability are also rising. Based on this, Yao Maoqing judged that embodied data suppliers may further concentrate, possibly even trending toward oligopoly.

Mifeng stated that supporting data production at the tens of millions of hours level requires hundreds of millions of yuan in fixed asset investment in collection devices, personnel settlement, data storage, and computing infrastructure.

The industry chain may in the future see further division of labor among hardware, operations, and data post-processing, but end-to-end coverage of the entire chain is not a business that an ordinary startup team can easily accomplish.

In addition, for model companies that need millions or even tens of millions of hours of data, managing dozens of suppliers at the same time means they must also handle the differences brought by different hardware, formats, and quality systems themselves.

The more suppliers there are, the higher the cost of data cleaning, conversion, and training validation.

In other words, scenarios can be dispersed, but the delivery interface is best centralized.

And whether expanding data volume or connecting scenarios, what is ultimately contested is not just current orders, but occupying the ecological niche of data production and scenario entry points in advance before demand continues to change.

Who will pay for the next hundred million hours?

The Ego bodyless collection data represented by Mifeng Pai crowdsourcing currently mainly serves the pre-training of embodied intelligence companies, world model companies, and general large model companies.

Models learn the relationships among objects, actions, and environments from human operations to expand coverage of different scenarios and skills.

Simply put, bodyless data solves whether the model has "seen enough," while real-machine data solves whether the robot can "actually do it."

Yao Maoqing said that if the goal is only to complete a specific task demo, using the target robot to collect real-machine data will be more direct; when the sample size is small, Ego data does not significantly improve a single task, and its value is more reflected after scale expands, as zero-shot generalization capability when facing new tasks and new environments.

As models begin to enter the physical world to make decisions, what needs to be supplemented is not only more hours, but also richer scenarios, higher spatial precision, and more complete action semantics. The stronger the model capability, the more complex the tasks it can learn and enter, and new data gaps may emerge accordingly.

At least under the current training paradigm, bodyless data is not a temporary substitute for when there are not enough robots.

The longer-cycle variable is that robots themselves begin to produce data.

For already mature tasks such as box moving, sorting, and fixed production line operations, robots in operation can directly record joint states, control commands, and success, failure, and abnormal situations.

This data is naturally consistent with the target body, and as robots are deployed at scale, the relative importance of human collection in mature tasks will decline.

For Mifeng, the more critical question is whether customers will still purchase from third-party platforms over the long term even if the industry still needs large amounts of data.

Once leading robot companies form their own installed base, scenario networks, and data feedback systems, the economics of self-collection will improve, and the right to produce data for mature tasks may also shift to robot operators.

Human data still has value, but that does not automatically mean third-party data collection platforms still have value.

Yao Maoqing positions Mifeng as a "turnkey" data service.

He believes that most embodied companies still tend to concentrate resources on hardware bodies and model algorithm research and development, while leaving heavy-asset, operations-intensive links such as devices, manpower, scenarios, and data operations to professional platforms. Some standardized data can also be sold multiple times across customers, reducing repeated collection of similar data.

How far human crowdsourcing can go ultimately depends on two things: how many worlds robots have not yet entered, and whether third-party platforms can find and produce this data at lower cost than customers themselves.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Most Discussed

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10