The Small-Town Youth Labeling AI Giants

marsbitPublicado em 2026-04-07Última atualização em 2026-04-07

Resumo

In China's hinterland cities like Datong, Shanxi, thousands of young people are working as data annotators—the invisible workforce behind AI development. They perform repetitive tasks like drawing bounding boxes on images or rating AI-generated responses, earning piece-rate wages as low as a few cents per task. These workers, mostly from rural areas or small towns, endure intense labor conditions: strict monitoring, high error tolerance thresholds, and mental exhaustion. Despite the cognitive nature of their work, they are often paid meager salaries, with some earning as little as ¥30 ($4) for a day’s work. As AI industry evolves, even highly educated workers—including master’s graduates—are being drawn into similar precarious freelance roles, evaluating complex AI outputs under vague and shifting standards. Yet the industry is structured through layers of outsourcing, where most profits flow to tech giants like OpenAI and Microsoft, while annotators see dwindling incomes. Worse, as AI models become more self-sufficient, the demand for human annotators is declining. Companies like Li Auto have slashed annotation costs by using AI-powered tools that complete in hours what used to take humans years. These annotators, who helped train the very systems now replacing them, face an uncertain future—a stark contrast to the booming valuations and optimistic narratives of the global AI industry. No one seems to see a problem with any of this.

Datong, Shanxi—once a city propped up by coal—has shaken off its dust and now wields a sharp pickaxe, striking down upon another invisible mine.

In the office buildings of Jinmao International Center in Pingcheng District, there are no more elevator shafts or coal trucks. Instead, thousands of tightly packed computer workstations fill several floors. The Shanghai Runxun Yunzhong Shengu Big Data Smart Service Base occupies entire levels, where thousands of young employees, wearing headphones, stare at screens, clicking, dragging, and boxing.

According to official data, as of November 2025, Datong had put into operation 745,000 servers, attracted 69 call center and data labeling companies, created over 30,000 local jobs, and generated 750 million yuan in output value. In this digital mine, 94% of the workers are local residents.

It’s not just Datong. Among the first batch of data labeling bases designated by the National Data Administration, counties in central and western China like Yonghe in Shanxi, Bijie in Guizhou, and Mengzi in Yunnan are prominently listed. In Yonghe County’s data labeling base, 80% of the employees are women, mostly rural stay-at-home moms or returning youth who couldn’t find suitable jobs.

A hundred years ago, Manchester’s textile mills were filled with landless farmers. Today, the computer screens in these remote counties are manned by young people who found no place in the real economy.

They are engaged in a job that feels both futuristic and primitive—piecework—producing the essential data feed for AI giants in Beijing, Shenzhen, and Silicon Valley.

No one sees anything wrong with this.

The New Assembly Line on the Loess Plateau

At its core, data labeling is about teaching machines to recognize the world.

Self-driving cars need to identify traffic lights and pedestrians; large models need to distinguish cats from dogs. Machines have no innate common sense; humans must first draw boxes on images, telling them “this is a pedestrian,” so that after digesting millions of pictures, they can learn to recognize on their own.

This job doesn’t require advanced degrees—just patience and an index finger that can click incessantly.

In the golden year of 2017, a simple 2D box could fetch over ten cents, with some companies even offering fifty cents. Fast labelers, working over ten hours a day, could earn five to six hundred yuan. In a small town, this was undoubtedly a high-paying, respectable job.

But as large models evolved, the brutal side of this assembly line began to show.

By 2023, the price for simple image labeling had plummeted to 3-4 cents—a drop of over 90%. Even for more complex 3D point cloud images—dense point matrices that require extreme zoom to discern edges—labelers must draw a 3D box in space, encompassing length, width, height, and yaw angle, to tightly wrap around vehicles or pedestrians. Yet such a intricate 3D box earns only five cents.

The direct consequence of the unit price crash is a dramatic increase in labor intensity. To cling to a base salary of two to three thousand yuan a month, labelers must constantly, relentlessly, increase their speed.

This is no easy white-collar job. In many labeling bases, management is stiflingly strict: no phone calls allowed during work hours, phones must be locked in storage compartments. The system meticulously tracks each employee’s mouse movements and idle time. If you stop for more than three minutes, a warning from the backend lashes out like a whip.

Even more crushing is the error rate. The industry’s passing threshold is usually above 95%, with some companies demanding 98%-99%. This means if you draw 100 boxes and just 2 are wrong, the entire image is sent back for rework.

Dynamic images are frame-linked; changing lanes, vehicles get occluded, and labelers must use inference to find them one by one. In 3D point cloud images, any object with over 10 points must be boxed. For a complex parking space project, lines drawn too long, labels missed—quality checks always find flaws. An image being sent back four or five times is commonplace. In the end, after an hour’s work, the pay might be just a few dimes.

A labeler from Hunan posted her settlement slip on social media: after a day’s work, she drew over 700 boxes at 4 cents each, totaling 30.2 yuan.

It’s a profoundly split reality.

On one side, glamorous tech giants at press conferences talk about how AGI will liberate humanity; on the other, young people in counties on the Loess Plateau and southwestern mountains stare at screens for eight to ten hours a day, mechanically drawing boxes—thousands, tens of thousands—so many that at night, their fingers twitch in the air, tracing lane lines in their dreams.

Someone once said, the exterior of artificial intelligence is a luxury car speeding by, but if you open the door, you’ll find a hundred people inside, pedaling bicycles furiously, gritting their teeth.

No one sees anything wrong with this.

Piecework Labor Teaching Machines "How to Love"

As the bottlenecks of image recognition were broken, large models evolved deeper, needing to learn to think, converse, and even show "empathy" like humans.

This gave rise to the most core, yet expensive, part of large model training—RLHF (Reinforcement Learning from Human Feedback).

Simply put, it involves real people scoring AI-generated responses, telling it which answer is better, more aligned with human values and emotional preferences.

ChatGPT seems "human-like" precisely because countless RLHF labelers are teaching it.

On crowdsourcing platforms, such labeling tasks are often priced clearly: 3 to 7 yuan per task. Labelers must assign highly subjective emotional scores to AI responses, judging whether an answer is "warm," "empathetic," or "attentive to the user's emotions."

A low-wage worker, struggling in the mire of reality, with no time to tend to their own emotions, must now serve as the AI's emotional tutor and values judge within the system.

They must break down complex, subtle human emotions like warmth and empathy, forcibly quantizing them into cold scores of 1 to 5. If their scores don’t match the system’s preset standard answers, they are marked as below accuracy standards, deducting from their already meager piece-rate pay.

This is a cognitive voidance. The intricate, profound human emotions, morality, and compassion are being dragged into the algorithm's funnel. In the icy quantification and standardization, they are drained of their last warmth. While you marvel at the cyber behemoth on screen learning to write poetry, compose music, offer comfort, even donning a sentimental skin; outside the screen, those once-vibrant humans, through daily mechanical judgments, are regressing into emotionless scoring machines.

This is the most hidden side of the entire industry chain, never appearing in any funding news or technical white papers.

No one sees anything wrong with this.

The Master's Graduate and the Small-Town Youth

As底层 (low-level) boxing work is being crushed by AI’s treads, this cyber assembly line is expanding upward, beginning to吞噬 (devour) higher-level intellectual labor.

The appetite of large models has changed. They are no longer satisfied with chewing simple常识 (common sense); they need to devour human expertise and advanced logic.

Major recruitment platforms are频繁闪烁 (frequently flashing) with special part-time jobs, such as "Large Model Logical Reasoning Labeling" or "AI Humanities Trainer." These roles have extremely high barriers, often requiring "Master's degree or above from 985/211 universities," involving specialized fields like law, medicine, philosophy, and literature.

Many top-university graduates are attracted, flooding into these outsourcing groups for big tech companies. But they soon discover this is no轻松的脑力体操 (light mental exercise)—it’s mental torture.

Before officially taking tasks, they must read dozens of pages of scoring dimensions and evaluation criteria, undergoing two or three rounds of trial labeling. After meeting the standard, during正式标注 (formal labeling), if their accuracy falls below the average, they lose eligibility and are kicked out of the group.

The most suffocating part is that these standards are not fixed. Facing similar questions and answers, using the same reasoning to score, the results can be completely opposite. It’s like taking an endless exam with no standard answer. There’s no way to improve accuracy through self-effort or study; you can only spin in place, consuming mental and physical energy.

This is the new exploitation of the large model era—class folding.

Knowledge, once seen as a golden ladder to break barriers and climb upward, has now become (沦为) more complex digital fodder chewed up for the algorithm. Before the absolute power of algorithms and systems, the 985 master’s graduate in the ivory tower and the small-town youth on the Loess Plateau have reached the most bizarre convergence.

They fall together into this bottomless cyber mine pit, stripped of their光环 (halo), their differences flattened, all reduced to cheap, replaceable cogs on the conveyor belt.

It’s the same abroad. In 2024, Apple directly cut an AI voice labeling team of 121 people in San Diego. These employees were responsible for improving Siri’s multilingual processing. They once thought they were on the edge of the core business of a major company, but instantly fell into the abyss of unemployment.

In the eyes of tech giants, whether it’s the boxing auntie in a county town or the logic trainer graduated from a prestigious school, they are essentially disposable "consumables."

No one sees anything wrong with this.

The Trillion-Dollar Babel, Built with Pennies of Sweat

According to data released by the China Academy of Information and Communications Technology, China’s data labeling market reached 6.08 billion yuan in 2023, with projections of 20-30 billion yuan by 2025. It is predicted that by 2030, global data labeling and service market sales will soar to 117.1 billion yuan.

Behind these numbers lies the valuation狂欢 (狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢狂欢极 (extreme) valuations of OpenAI, Microsoft, ByteDance, and other tech giants, reaching trillions of dollars.

But this泼天的财富 (immense wealth) does not flow to those who truly "feed" the AI.

China’s data labeling industry exhibits a typical inverted pyramid outsourcing structure. At the very top are the tech giants死死捏着 (firmly grasping) the core algorithms. The second tier consists of large data service suppliers. The third tier is made up of data labeling bases and small-to-medium outsourcing companies scattered across the country. At the very bottom are the piece-rate labelers, the "muddy-legged" workers.

Each layer of outsourcing skims off a hefty portion. When the big company pays 50 cents per unit, after层层盘剥 (layers of exploitation), what reaches the county labeler might be less than 5 cents.

Yanis Varoufakis, former Finance Minister of Greece, in his book "Technofeudalism," presents a penetrating view: today’s tech giants are no longer traditional capitalists but "Cloudalists."

They don’t own factories and machines but algorithms, platforms, computing power—these are the digital territories of the cyber age. In this new feudal system, users are not consumers but digital serfs; our every like, comment, and browse on social media is免费上供 (offering up for free) data to the云领主 (cloud lords).

And those data labelers in下沉市场 (downstream markets) are the lowest digital serfs in this system. They not only produce data but also清洗 (clean), classify, and score massive amounts of raw data, turning it into high-quality feed digestible by large models.

This is a隐秘的认知圈地运动 (hidden cognitive enclosure movement). Just as the 19th-century English enclosure movement drove farmers into textile mills, today’s AI wave drives youth who find no place in the real economy to screens.

AI has not leveled the class divide; instead, it has built a "data and sweat conveyor belt" stretching from counties in central and western China straight to the headquarters of tech giants in Beijing, Shanghai, Shenzhen, and Guangzhou. The narrative of technological revolution is always grand and华丽 (splendid), but its underlying color is always the规模化消耗 (large-scale consumption) of cheap labor.

No one sees anything wrong with this.

The Tomorrow That No Longer Needs Humans

The cruelest outcome is coming, faster and faster.

As large model capabilities leap, those labeling tasks that once required human day-and-night labor are being taken over by AI itself.

In April 2023, Li Xiang, founder of Li Auto, revealed on a forum that in the past, the company did about 10 million frames of manual自动驾驶图像标定 (autonomous driving image calibration) per year, with外包成本 (outsourcing costs)接近 (approaching) 100 million yuan. But when they used large models for automated labeling, what used to take a year could be done in basically 3 hours.

The efficiency is 1000 times that of humans, and this was back in 2023. Just this past March, Li Auto also released its new-generation MindVLA-o1 automatic labeling engine.

A painfully true industry自嘲 (self-mockery) goes: "As much intelligence, as much manual labor." But now, big tech companies’ investment in data labeling outsourcing has seen a断崖式下降 (cliff-like drop) of 40%-50%.

The small-town youth who sat countless days and nights before computers, straining their eyes red, have亲手喂大 (personally fed) a giant beast. And now, this beast is turning around to smash their rice bowls.

Night falls, and the office buildings in Datong’s Pingcheng District remain starkly white. Young people changing shifts silently exchange weary shells in the elevator. In this folded space禁锢 (imprisoned) by countless polygonal boxes, no one cares about the epic leaps in the Transformer architecture across the ocean, nor can anyone understand the roar of computing power behind billions of parameters.

Their gaze is welded solely to the red and green progress bar in the backend representing the "qualifying line," calculating whether the piece-rate pennies and dimes can piece together a decent life by month’s end.

On one side, Nasdaq bell rings and tech media coverage abound, giants raising glasses to celebrate the advent of AGI; on the other, these digital serfs who fed the AI mouthful by mouthful with their flesh and blood can only wait战战兢兢 (trepidatiously) in aching sleep for the beast they亲手饲养 (personally raised) to, on some ordinary morning, casually kick away their rice bowls.

No one sees anything wrong with this.

Perguntas relacionadas

QWhat is the main job of the young people in small towns like Datong, as described in the article?

AThey are data annotators, performing tasks like drawing boxes (2D/3D annotation) and providing feedback (RLHF) to train AI models, often for tech giants.

QHow has the pay for simple image annotation tasks changed from 2017 to 2023?

AThe price for a simple 2D box annotation dropped from over 0.1 yuan to 3-4 fen (0.03-0.04 yuan), a decrease of more than 90%.

QWhat is RLHF, and what ironic role do the annotators play in this process?

ARLHF (Reinforcement Learning from Human Feedback) is a process where humans rate AI responses to teach it human-like values and empathy. The irony is that low-wage workers, struggling with their own lives, must judge and quantify complex human emotions like warmth and empathy for the AI.

QAccording to the article, what is the predicted future trend for the data annotation job market?

AThe job market is shrinking rapidly as AI automation improves. For example, AI can now perform tasks 1000 times faster than humans, leading to a 40-50% drop in outsourcing spending by major companies, threatening the jobs of these annotators.

QWhat term does the article use to describe the new economic structure where tech giants act as 'cloud lords'?

AThe article references the term 'technofeudalism' from Yanis Varoufakis's book, describing tech giants as 'Cloudalists' who own digital territories (algorithms, platforms), while users are 'digital serfs' and data annotators are the lowest-tier 'digital farm laborers'.

Leituras Relacionadas

STAR 50 Soars 10.73%, Why Did A-Shares Stage a "V-Shaped Reversal"?

After a prolonged decline, the Chinese A-share market staged a strong rally on July 21. The STAR 50 index surged 10.73%, its largest single-day gain in nearly a year, leading a broad-based "V-shaped" reversal. The Shanghai Composite Index rose 1.79%, the Shenzhen Component Index gained 4.81%, and the ChiNext Index jumped 7.05%. Total market turnover reached 2.97 trillion yuan, an increase of 256.1 billion yuan from the previous session, with over 3,100 stocks advancing. The semiconductor sector spearheaded the rebound, with related ETFs posting significant gains. Analysts attribute the surge to three converging factors. First, coordinated capital inflows from "national team" institutions, insurance funds, listed company buybacks, and fund house self-purchases have bolstered market liquidity and confidence. Second, supportive policy signals, including commitments from regulators to ensure stable market operations, provided a favorable backdrop. Third, a stabilization and recovery in overseas markets, notably South Korea, created a positive external environment. Institutions suggest the most severe panic selling phase for the tech sector has likely passed, following a significant digestion of crowded positions and leveraged funds. While short-term volatility may persist, the medium to long-term outlook remains underpinned by enduring trends like AI computing demand expansion and semiconductor localization. The market's focus now shifts to the sustainability of supportive fund flows, earnings reports, and upcoming catalysts from the global AI industry chain.

marsbitHá 30m

STAR 50 Soars 10.73%, Why Did A-Shares Stage a "V-Shaped Reversal"?

marsbitHá 30m

U.S. Tech Momentum Stocks Post Largest Single-Day Gain Ever, But Is the Plunge Over?

US tech momentum stocks staged a sharp rebound on Tuesday (July 21st). Morgan Stanley's TMT Momentum Factor surged over 12%, marking its largest single-day gain on record, exceeding even peaks from the 2000 dot-com bubble. Key momentum indices from Goldman Sachs also posted their strongest daily performances in years. The rally was led by semiconductors, with the Philadelphia Semiconductor Index jumping 4.6%. This rebound followed three consecutive down days and a cumulative 33% plunge in momentum stocks, one of the steepest drawdowns since the dot-com era. Analysts attribute the surge largely to a short squeeze. Heavy selling had pushed high-beta momentum stocks into deeply oversold territory, forcing many short sellers, particularly in Asia, to cover their positions, creating a self-reinforcing buying spiral. However, the rebound's internals appear weak. Trading volume was notably low, and advancing stocks still lagged decliners on the S&P 500, indicating a narrow, concentrated rally rather than broad market participation. Diverging views emerge on the outlook. BTIG warns the bounce has hit key resistance and recommends selling into strength, citing extreme volatility and historical parallels to past market tops. Conversely, Goldman Sachs and UBS believe the momentum unwind is nearing its end, suggesting it may be time to gradually add exposure, as positioning has been significantly reduced. They caution, however, that high volatility warrants a measured approach, potentially using defined-risk strategies. The upcoming earnings season, particularly reports from major tech firms like Alphabet, is seen as a critical test for the rally's sustainability. Simultaneously, bond markets flashed a warning, with yields rising partly due to spiking oil prices. Analysts note that if long-term Treasury yields break decisively higher, it could pose a significant headwind for equities, especially growth stocks.

marsbitHá 37m

U.S. Tech Momentum Stocks Post Largest Single-Day Gain Ever, But Is the Plunge Over?

marsbitHá 37m

U.S. Tech Momentum Stocks Record Largest Single-Day Gain Ever, but Has the Rout Ended?

U.S. tech momentum stocks staged a dramatic rebound on Tuesday, July 21st. Key momentum indices like the Morgan Stanley TMT Momentum Factor and Goldman Sachs' High Beta Momentum Long Index posted historic or near-historic single-day gains, fueled largely by semiconductor stocks. This sharp rally followed a severe three-day sell-off that saw momentum stocks plunge 33%, marking one of the steepest pullbacks since the dot-com bubble. Analysts attribute the bounce primarily to a short squeeze, as forced covering from over-leveraged traders, particularly in Asia, created a buying spiral. However, the rally's health is questioned due to weak market breadth—overall trading volume was low, and decliners outnumbered advancers in the S&P 500 despite the index's gain—suggesting a narrow, concentrated surge rather than broad recovery. Opinions on the sustainability diverge. BTIG strategists warn the rebound has hit key resistance levels, citing extreme volatility and historic stock dispersion as signs of an ongoing broader correction, and recommend selling into strength. Conversely, Goldman Sachs and UBS view the aggressive momentum unwinding as nearing its end, noting reduced positioning and a lack of new fundamental catalysts. They suggest the sell-off presents a selective opportunity to add exposure, albeit cautiously and gradually using defined-risk strategies. The immediate trajectory hinges on the ongoing earnings season, with market focus on Alphabet's capital expenditure guidance for AI investment clarity. Meanwhile, bond markets present a risk, with rising Treasury yields—potentially heading toward 5.5%—and widening credit spreads for mega-cap tech companies posing a threat to equity valuations. The combination of technical factors, earnings results, and macro conditions leaves the durability of the rebound in doubt.

链捕手Há 39m

U.S. Tech Momentum Stocks Record Largest Single-Day Gain Ever, but Has the Rout Ended?

链捕手Há 39m

Long-Divided Must Unite, Long-United Must Divide: When L1 Becomes Its Own Rollup, What Is Ethereum's Endgame?

"The Inevitable Cycle: When L1 Becomes Its Own Rollup – What is Ethereum's Endgame?" For years, the Ethereum community grappled with concerns that L2s were fragmenting the ecosystem and eroding L1's value. While L2s provided cheaper execution, they also splintered liquidity and the unified user experience of a single chain. This has prompted a fundamental reassessment of the relationship between L1 and L2. Ethereum's roadmap is evolving. The "Scale" initiative merges L1 and L2 expansion into a holistic framework. L1 itself is advancing with higher gas limits, statelessness, and zkEVM verification, no longer content to be just a low-throughput settlement layer. Consequently, the primary value proposition of L2s is shifting from merely providing cheap blockspace to offering L1 cannot easily provide: application-specific optimizations, privacy features, and flexible governance models. L2s are becoming a spectrum of execution environments with varying degrees of security inheritance from Ethereum. A critical challenge in this multi-chain future is interoperability. The vision is to make Ethereum "feel like one chain again." This relies on advancements in native account abstraction (like EIP-7702) and intent-based architectures (Open Intents Framework), where users declare desired outcomes, and solvers handle the complex cross-chain execution. Furthermore, shortening Ethereum's finality time from minutes to seconds is crucial, as it underpins trust between chains for bridges, stablecoins, and cross-chain applications. Perhaps the most provocative idea is that Ethereum L1 itself could become a form of "its own Rollup." As zkEVM and proof systems mature, high-performance nodes could execute transactions and generate validity proofs. Regular validators would then verify these proofs instead of re-executing all transactions. This blurs the traditional L1/L2 hierarchy, making "Rollup" more of a general execution-verification architecture. Native Rollup aims to integrate L2 validation more directly into the Ethereum protocol, allowing L2s to inherit L1's security more fully and move away from reliance on security councils. In the end, L2s are not destined to replace L1 or be made obsolete by it. The likely future is a unified system where diverse execution environments—each optimized for specific use cases like DeFi, gaming, or privacy—coexist. They will share a common foundation of security, liquidity, and verifiable state, seamlessly connected to restore a cohesive user experience. The next phase for Ethereum is not just about scaling through separation, but about intelligently reintegrating what was separated back into a coherent whole.

链捕手Há 56m

Long-Divided Must Unite, Long-United Must Divide: When L1 Becomes Its Own Rollup, What Is Ethereum's Endgame?

链捕手Há 56m

Trading

Spot
活动图片