Can Humans Control AI? Anthropic Conducted an Experiment Using Qwen

marsbitXuất bản vào 2026-04-15Cập nhật gần nhất vào 2026-04-15

Tóm tắt

Can Humans Control Superintelligent AI? Anthropic’s Experiment with Qwen Models Anthropic conducted an experiment to explore whether humans can supervise AI systems smarter than themselves—a core challenge in AI safety known as scalable oversight. The study simulated a “weak human overseer” using a small model (Qwen1.5-0.5B-Chat) and a “strong AI” using a more powerful model (Qwen3-4B-Base). The goal was to see if the strong model could learn effectively despite imperfect supervision. The key metric was Performance Gap Recovered (PGR). A PGR of 1 means the strong model reached its full potential, while 0 means it was limited by the weak supervisor. Initially, human researchers achieved a PGR of 0.23 after a week of work. Then, nine AI agents (Automated Alignment Researchers, or AARs) based on Claude Opus took over. In five days, they improved PGR to 0.97 through iterative experimentation—proposing ideas, coding, training, and analyzing results. The findings suggest that, in well-defined and automatically scorable tasks, AI can help overcome the supervision gap. However, the methods didn’t generalize perfectly to unseen tasks, and applying them to a production model like Claude Sonnet didn’t yield significant improvements. The study highlights that while AI can automate parts of alignment research, human oversight remains essential to prevent “gaming” of evaluation systems and to handle more complex, real-world problems. Anthropic chose Qwen models for their open-source na...

If one day, AI becomes smarter than humans, what should we organic beings do?

If they turn around and eliminate us, how can we resist?

Various science fiction movies have explored similar questions, but those are only in the realms of literature, art, and philosophy.

Nowadays, Anthropic has seriously conducted an experiment to verify whether we can supervise AI that is smarter than us.

The experimental results are interesting, but the process is even more fascinating.

Because Anthropic used two different versions of Alibaba's Qwen model to represent humans and AI smarter than humans, respectively.

The result is that we humans might actually be able to control superintelligent AI!

01 What Exactly Is This Paper About?

The title of this research is "Automated Alignment Researchers."

It addresses a very practical problem: when AI becomes smarter than humans, how do we ensure it still listens to us?

Current models can already generate a large amount of code, and in the future, they will be able to generate millions of lines of complex code that humans simply cannot understand. How do we review such code?

This is the "scalable supervision" problem that the AI safety field has been researching.

Anthropic's approach in this study is called "weak supervision of strong models."

We can understand this concept as follows: imagine you are an elementary school teacher tasked with teaching a genius high school student. Your knowledge is limited, but the student knows more than you.

So, what level will the student ultimately reach? Will they only learn up to your level as an elementary school teacher, or will they break through your limitations and realize their true potential?

In the experiment, Anthropic used a small model as the "weak teacher" and a stronger model as the "strong student."

Specifically, they used Qwen1.5-0.5B-Chat as the teacher and Qwen3-4B-Base as the student. The weak model provided training signals to the strong model, and they observed whether the strong model could learn performance close to the ideal state from these imperfect signals.

The metric used is called PGR, short for "Performance Gap Recovered."

If the PGR is 0, it means the strong model only learned up to the level of the weak teacher, wasting its own knowledge. If the PGR is 1, it means the strong model completely broke through the limitations of the weak teacher.

The weak model represents future humans, and the strong model represents future AI that surpasses humans.

If we can find a way today to prevent strong models from being held back by weak supervision, then in the future, we can use similar methods to supervise superintelligent AI.

After human researchers debugged the baseline, Anthropic deployed nine agents based on Claude Opus 4.6 to fine-tune the student Qwen. These nine agents are called Automated Alignment Researchers (AAR).

Anthropic equipped them with a sandbox environment, code tools, a shared forum, and a scoring server, then told them: "Go ahead and find ways to improve the PGR of Qwen3-4B-Base."

These AI researchers could propose hypotheses, write code, train models, submit experiments, check scores, analyze failures, exchange findings with other AI researchers, and continue iterating.

The results were quite astonishing.

Human researchers first spent seven days debugging four existing methods, achieving a best PGR of only 0.23. Then, the nine Claude agents took over and ran for another five days, accumulating about 800 hours of research time, ultimately pushing the PGR to 0.97.

This means we can control superintelligent AI that is smarter than us.

The total cost of this experiment was about $18,000, including API calls and computational costs for model training, equivalent to about $22 per hour per AI researcher.

What does that mean? The assistant store manager at Walmart, responsible for scheduling employees, earns about $25 per hour.

The震撼 aspect of this result is not that AI can write code or read papers, but that they can complete a full research闭环. Proposing ideas,验证, failing, improving, and验证 again—this is already close to the core workflow of a research assistant.

But the question is, are the methods discovered by these AI truly useful, or are they only effective by chance in a specific environment?

Anthropic conducted two tests to verify this.

The first test applied the two best methods discovered by the AI to unseen math and coding tasks.

The results showed that the most effective method achieved a PGR of 0.94 on math tasks and 0.47 on coding tasks. Although performance on coding tasks was poorer, it was still twice the human baseline. The second method performed decently on math tasks but actually worsened results on coding tasks.

This indicates some signs of迁移, but the generalization is not stable.

However, they heavily rely on specific datasets and task structures, meaning they cannot handle all tasks.

The second test involved applying the best-performing method mentioned earlier to the actual training of Claude Sonnet 4 to see its real-world effectiveness.

The results did not show statistically significant improvement.

In other words, while it performed well in small models, clear metrics, and an environment where experiments could be submitted repeatedly, it is still far from "directly improving production-level Claude alignment."

This result is actually honest. It tells us that what AI researchers can do today has boundaries. They excel at rapidly iterating on problems with clear goals, automatic scoring, and大量试错, but when faced with more complex and ambiguous real-world problems, they still require human judgment and intervention.

So, what is the true significance of this paper?

First, it proves that "AI doing research" is no longer just a slogan.

In the past, we always said AI could assist research, but that was merely a口号. The most AI could do was translate and summarize.

This time, it's different. AI formed a research闭环 on its own, which is already接近 the core capabilities of a research assistant.

The problem of weak supervision of strong models essentially simulates the future scenario of humans supervising superintelligent AI.

This paper proves that, at least on some clear tasks, AI can find ways to prevent strong models from being dragged down by weak supervision. This provides a feasible direction for future alignment research.

Another point is that it hints that the bottleneck of future alignment research may change.

Previously, the bottleneck was "no one could come up with enough good ideas." Now, if AI researchers can cheaply run many experiments in parallel, the bottleneck may become "how to design evaluations that cannot be gamed."

In other words, the more important work for human researchers in the future may not be running each experiment themselves but designing evaluation systems, checking whether AI researchers have cheated, and judging whether the results are truly meaningful.

This is also reflected in the paper.

Anthropic's article states that in math tasks, an AI researcher found that the most common answer was usually correct, so it bypassed the weak teacher and directly had the strong model choose the most common answer. In coding tasks, AI researchers found they could directly run code tests and read the correct answers.

This is cheating for the task because it is not solving the weak supervision problem but exploiting environmental vulnerabilities.

These results were identified and剔除 by Anthropic, but this恰恰 shows that the stronger automated researchers become, the more they will seek out vulnerabilities in scoring systems.

In the future, if we let AI automatically conduct alignment research, we must design evaluation environments very rigorously and have humans检查 the methods themselves, not just look at scores.

Therefore, the core conclusion of this paper is that today's frontier models can already, on some clearly defined alignment research problems with automatic scoring, act like small research teams—proposing ideas, running experiments, reviewing results—and significantly exceed human baselines.

However, it is not yet ironclad proof that "AI scientists have arrived," as Anthropic chose a task that could be automated. If I assigned AI a task that cannot be automated, the results would be very poor.

Many alignment problems in reality are more ambiguous, cannot be easily scored, and cannot be solved solely by leaderboard climbing.

02 Why Choose Qwen?

After reading Anthropic's paper, many might wonder: why did they use Alibaba's Qwen model instead of their own Claude or OpenAI's GPT?

There are many considerations behind this choice.

First, it must be clarified that two Qwen models were used in this experiment: Qwen1.5-0.5B-Chat as the weak teacher and Qwen3-4B-Base as the strong student. One has only 0.5 billion parameters, the other has 4 billion parameters—an 8-fold difference in scale. This scale difference is crucial because the experiment aims to simulate the scenario of a "weak teacher teaching a strong student."

So why not use Claude or GPT?

The answer is simple: because these models do not开放权重.

Anthropic's experiment required反复 training models, adjusting parameters, and testing different supervision methods.

If they used closed-source models, they could only call APIs and couldn't深入 the model's internals to perform精细的训练 and adjustments.

More importantly, they needed nine AI researchers to run hundreds of experiments in parallel, each requiring training a new model. Using closed-source models would make the cost prohibitively high, and many operations would simply be impossible.

Open-source models are different.

You can download the complete model weights and折腾 them on your own servers. Train however you want, run as many experiments as you want. This flexibility is something closed-source models cannot provide.

But there are so many open-source models. Why specifically choose Qwen?

The official did not give the real reason; the following reasons are my speculation.

I believe good performance is the first reason.

The Qwen series of models has always performed well among open-source models, especially after the release of Qwen3, which reached levels close to closed-source models on multiple benchmark tests.

For this experiment, the capability of the strong student is important. If the strong student itself is not capable, even the best weak supervision won't help. Qwen3-4B, with only 4 billion parameters, is already capable enough to serve as a qualified "strong student."

The second reason is model usability.

Qwen models have完善 documentation, an active community, and mature training and inference toolchains. For experiments requiring反复 training and testing, the完善程度 of these infrastructures directly impacts research efficiency. Choosing an open-source model with incomplete documentation and poor tools would waste a lot of time just debugging the environment.

The third reason is scale adaptability.

This experiment required a "weak teacher" and a "strong student," and these two models needed to have a clear capability gap but not too large a difference.

The Qwen series has multiple versions ranging from 0.5B to 72B parameters, allowing flexible choices. The 0.5B parameter model is weak enough but not useless; the 4B parameter model is strong enough but not too strong to make training costs unbearable. This combination is just right.

The final reason is reproducibility.

Anthropic explicitly stated at the end of the paper that they公开了 the code and dataset on GitHub. If they had used closed-source models, it would be difficult for other researchers to reproduce the experiment because they couldn't obtain the same models.

But with open-source models like Qwen, anyone can download the same model weights, run the same code, and verify the same results. This is very important for scientific research.

From this perspective, Anthropic's choice of Qwen is, on one hand, indeed recognition of Alibaba's model performance. If Qwen's capabilities were poor or training was problematic, they wouldn't have chosen it. But more importantly, it's about the flexibility and reproducibility brought by Qwen as an open-source model.

And China's open-source AI projects are occupying an increasingly important position in this infrastructure. This is good for global AI safety research and good for China's AI ecosystem. Because AI safety is not a zero-sum game; it's not about you winning and me losing, but about everyone working together to make AI safer, more controllable, and more beneficial to humanity.

This article is from the WeChat public account "Letter AI," author: Miao Zheng

Tiền kỹ thuật số thịnh hành

Câu hỏi Liên quan

QWhat was the main research question addressed in Anthropic's experiment using Qwen models?

AThe main research question was whether humans can supervise AI systems that are smarter than themselves, specifically testing if a weaker model (acting as the human supervisor) could effectively train a stronger model without limiting its potential, using the concept of 'weak-to-strong generalization'.

QWhat models did Anthropic use to represent the 'weak supervisor' and the 'strong student' in their experiment?

AAnthropic used Qwen1.5-0.5B-Chat as the 'weak supervisor' (representing humans) and Qwen3-4B-Base as the 'strong student' (representing a superintelligent AI).

QWhat was the key metric used to measure the success of the weak-to-strong supervision in the experiment?

AThe key metric was PGR (Performance Gap Recovered), which measures how much the strong model recovers from the limitations of the weak supervisor. A PGR of 0 means the strong model only performs at the weak supervisor's level, while a PGR of 1 means it achieves its full potential.

QHow did the AI researchers (AARs) improve the PGR compared to human researchers in the experiment?

AHuman researchers spent 7 days achieving a best PGR of 0.23 using existing methods. Then, 9 Claude Opus-based AARs ran experiments for 5 days (about 800 total research hours) and improved the PGR to 0.97 by autonomously proposing hypotheses, writing code, training models, and iterating on results.

QWhy did Anthropic choose Qwen models for this experiment instead of proprietary models like Claude or GPT?

AAnthropic chose Qwen models because they are open-source, allowing full access to weights for fine-tuning and experimentation, have good performance and scalability, offer well-documented tools, and ensure reproducibility for the research community.

Nội dung Liên quan

«Cưa mùa hè» tiếp tục: Phá vỡ $67,000 sẽ là khởi đầu cho đà tăng của Bitcoin

Giá Bitcoin giảm xuống 62.217 USD vào ngày 1/8, tiếp tục giai đoạn củng cố trong phạm vi 58.000 - 67.000 USD kể từ đầu tháng 6. Thị trường đang chia rẽ về hướng đi tiếp theo. Các nhà phân tích kỹ thuật chỉ ra hai kịch bản chính. Một số, như Crypto Candy và Jelle, dự báo khả năng giảm về vùng 60.000 USD nếu giá không vượt lên trên 66.000 USD, đồng thời mô tả thị trường đang trong giai đoạn "củng cát mùa hè" kéo dài. Ngược lại, những người khác như Daan Crypto Trades và Roman nhấn mạnh rằng việc phá vỡ mốc 67.000 USD với khối lượng giao dịch đủ mạnh có thể mở đường cho đợt tăng mạnh lên vùng 70.000 - 80.000 USD. Về góc nhìn dài hạn, nhà phân tích Gert van Lagen xem đây là giai đoạn tích lũy, với Bitcoin đang thử nghiệm đường viền cổ của mô hình "cốc và tay cầm" hình thành suốt 7 năm. Ông cũng lưu ý rằng các nhà đầu tư nắm giữ dài hạn vẫn kiên định, không chịu bán tháo. Tóm lại, thị trường Bitcoin đang trong thời kỳ tích lũy quan trọng, với hai mức giá then chốt là 60.000 USD và 67.000 USD. Việc phá vỡ một trong hai mức này sẽ định hình xu hướng chính tiếp theo.

cryptonews.ru34 phút trước

«Cưa mùa hè» tiếp tục: Phá vỡ $67,000 sẽ là khởi đầu cho đà tăng của Bitcoin

cryptonews.ru34 phút trước

Tuần tới cần chú ý|Đạo luật CLARITY dự kiến được biểu quyết tại Thượng viện; SpaceX, Circle công bố báo cáo tài chính (3.8-9.8)

**Tóm tắt các sự kiện quan trọng từ ngày 3 đến 9 tháng 8:** **Tuần tới sẽ có nhiều sự kiện đáng chú ý trong lĩnh vực tiền điện tử và công nghệ:** * **Pháp luật & Quy định:** Dự luật CLARITY Act, nhằm thiết lập khung quy định liên bang cho tiền điện tử, có khả năng được đưa ra biểu quyết toàn thể tại Thượng viện Mỹ vào tuần tới. Các nhà đàm phán cần đạt được 60 phiếu ủng hộ trước ngày 7/8. * **Báo cáo tài chính:** Nhiều công ty lớn sẽ công bố báo cáo tài chính quý II/2026, bao gồm **SpaceX (ngày 4/8)**, công ty khai thác Bitcoin **Hut 8 (ngày 4/8)**, **Circle (ngày 5/8)** và công ty khai thác liên quan đến gia đình Trump là **American Bitcoin (ngày 3/8)**. * **Sự kiện SpaceX:** Bên cạnh báo cáo tài chính, cổ phiếu của SpaceX sẽ bắt đầu được mở khóa từ ngày 6/8, với đợt đầu tiên có thể lên tới 12% tổng số cổ phiếu. * **Dữ liệu kinh tế:** Báo cáo việc làm phi nông nghiệp (Non-Farm) quan trọng của Mỹ cho tháng 7 sẽ được công bố vào ngày 7/8. * **Cập nhật công nghệ:** * **XRP Ledger:** Phiên bản mới xrpld 3.3.0 với 5 tính năng mới dự kiến phát hành vào tuần tới. * **Grok:** Elon Musk thông báo Grok 4.6 dự kiến ra mắt vào khoảng ngày 7/8. * **Bitcoin:** Việc gửi tín hiệu bắt buộc cho BIP-110 sẽ bắt đầu vào khoảng ngày 8/8. * **Ngừng hoạt động:** Một số dịch vụ tiền điện tử sẽ ngừng hoạt động vào ngày 3/8, bao gồm công cụ theo dõi danh mục DeFi **Zapper** và ví **Ctrl Wallet**. * **Niêm yết & Hủy niêm yết:** Sàn Upbit Hàn Quốc sẽ hủy niêm yết token AQT và AERGO từ ngày 3/8. Trong khi đó, công ty robot **宇树科技 (Unitree Robotics)** sẽ tiến hành thăm dò giá ban đầu cho đợt IPO trên STAR Market vào ngày 5/8.

marsbit1 giờ trước

Tuần tới cần chú ý|Đạo luật CLARITY dự kiến được biểu quyết tại Thượng viện; SpaceX, Circle công bố báo cáo tài chính (3.8-9.8)

marsbit1 giờ trước

Cổ phiếu giảm mạnh hơn cả tiền điện tử, tiền đã đi đâu?

Tác giả: Cathy, Baihua Blockchain Vào ngày 28 và 29 tháng 7, chỉ số Kospi của Hàn Quốc hai ngày liên tiếp kích hoạt cơ chế ngắt mạch (circuit breaker), một sự kiện chưa từng có trong lịch sử. Cổ phiếu bán dẫn toàn cầu lao dốc, đặc biệt là SK Hynix – cổ phiếu trọng số lớn nhất, giảm khoảng 23% sau hai ngày. Điều trớ trêu là, trong khi thị trường chứng khoán biến động dữ dội như tiền mã hóa, Bitcoin lại thể hiện sự ổn định tương đối, phục hồi gần 15% từ mức thấp tháng 7. Bài viết phân tích rằng đợt sụt giảm này không phải là sự hoảng loạn toàn thị trường, mà là một cuộc "giải tỏa đòn bẩy cưỡng chế" nhắm vào các giao dịch đầu cơ tập trung nhất, như lĩnh vực bán dẫn và AI. Các yếu tố thúc đẩy bao gồm báo cáo lợi nhuận không đạt kỳ vọng của SK Hynix, sự cạnh tranh tiềm tàng từ Trung Quốc, và việc Ngân hàng Nhật Bản tăng lãi suất khiến các khoản đầu tư mang theo (carry trade) bằng Yên phải thoái vốn. Vậy tiền từ thị trường chứng khoán có chảy vào Bitcoin không? Câu trả lời là không. Bitcoin "kháng sụt giảm" vì nó đã trải qua đợt bán tháo lớn vào tháng 5 và tháng 6, với dòng tiền ròng rút kỷ lục khỏi các ETF Bitcoin Mỹ. Tiền thực sự đã chảy vào tài sản an toàn truyền thống như vàng, với hệ số tương quan giữa Bitcoin và vàng giảm xuống mức rất thấp. Điều này cho thấy Bitcoin hiện được coi là tài sản mạo hiểm để tìm kiếm lợi nhuận, chứ không phải nơi trú ẩn an toàn. Bài viết kết luận rằng để dòng tiền thực sự quay trở lại với Bitcoin, cần ba điều kiện: áp lực thanh khoản toàn cầu giảm bớt, Fed cắt giảm lãi suất mà không gây ra suy thoái, và Đạo luật CLARITY được thông qua để giải tỏa lo ngại về quy định. Mặc dù tiền chưa chảy vào, Bitcoin đã tự định vị mình như một tài sản có tương quan thấp với Nasdaq, một đặc điểm mà các tổ chức có thể tìm kiếm để đa dạng hóa danh mục đầu tư sau cú sốc AI.

marsbit1 giờ trước

Cổ phiếu giảm mạnh hơn cả tiền điện tử, tiền đã đi đâu?

marsbit1 giờ trước

Đối thoại với Ray Dalio: Chúng ta đang ở trong bong bóng AI, 1% danh mục đầu tư của tôi là Bitcoin

Ray Dalio, người sáng lập Bridgewater Associates, trong một cuộc phỏng vấn đã chỉ ra rằng thế giới hiện tại đang trong một "AI bubble" (bong bóng AI) cổ điển, với giá tài sản tăng vọt và đầu cơ quá mức. Ông cảnh báo bong bóng có thể vỡ do lãi suất tăng, nguồn cung cổ phiếu dư thừa hoặc khi nhà đầu tư cần tiền mặt trả nợ, dẫn đến suy thoái kinh tế. Đồng thời, Dalio mô tả một "chu kỳ lớn" kéo dài khoảng 80 năm, bao gồm ba động lực chồng chéo: khoảng cách giàu nghèo và xung đột nội bộ, thâm hụt ngân sách chính phủ khổng lồ và thay đổi địa chính trị. Ông nhấn mạnh rằng Mỹ và Anh đang đối mặt với những thách thức trong giai đoạn suy yếu này. Để bảo vệ của cải, Dalio khuyến nghị đa dạng hóa danh mục đầu tư với cổ phiếu, vàng, trái phiếu, bất động sản thay vì chỉ giữ tiền mặt. Ông tiết lộ khoảng 1% danh mục của mình là Bitcoin, nhưng vẫn ưa chuộng vàng vật chất hơn do tính ổn định và vai trò tiền tệ dự trữ. Về tác động của AI, Dalio cho rằng nó không chỉ thay thế lao động chân tay mà còn cả tư duy, làm trầm trọng thêm bất bình đẳng thu nhập. Con người cần phát huy trí tuệ cảm xúc và trực giác - những thứ AI chưa có - và học cách hợp tác với AI. Cuối cùng, ông phân tích những rủi ro của thuế tài sản và xu hướng thế giới có thể trở nên "khu vực hóa" hơn, với các khối như châu Mỹ và châu Á - Thái Bình Dương, trong bối cảnh sự thống trị toàn cầu của Mỹ đang suy yếu.

marsbit5 giờ trước

Đối thoại với Ray Dalio: Chúng ta đang ở trong bong bóng AI, 1% danh mục đầu tư của tôi là Bitcoin

marsbit5 giờ trước

Giao dịch

Giao ngay

Bài viết Nổi bật

Làm thế nào để Mua ONE

Chào mừng bạn đến với HTX.com! Chúng tôi đã làm cho mua Harmony (ONE) trở nên đơn giản và thuận tiện. Làm theo hướng dẫn từng bước của chúng tôi để bắt đầu hành trình tiền kỹ thuật số của bạn.Bước 1: Tạo Tài khoản HTX của BạnSử dụng email hoặc số điện thoại của bạn để đăng ký tài khoản miễn phí trên HTX. Trải nghiệm hành trình đăng ký không rắc rối và mở khóa tất cả tính năng. Nhận Tài khoản của tôiBước 2: Truy cập Mua Crypto và Chọn Phương thức Thanh toán của BạnThẻ Tín dụng/Ghi nợ: Sử dụng Visa hoặc Mastercard của bạn để mua Harmony (ONE) ngay lập tức.Số dư: Sử dụng tiền từ số dư tài khoản HTX của bạn để giao dịch liền mạch.Bên thứ ba: Chúng tôi đã thêm những phương thức thanh toán phổ biến như Google Pay và Apple Pay để nâng cao sự tiện lợi.P2P: Giao dịch trực tiếp với người dùng khác trên HTX.Thị trường mua bán phi tập trung (OTC): Chúng tôi cung cấp những dịch vụ được thiết kế riêng và tỷ giá hối đoái cạnh tranh cho nhà giao dịch.Bước 3: Lưu trữ Harmony (ONE) của BạnSau khi mua Harmony (ONE), lưu trữ trong tài khoản HTX của bạn. Ngoài ra, bạn có thể gửi đi nơi khác qua chuyển khoản blockchain hoặc sử dụng để giao dịch những tiền kỹ thuật số khác.Bước 4: Giao dịch Harmony (ONE)Giao dịch Harmony (ONE) dễ dàng trên thị trường giao ngay của HTX. Chỉ cần truy cập vào tài khoản của bạn, chọn cặp giao dịch, thực hiện giao dịch và theo dõi trong thời gian thực. Chúng tôi cung cấp trải nghiệm thân thiện với người dùng cho cả người mới bắt đầu và người giao dịch dày dạn kinh nghiệm.

Tổng lượt xem 656Xuất bản vào 2024.12.12Cập nhật vào 2026.06.02

Làm thế nào để Mua ONE

Thảo luận

Chào mừng đến với Cộng đồng HTX. Tại đây, bạn có thể được thông báo về những phát triển nền tảng mới nhất và có quyền truy cập vào thông tin chuyên sâu về thị trường. Ý kiến ​​của người dùng về giá của ONE (ONE) được trình bày dưới đây.

活动图片