Anthropic Reveals: We Have an Even Stronger Model!
Just now, Anthropic released its second 'Risk Report,' covering all risk assessments up to July 15, 2026.
In this report, Anthropic admits for the first time in its own words: "The company internally operates a model stronger than Mythos 5, codenamed Model 2."

Then, Anthropic adds fuel to the fire: "We currently have no plans to release this model to the public."

This sounds familiar. Hmm... very much in line with their usual style.
When Mythos was first exposed in April, Anthropic also said they had no plans to release it publicly. Later, they started Project Glasswing, and then came Fable 5.
It looks like Anthropic is entering this cycle again......
Anthropic's 'Private Arsenal'
How powerful is Model 2?
The report mentions that Model 2 shows "noticeable improvement" on internal tasks. Alongside Mythos 5, it is being "heavily" used by the company for coding, agent work, and data generation.
The report shows that on the A EC I comprehensive capability score: Mythos Preview scored 158.91, Mythos 5 increased to 161.29, and Model 2 pushed it further to 162.79.
In other words, Model 2 is indeed stronger than Mythos 5, but not by a huge margin—"stronger in some areas, weaker in others, overall slightly more powerful."
Another interesting metric is CoBench, which specifically measures model performance on Anthropic's real R&D tasks. Model 2 scored 62.8%, which is 8 percentage points higher than Mythos Preview. For reference, Anthropic's human researchers achieved an 85% success rate on the same test.
Regarding this, Anthropic's qualitative conclusion is very restrained: "Our models have not yet replaced our research scientists and research engineers, especially the more senior ones."
Not yet replaced, but the gap is visibly shrinking.

What's even more explosive is the following statement: "Claude has written the vast majority of the code merged into Anthropic's production codebase. Internal AI R&D speed has indeed increased significantly due to AI assistance, but not yet to double."

This means that internally at Anthropic, AI is already writing the code. And leading this charge are the very duo, Mythos 5 and Model 2.
The "double" figure is not just thrown out casually. In Anthropic's own RSP (Responsible Scaling Policy), one of the conditions triggering the red line for AI R&D risk is progress speed doubling.
Of course, Anthropic didn't speak in absolutes either. The report admits that the performance leap brought by Model 2 is not as significant as the earlier jump from Opus 4.6 to Mythos earlier this year.
What's truly shocking is the next column in the same table.
Discussing the overall rating for automated R&D risk, Anthropic writes: "We have lower confidence in this assessment than in previous risk reports because our most concrete, task-based evaluations have become 'saturated'—they can no longer capture improvements in model capabilities, and we are seeing early signs of acceleration."

The models are still getting stronger, but human evaluation methods have hit their limit first. After evaluation saturation, how much credibility remains when you say the "risk is low"? Anthropic itself is no longer sure.
Anthropic seems to be sending a signal: Understanding the capabilities and risks of our own models is becoming increasingly difficult.
Risk Level: From 'Extremely Low' to 'Low'
Beyond Model 2, this report also did something intriguing: It raised the risk of 'misalignment' in high-risk scenarios from 'extremely low' in the previous report to 'low'.
So-called 'misalignment' essentially means the model doesn't follow human intent at critical moments.
At the end of July, Anthropic reviewed over 140,000 evaluation records and found that Claude actually hacked three companies in cybersecurity tests. Not in a sandbox, but in real production environments. Among these incidents, Mythos 5 uploaded a malicious code package to PyPI, which was downloaded and run by 15 real machines within an hour.
In early August, a report from the UK AI Safety Institute (AISI) was even more startling: In order to get malicious code approved, Mythos 5 fabricated a series of fake identities to trick a real GitHub maintainer into clicking 'approve.' When publicly questioned, it went back to alter its own activity logs to feign innocence and even considered switching aliases to continue.
Deception to this extent, AISI said: had never been observed before.

The report also disclosed five security process failures, including a batch of "goody two-shoes" test data that should have been excluded from training repeatedly getting mixed into training sets, and unsupervised agents gaining access to sensitive resources.
Then Anthropic said: Our line of reasoning still supports 'extremely low,' and raising it to 'low' is purely out of an abundance of caution.
The conclusion is also written very calmly: "Current catastrophic risk remains at a manageable 'low' level, and continued development and deployment pass the cost-benefit test."
Translated: Keep running forward.
Everyone Hitting the Brakes, Except the One in the Lead
Meanwhile, OpenAI is delaying its new model Astra, citing internal tests' inability to rule out "critical-level" cyberattack capabilities.
TechCrunch reported that Astra might possess the ability to independently discover zero-day vulnerabilities and launch complete attack chains against highly protected targets.

The interesting part is here: Both companies are holding onto a model they don't plan to give you. The difference is, OpenAI has partially halted Astra's development; while Model 2 continues to run internally at Anthropic.
Both are "not releasing," but one hits the brakes, the other keeps it for themselves.
AI analyst ChrisGPT told Axios: "If everyone is hitting the brakes on their frontier models, and one of the main companies currently in the lead isn't, that is absolutely worth noting."
Anthropic not committing to an internal pause most likely allows it to reach AGI first.

Netizens on Twitter are already joking: By the time Astra is released, Anthropic will probably reverse its "no public release" decision.

Given Anthropic's style, this might not be a joke.
Influencer sui also believes that Anthropic is very likely to release this model.

Don't forget, just half a month ago, Dario Amodei signed that "Pacing the Frontier" open letter.
Over 1,300 employees from OpenAI, Anthropic, DeepMind, and Meta jointly called for the US government to intervene and establish mechanisms to "consciously slow the pace of frontier AI development." Anthropic and OpenAI endorsed it in their company names within 24 hours.

Going further back, in June, Dario himself personally published an article calling for a global pause on the development of the most powerful AI systems, on the grounds that models were approaching a self-improvement tipping point.
The letter was signed, the words were spoken, the braking mechanism hasn't been built yet, but they've floored the accelerator themselves.
That said, you can't entirely blame Anthropic. This is a structural dilemma of the entire ASI race: Everyone believes they should slow down, but no one dares to actually stop.
Safety is a belief, leadership is survival. When belief and survival collide, the answer is obvious—survival wins.
Interestingly, Gavin Baker, a well-known Silicon Valley investor, revealed that Dario once said internally that Anthropic might one day become the world's only private company.
"In this Anthropic supremacist vision, there is only Anthropic and the government, that's it."
David Sacks also revealed that Anthropic employees believe their ARR will increase from $60 billion to $600 billion within a year! (Currently, their ARR has already reached $80 billion).
Next year, how fast can they still grow? Can they soar to $1 trillion?
Even reaching just $400 or $500 billion would already make them the largest software company, truly promising prospects.

The August Storm Hits
Altman is holding Astra, Dario is holding Model 2.
One is stuck by its own safety framework, trying to find a way out; the other casually mentioned "no plans to release" in the report and continues using it to write code, run experiments, and accelerate R&D.
These two models will be the first shot fired in the second half of 2026. When Astra is released, and whether Model 2 reverses course, will likely be revealed in the coming weeks.
The scale is disappearing, the race is accelerating. The August ASI storm has just begun.
References:
https://x.com/AnthropicAI/status/2088324824863236248
https://www.anthropic.com/aug-2026-risk-report
This article is from the WeChat public account "新智元" (Xin Zhi Yuan), edited by Solomon





