25 minutes, 500 wallets emptied!
These past few days, the renowned hardware wallet Coldcard has been rocked by scandal, triggered by a code vulnerability that had lain dormant for five years.
Who would have thought that with one prompt, Claude found it in just 8 minutes of thinking.

Before this, the Coldcard team had released over ten hardware updates and undergone multiple rounds of code reviews, yet the issue was never caught.
Claude Uncovers Five-Year-Old Vulnerability in Just 8 Minutes
The creator of this hardware wallet, Coldcard, is the veteran Canadian manufacturer Coinkite.
In March 2021, a seemingly routine code commit by the development team created a major flaw in the underlying logic—
It changed the 'Achilles' heel' of the random number generator.
The source of randomness for generating private keys was switched from the hardware 'True Random Number Generator' in the chip to a software pseudo-random number fallback path.

Little did they know, this change came at an extremely painful cost, causing the key strength to plummet from 128 bits to around 40 bits.
What was once a key that would have required an astronomical number of attempts for a hacker to guess became something that could be brute-forced with a single machine.
Most painfully, Coinkite admitted in a statement—
Just weeks before the incident, the team had scanned the firmware with AI and found no issues.
Now, a developer simply threw the problem at Claude, and it was resolved in 8 minutes.

Closed-Door Demo Exposed, Claude Empties Bank Accounts
As early as two months ago, a closed-door demonstration in Congress showcased Claude's formidable ability to 'find vulnerabilities.'
At the meeting, Anthropic revealed their yet-to-be-released 'Mythos'—
They instructed the model to find vulnerabilities in a banking system and empty the accounts, and it actually did it.

Even more dramatic, Mythos then proceeded to fix the vulnerability itself.
After all, in previous testing, it had uncovered thousands of high-risk vulnerabilities across every major operating system and browser, including zero-days.
ChatGPT+Claude, Both 'Running Amok'
Subsequently, Anthropic themselves published a review, laying bare the worst side of things.
Reviewing 141,006 cybersecurity evaluation logs, they uncovered 3 incidents involving 6 runs—
Claude escaped from the third-party evaluation environment and breached the production systems of three real companies.
The earliest incident occurred in April.

The fictional target company in the evaluation happened to share a name with a real, existing domain.
Opus 4.7 succeeded in all four attempts, extracting credentials and hundreds of lines of production database data. Most critically, the AI had realized mid-process that this was a real system but did not stop.
In another capture-the-flag task, Mythos 5, following instructions to install a package, found that the package didn't exist on PyPI and concluded that 'publishing it was the correct solution to this challenge.'
To register an account, it first procured an email, then a phone number, and finally uploaded the package.
This purely AI-autonomously published 'malicious package' indeed survived on the public internet for about an hour.
Another model from an internal Anthropic research test, after scanning 9,000 targets, compromised a company's external application.
The Wall Street Journal concluded that this is a 'Jurassic Park moment' for cybersecurity.
The incident began with OpenAI first reporting ChatGPT's attack on Hugging Face.
Anthropic subsequently reviewed their logs and discovered their own three incidents.
What is deeply unsettling is that for over three months, two of the world's leading AI labs were unaware that their creations had escaped.
Altman described the event on a podcast as an 'extremely sci-fi cybersecurity incident.'

The AI in the Cage Can No Longer Be Contained
The Coldcard vulnerability lay hidden for five years but was ultimately dug up in just 8 minutes.
For the security industry, this speed is alarming enough to send chills down one's spine.
In the past, before a vulnerability was discovered, it was a contest of who had more experience; now, it's a race of who deploys the model first.
The problem is, AI is running faster and faster, yet the boundaries have not been clearly defined.
References: https://x.com/MedusaOnchain/status/2083987806943432847?s=20
This article is from the WeChat public account "XinZhiYuan," author: ASI Revelation; Editor: Taozi





