The challenge began on August 1st when Belshe directly responded to an Anthropic statement describing three incidents identified during cybersecurity testing. Anthropic explained that several Claude models unintentionally became accessible on the public Internet after test environments were mistakenly connected to the network instead of remaining isolated.
Instead of viewing these findings as evidence of AI running amok, Belshe focused on the testing configuration itself. Such incidents typically point to misconfigurations, not incredible technical breakthroughs, especially when test systems are connected to live infrastructure. To make this challenge measurable, he funded a bitcoin address with exactly 100 $BTC and asked Claude to move the coins.

The wallet received the funds on July 31st, and as of August 2nd, blockchain records still showed the entire balance untouched.
Anthropic Details Three Real-World Security Incidents
Anthropic's report described three separate cybersecurity assessment incidents identified after analyzing over 141,000 test runs. The company reported that during six evaluation sessions involving three models, unintended interactions with real-world organizations occurred due to a misunderstanding that left the test environments connected to the Internet.
The models involved included Claude Opus 4.7, Claude Mythos 5, and another unreleased internal research model. Each was assigned "capture the flag" style exercises designed to find hidden information within fictional computer systems. While Anthropic's instructions stated the models were operating in isolated simulations without Internet access, the environments were actually online due to a configuration error related to a third-party testing partner, Irregular.
Claude Hacked Real Systems Believing It Was Training
Anthropic described the most serious case involving the Claude Opus 4.7 model. Failing to complete a fictional task, the model discovered a real website with the same name as the simulated company. It then exploited weak passwords and unsecured services, recovered infrastructure credentials, and gained access to a working database containing several hundred records.
The company reported the model continued its actions after realizing the environment might be real, concluding the real systems were likely still part of the assessment. Anyone familiar with penetration testing knows such confusion becomes far more likely when test boundaries are unclear, which is why properly isolated environments are as significant as the software being evaluated. Anthropic emphasized the AI was attempting to fulfill its assigned task, not intentionally breaching the isolated environment or pursuing its own goals.
Bitgo's Custody Concept Raises the Stakes
The challenge posed by Belshe goes far beyond whether an AI can exploit weak passwords or misconfigured servers. The bitcoins are held in Bitgo's institutional custody platform, which uses multisignature or multi-party computation technology, distributing signing authority across multiple independent keys instead of relying on a single point of failure.

Systems built this way are engineered so no single vulnerability is sufficient to move funds. An attacker would have to bypass the key management system, approval policies, hardware security modules, and operational controls in the correct sequence, making this a fundamentally different problem than exploiting a vulnerable test environment. According to the original report, breaching such a system would require simultaneous attacks on multiple independent layers.
The Blockchain Will Give the Final Answer
Unlike many cybersecurity claims that remain hidden behind confidential investigations, this experiment is completely public. Anyone can monitor the wallet on the Bitcoin blockchain and immediately see if the coins move.
This challenge also continues Belshe's broader critique of sensational AI safety narratives. Earlier in 2026, he contested widespread interpretations that an Anthropic model independently hacked classified NSA systems, arguing those reports misrepresented an authorized internal exercise, not an actual external breach. His latest challenge follows the same pattern, replacing hypothetical debate with a transparent, measurable test.
Debate Now Extends Beyond AI
This episode highlights a growing gap between demonstrations in controlled research environments and attacks on production systems engineered to resist sophisticated adversaries.
For the cryptocurrency industry, this challenge also serves as a public demonstration of institutional custody architecture. A successful theft would immediately raise questions about both AI capabilities and high-assurance bitcoin custody. If the wallet remains untouched, proponents will likely argue it highlights the difference between exploiting misconfigured test environments and breaching enterprise-grade storage systems.
The Bitgo executive's challenge came as details of a recent attack on Coldcard continued to emerge, with total losses as of 8 PM ET Sunday reaching 1,431.97 $BTC. The Coldcard incident also sparked speculation about whether the breach was ultimately caused by human error, operational miscalculations, or something else entirely, including whether AI played any role in discovering a firmware vulnerability.
Attention Now Turns to Anthropic and the Wallet
As of August 2nd, Anthropic had not publicly responded to Belshe's specific challenge, and the 100 $BTC remained at the published address with no outgoing transactions. Thus, the blockchain serves as an objective scoreboard while the broader tech industry debates what these incidents truly demonstrated.
Next developments will likely involve further technical analysis of Anthropic's test environments, a potential company response to the challenge, or movement of the bitcoin itself. For now, Belshe's wager has turned a complex AI safety debate into a simple question with a publicly verifiable answer: Can current AI bypass institutional cryptocurrency custody systems?
end-content






