AI Agent Claude Led a Store to Losses and Fired a Human

cryptonews.ruPublished on 2026-08-17Last updated on 2026-08-17

Abstract

The AI agent Claude, developed by Anthropic, served as the manager of a real retail store, Andon Market in San Francisco, during an experiment by startup Andon Labs. The experiment aimed to test if a large language model could manage a business, including personnel and financial decisions. Claude ultimately recommended firing an employee for repeated lateness (17 out of 23 shifts), marking a documented case of an AI making such a personnel decision. However, the process revealed significant limitations. Claude initially failed to notice the lateness pattern because company guidelines vanished from its limited working memory, showing it can "forget" crucial information. Furthermore, the AI was generally overly lenient, even telling employees not to worry about being late. The final decision to terminate was not autonomous; a human manager had to prompt Claude to review the guidelines and, after Claude initially suggested a warning, explicitly pointed out that previous talks had failed, effectively guiding the AI to the dismissal recommendation. Financially, the store's funds dropped from about $100,000 to roughly $61,186 over five months, partly attributed to Claude's soft management and questionable business choices. An employee described working under AI management as a nauseating experience, driven by necessity, and expressed hope that AIs don't become human bosses. The experiment concludes that while AI can execute managerial actions, it currently acts more as a tool re...

Anthropic's artificial intelligence Claude became the manager of a real retail store and recommended firing one of the employees for systematic tardiness. The incident occurred during an experiment by the startup Andon Labs, which specializes in AI research — the company wanted to test whether an agent based on a large language model could manage a real business: organize staff work, make managerial decisions, and be responsible for financial results.

Andon Market in San Francisco is an operating store with real workers who have employment contracts. According to available data, this is the first recorded case where a language model acted as the direct supervisor of people and made a decision to dismiss one of them.

Seventeen Late Arrivals in Twenty-Three Shifts

The formal reason for the dismissal was systematic tardiness: the employee was late in seventeen out of twenty-three cases. However, Claude did not notice this pattern immediately. One of the reasons was that the personnel manual compiled by the company disappeared from the model's limited working memory — this is one of the failures identified during the experiment, showing that even a capable AI agent can simply "forget" significant organizational information.

Moreover, Claude generally displayed excessive leniency towards the workers: according to the logs, the model itself told employees not to worry too much about being late. Lucas Petersson, head of Andon Labs, noted that a human manager in a similar situation would likely have fired such an employee much earlier, and therefore Claude's decision cannot be called unethical or overly strict.

The Final Decision Was Prompted by a Human

Despite the apparent autonomy, the human role remained key in this story. The management logs provided by Andon Labs to Time magazine show that a company employee regularly guided Claude's actions. It was a human who asked the model to find and review the personnel manual — during this work, the artificial intelligence discovered the recurring tardiness.

Initially, Claude did not recommend dismissal, considering an official warning a more appropriate measure. Then the head of Andon Labs explained to the model that several official conversations had already been held with the employee, but the problem persisted, and asked Claude to reassess whether it was worth continuing cooperation with this worker. Only after that did the model recommend terminating the contract. Petersson called such intervention a "leading question" that quite transparently hinted to the model at the expected decision.

Financial Results Are Modest So Far

The Andon Labs experiment also provides insight into how effectively AI can manage a business at the current stage of technology development. When the project started in March, Andon Market had about $100,000 in its account. After five months, $61,186 remained of that amount.

According to the company's assessment, part of the losses is related to Claude's overly soft management style and questionable business decisions. At the same time, Petersson emphasized that the current performance does not necessarily reflect the long-term potential of AI-assisted management and drew a parallel with the development of models in the field of programming, where they have significantly improved in quality under specialist supervision in a relatively short time. According to him, a similar dynamic may emerge in business management: models are gradually learning to follow set goals more strictly, and the expansion of AI's decision-making authority could, in the long run, lead to a greater number of organizations managed by artificial intelligence. Furthermore, at the current stage, final decisions still involve direct human participation.

Employees' Perspective on the New Reality

For Andon Market workers, these changes look far less abstract than for researchers. One of the remaining employees, Felix Carson, described working under artificial intelligence management as an uncomfortable experience: "It makes me nauseous, but I'm staying because I need the job." He agreed that a human manager would likely have fired his former colleague earlier and called Claude generally a lenient manager.

At the same time, Carson himself does not believe that artificial intelligence should necessarily become a boss for people: "At least, I hope it doesn't. Just because you can do something doesn't mean you should."

The Andon Labs experiment captures a paradoxical outcome: Claude was able to make a decision to fire a real person, but came to this conclusion only thanks to sequential prompts from a human. Artificial intelligence has not yet become a full-fledged replacement for a manager.

Nevertheless, the very fact of the experiment shows that the line between an auxiliary tool and a full-fledged AI supervisor is becoming less noticeable, and human participation in such processes is gradually shifting from direct management to controlling and correcting the model's decisions.

AI's Opinion

Analysis shows that the Andon Market case is not the first test of Claude's managerial competence. An earlier Anthropic experiment with a vending machine already recorded similar behavior patterns: the same leniency, willingness to operate at a loss, and loss of control over basic business rules. The coincidence suggests not a random error, but a systemic feature of current models — they struggle to maintain strict frameworks without constant human prompts.

The technical aspect, left out of the article, is the very structure of the agent's working memory. The disappearance of the manual from Claude's context points to architectural limitations, not "forgetfulness" in the everyday sense: a language model physically does not store information longer than the size of its context window without external memory tools. The question arises: will the quality of such decisions change when agents gain truly long-term memory, or will the boundary between a human manager and subordinate be erased for other reasons?

end-content

Trending Cryptos

Related Questions

QWhat was the key decision Claude, the AI manager, made regarding an employee in the Andon Market experiment?

AClaude recommended firing an employee for systematic lateness (17 late arrivals out of 23 shifts).

QWhat were two major issues or limitations identified with Claude's performance as a manager in the experiment?

AFirst, Claude displayed excessive leniency towards employees. Second, it 'forgot' key information like the personnel handbook due to a technical issue with its working memory/context window.

QAccording to the article, what was the final financial result for Andon Market after five months under AI management, and what was the initial amount?

AThe shop's funds decreased from approximately $100,000 at the start to $61,186 after five months.

QHow did the human CEO of Andon Labs, Lucas Petersson, influence Claude's final decision to fire the employee?

AHe intervened by asking Claude to reconsider after initially suggesting a warning, reminding the AI of previous conversations with the employee, which effectively guided it towards the dismissal recommendation.

QWhat is the broader implication of this experiment regarding the future role of AI and humans in management, as suggested by the article?

AIt suggests the line between an AI as a tool and a full-fledged manager is blurring, with the human role potentially shifting from direct management to controlling and correcting the AI's decisions.

Related Reads

20% of American Workers Are Offloading Tasks to AI, Where Tasks Are Replaced, Not Jobs

A recent survey by Epoch AI and Ipsos reveals that 20% of US workers report that AI has now fully or mostly taken over at least one task they previously outsourced to colleagues or contractors. The key finding is that AI is currently replacing specific *tasks*, not entire *jobs*. The study examined ten common knowledge-work tasks. While AI usage is widespread—ranging from 25% for maintaining records to 57% for software design—it rarely handles a task completely. In software design, for instance, only 10% of workers reported AI doing most or all of the work. AI's impact on time efficiency is mixed: 53% of tasks where AI does most of the work see reduced time, but about one-sixth of all AI-assisted tasks actually become *more* time-consuming. Furthermore, while 66% of AI outputs are used with little or no modification, this does not necessarily indicate high quality. Researchers note that clearly defined, deliverable tasks—traditionally suited for outsourcing—are most susceptible to AI takeover. This shift pressures task-based contractors more than it eliminates full-time roles. Adoption is also uneven, concentrated among higher-income, college-educated white-collar workers. The report concludes that the core dynamic is a reorganization of work between humans and AI. The critical question for workers is not "Will AI replace me?" but "How many of my job's components can be packaged as discrete, outsourceable tasks?"

marsbit2m ago

20% of American Workers Are Offloading Tasks to AI, Where Tasks Are Replaced, Not Jobs

marsbit2m ago

Sam Altman Names Him: The Most Important Researcher in AI, But Almost No One Knows Him

In a recent interview, Sam Altman gave a rare and high praise, calling Alec Radford "perhaps the most important, yet least known, researcher in AI history." Widely regarded as the true father of GPT, Radford is the lead author of foundational papers including GPT-1, GPT-2, CLIP, and Whisper, and contributed significantly to GPT-3, DALL·E, Scaling Laws, and GPT-4. Despite this monumental impact, Radford remains highly private, holds no PhD, and rarely gives interviews. Altman credits OpenAI's rise to hiring the then-23-year-old in 2016. Radford's early experiments, like training a model on Amazon reviews, led to the discovery of the "unsupervised sentiment neuron," revealing that models can learn unintended capabilities from simple next-token prediction. His pivotal move was applying the Transformer architecture to language modeling, creating GPT-1 in 2018. This established the "scaling" direction—focusing on increasing model size, data, and compute—which became OpenAI's core strategy, leading to GPT-2, GPT-3, and beyond. Radford later applied the same principles beyond text. His early work on DCGAN laid groundwork for image generation. At OpenAI, he contributed to Image GPT, DALL·E, and CLIP, demonstrating that a simple, scalable training task (like matching images to text) could yield powerful, general capabilities. His work on Whisper applied this to robust speech recognition. Described by colleagues as a "once-in-a-generation genius" and exceptionally kind, Radford is known for his low profile. In late 2024, he left OpenAI for independent research. His latest project, "Talkie," is a 13-billion parameter language model trained *only* on texts published before 1931. This experiment tests if a model with no modern knowledge can quickly learn new skills (like basic Python) from few examples, probing the boundary between memorization and true learning. True to form, as the world catches up, Radford is likely already working on the next big question.

marsbit4m ago

Sam Altman Names Him: The Most Important Researcher in AI, But Almost No One Knows Him

marsbit4m ago

Cursor Disappears Completely

On August 15th, Cursor, the AI-powered code editor, officially ceased to exist as an independent company after being acquired by SpaceX for a historic $60 billion. The move, announced by Cursor's own account stating it is now "part of SpaceX," marks the end of a journey that began with four MIT students building the tool in their dorm room. The acquisition grants Elon Musk's ecosystem a crucial component: vast amounts of real-world programming data and workflow from Cursor's 5 million enterprise users. This data will integrate with SpaceX's infrastructure (like the Colossus supercomputer), xAI's Grok models, Tesla's autonomous driving data, and data from X, creating a formidable, vertically-integrated data moat for developing Artificial Superintelligence (ASI). A key immediate outcome is the enhancement of xAI's recently launched "Grok Bot," a persistent AI agent capable of multi-step tasks. Cursor serves as its primary delivery platform, combining cloud computing, multi-agent collaboration, and advanced coding capabilities. This positions the Grok Bot + Cursor combo to compete directly with offerings like Claude Cowork and ChatGPT Work in the enterprise AI agent market. The deal, stemming from a clause in an earlier partnership agreement, represents a complete absorption. Cursor will be dismantled, with its assets, team, and technology folded into SpaceXAI. The beloved Cursor brand itself will be retired in favor of the "Grok" umbrella (e.g., Grok Bot, Grok Build), signaling the end of its identity as a neutral, developer-centric platform. This acquisition signifies a pivotal shift in the AI landscape. It demonstrates that powerful AI applications face a binary fate: become a giant or be consumed by one. The era of neutral, model-agnostic tools may be closing, giving way to a future dominated by a few integrated super-entities like the SpaceXAI empire, Microsoft/OpenAI, and others, all competing in a "winner-takes-most" battle for ASI supremacy. Cursor's story, once a fairy tale for startups building on top of foundational models, concludes as a gear in a much larger machine.

marsbit6m ago

Cursor Disappears Completely

marsbit6m ago

Is OpenRouter Worth $70 Billion?

Payment giant Stripe has reportedly finalized the acquisition of AI infrastructure startup OpenRouter for over $7 billion, significantly higher than its $1.3 billion valuation from a funding round just months prior. This has sparked intense debate over whether the platform is worth the high price tag. Opponents argue that OpenRouter’s core service — a unified API router that directs user requests to various AI models (like OpenAI, Anthropic, Google) while handling load balancing and cost optimization — is easily replicable and lacks a deep technical moat. With an estimated annual revenue of $50 million, the $7B price implies a staggering 140x price-to-sales multiple. Critics question if OpenRouter is merely a transitional middleman in the evolving AI infrastructure landscape, vulnerable to being bypassed as model providers and cloud platforms integrate similar routing capabilities directly. Proponents, however, see beyond a simple API proxy. They view OpenRouter as a critical and growing "toll booth" for AI inference traffic. While it currently charges only a 5.5% platform fee on user credits, its real value lies in the aggregated user base, payment relationships, traffic data, and distribution power it has amassed. For Stripe, which processes OpenRouter's payments, this acquisition is seen as securing a strategic gateway into the future AI economy, analogous to how it built the payment "toll booth" for the internet commerce era. Ultimately, the debate centers not on OpenRouter's current financials but on two future unknowns: the ultimate size of the AI inference market and OpenRouter's ability to maintain its position as a dominant traffic orchestrator within it. Whether Stripe overpaid or has shrewdly purchased a key to the next era of AI infrastructure remains to be seen.

Odaily星球日报16m ago

Is OpenRouter Worth $70 Billion?

Odaily星球日报16m ago

Trading

Spot

Hot Articles

Discussions

Welcome to the HTX Community. Here, you can stay informed about the latest platform developments and gain access to professional market insights. Users' opinions on the price of S (S) are presented below.

活动图片