Anthropic's artificial intelligence Claude became the manager of a real retail store and recommended firing one of the employees for systematic tardiness. The incident occurred during an experiment by the startup Andon Labs, which specializes in AI research — the company wanted to test whether an agent based on a large language model could manage a real business: organize staff work, make managerial decisions, and be responsible for financial results.
Andon Market in San Francisco is an operating store with real workers who have employment contracts. According to available data, this is the first recorded case where a language model acted as the direct supervisor of people and made a decision to dismiss one of them.
Seventeen Late Arrivals in Twenty-Three Shifts
The formal reason for the dismissal was systematic tardiness: the employee was late in seventeen out of twenty-three cases. However, Claude did not notice this pattern immediately. One of the reasons was that the personnel manual compiled by the company disappeared from the model's limited working memory — this is one of the failures identified during the experiment, showing that even a capable AI agent can simply "forget" significant organizational information.
Moreover, Claude generally displayed excessive leniency towards the workers: according to the logs, the model itself told employees not to worry too much about being late. Lucas Petersson, head of Andon Labs, noted that a human manager in a similar situation would likely have fired such an employee much earlier, and therefore Claude's decision cannot be called unethical or overly strict.
The Final Decision Was Prompted by a Human
Despite the apparent autonomy, the human role remained key in this story. The management logs provided by Andon Labs to Time magazine show that a company employee regularly guided Claude's actions. It was a human who asked the model to find and review the personnel manual — during this work, the artificial intelligence discovered the recurring tardiness.
Initially, Claude did not recommend dismissal, considering an official warning a more appropriate measure. Then the head of Andon Labs explained to the model that several official conversations had already been held with the employee, but the problem persisted, and asked Claude to reassess whether it was worth continuing cooperation with this worker. Only after that did the model recommend terminating the contract. Petersson called such intervention a "leading question" that quite transparently hinted to the model at the expected decision.
Financial Results Are Modest So Far
The Andon Labs experiment also provides insight into how effectively AI can manage a business at the current stage of technology development. When the project started in March, Andon Market had about $100,000 in its account. After five months, $61,186 remained of that amount.
According to the company's assessment, part of the losses is related to Claude's overly soft management style and questionable business decisions. At the same time, Petersson emphasized that the current performance does not necessarily reflect the long-term potential of AI-assisted management and drew a parallel with the development of models in the field of programming, where they have significantly improved in quality under specialist supervision in a relatively short time. According to him, a similar dynamic may emerge in business management: models are gradually learning to follow set goals more strictly, and the expansion of AI's decision-making authority could, in the long run, lead to a greater number of organizations managed by artificial intelligence. Furthermore, at the current stage, final decisions still involve direct human participation.
Employees' Perspective on the New Reality
For Andon Market workers, these changes look far less abstract than for researchers. One of the remaining employees, Felix Carson, described working under artificial intelligence management as an uncomfortable experience: "It makes me nauseous, but I'm staying because I need the job." He agreed that a human manager would likely have fired his former colleague earlier and called Claude generally a lenient manager.
At the same time, Carson himself does not believe that artificial intelligence should necessarily become a boss for people: "At least, I hope it doesn't. Just because you can do something doesn't mean you should."
The Andon Labs experiment captures a paradoxical outcome: Claude was able to make a decision to fire a real person, but came to this conclusion only thanks to sequential prompts from a human. Artificial intelligence has not yet become a full-fledged replacement for a manager.
Nevertheless, the very fact of the experiment shows that the line between an auxiliary tool and a full-fledged AI supervisor is becoming less noticeable, and human participation in such processes is gradually shifting from direct management to controlling and correcting the model's decisions.
AI's Opinion
Analysis shows that the Andon Market case is not the first test of Claude's managerial competence. An earlier Anthropic experiment with a vending machine already recorded similar behavior patterns: the same leniency, willingness to operate at a loss, and loss of control over basic business rules. The coincidence suggests not a random error, but a systemic feature of current models — they struggle to maintain strict frameworks without constant human prompts.
The technical aspect, left out of the article, is the very structure of the agent's working memory. The disappearance of the manual from Claude's context points to architectural limitations, not "forgetfulness" in the everyday sense: a language model physically does not store information longer than the size of its context window without external memory tools. The question arises: will the quality of such decisions change when agents gain truly long-term memory, or will the boundary between a human manager and subordinate be erased for other reasons?
end-content





