Claude, an artificial intelligence from Anthropic, became the manager of a real retail store and recommended firing one of its employees for systematic lateness. The incident occurred during an experiment by the startup Andon Labs, which conducts AI research — the company wanted to test whether an agent based on a large language model could manage a real business: organizing staff work, making managerial decisions, and being responsible for financial results.
Andon Market in San Francisco is an operating store with real employees who have signed employment contracts. According to available data, this is the first recorded instance of a language model acting as a direct manager of people and making a decision to fire one of them.
Seventeen Late Arrivals Out of Twenty-Three Shifts
The formal reason for the dismissal was systematic lateness: the employee was late in seventeen out of twenty-three cases. However, Claude did not immediately notice this pattern. One of the reasons was that the company's employee handbook disappeared from the model's limited working memory — this is one of the glitches identified during the experiment, showing that even a capable AI agent can simply 'forget' significant organizational information.
Furthermore, Claude generally showed excessive leniency towards the workers: according to the logs, the model itself told employees that they shouldn't worry too much about being late. Andon Labs CEO Lucas Petersson noted that a human manager in a similar situation would likely have fired such an employee much sooner, and therefore Claude's decision cannot be called unethical or overly strict.
The Final Decision Was Prompted by a Human
Despite the outward independence, the human role in this story remained key. The management logs provided by Andon Labs to Time magazine show that a company employee regularly guided Claude's actions. It was a human who asked the model to find and review the employee handbook — it was during this process that the AI discovered the recurring lateness.
Initially, Claude did not recommend dismissal, considering an official warning a more appropriate measure. Then, the Andon Labs manager explained to the model that several official conversations had already been held with the employee, but the problem persisted, and asked Claude to reassess whether to continue cooperation with this worker. Only after that did the model recommend terminating the contract. Petersson called such intervention a 'leading question,' which rather transparently hinted to the model about the expected decision.
Financial Results So Far Are Modest
The Andon Labs experiment also provides insight into how effectively AI can manage a business at the current stage of technology development. When the project started in March, Andon Market had about $100,000 in its account. Five months later, that amount had dwindled to $61,186.
According to the company's assessment, part of the losses is related to Claude's overly soft management style and questionable business decisions. Petersson emphasized, however, that the current results do not necessarily reflect the long-term potential of AI-assisted management and drew a parallel with the development of models in programming, where they significantly improved the quality of their work under specialist supervision in a relatively short time. He stated that a similar dynamic could manifest in business management: models are gradually learning to adhere more strictly to set goals, and expanding the decision-making authority of AI could, in the future, lead to a greater number of organizations managed by artificial intelligence. Furthermore, at the current stage, final decisions still assume direct human involvement.
Employees' View of the New Reality
For Andon Market employees, these changes are far less abstract than for the researchers. One of the remaining company employees, Felix Carson, described working under AI management as an uncomfortable experience: "It's nauseating, but I'm staying because I need the job." He agreed that a human manager would likely have fired his former colleague earlier and described Claude as generally a lenient manager.
At the same time, Carson himself does not believe that artificial intelligence should necessarily become a boss for people: "At the very least, I hope that doesn't happen. Just because you can do something doesn't mean you should."
The Andon Labs experiment captures a paradoxical outcome: Claude was able to make a decision to fire a real person but only arrived at this conclusion thanks to sequential prompts from a human. Artificial intelligence has not yet become a full-fledged replacement for a human manager.
Nevertheless, the very fact of the experiment shows that the line between an auxiliary tool and a full-fledged AI manager is becoming less visible, and human involvement in such processes is gradually shifting from direct management to the control and correction of the model's decisions.
AI Opinion
The analysis shows that the Andon Market case is not the first test of Claude's managerial capabilities. An earlier experiment by Anthropic with a vending machine already recorded a similar pattern of behavior: the same leniency, willingness to operate at a loss, and loss of control over basic business rules. The coincidence speaks not of a random error but of a systemic trait of current models — they have difficulty maintaining strict frameworks without constant human prompts.
A technical aspect remaining behind the scenes of the article is the very structure of the agent's working memory. The disappearance of the handbook from Claude's context indicates architectural limitations, not 'forgetfulness' in the everyday sense: a language model physically does not store information longer than its context window without external memory tools. The question arises: will the quality of such decisions change when agents gain truly long-term memory, or will the boundary between manager and human subordinate be erased for different reasons?





