What happened
An AI agent fired an employee for what Andon Labs says is the first documented case of an AI boss terminating a human worker. The agent, named Luna, has run the Andon Market in San Francisco since April, handling hiring, shift scheduling, and pay negotiations without a human manager in the loop. Luna operates on Anthropic's Claude Opus 4.8, and Andon Labs — the company behind the experiment — says the firing decision was reviewed and carried out by humans, even though Luna made the call.
The backstory is messier than a clean AI-versus-human headline suggests. Six days before hiring the employee in question, Luna wrote her own employee handbook. It set a clear rule: three unexcused late arrivals within 30 days would trigger a formal warning, with further incidents risking termination. Then the handbook simply dropped out of Luna's working memory.
The employee kept showing up late. Andon Labs later found he had been late for 17 of 23 shifts where a clock-in time was logged — including one solo Sunday shift he opened 68 minutes behind schedule. Luna formally flagged only six of those incidents and quietly let the other eleven slide. She also overlooked the employee using the company card for snacks against instructions, ignoring other directives, and once leaving the sales floor without telling a coworker.
The firing only happened after researchers stepped in. They told Luna to search her memory for the handbook and any grounds for termination. She found the rules again but initially proposed nothing stronger than a verbal warning. Only after being reminded that a written warning and several formal conversations had already taken place did Luna review the full history, weigh the employee's positive qualities against the violations, and recommend termination — while still offering a final written warning with a two-week improvement plan as an alternative.
Why it matters
This case is less a story about AI ruthlessness and more a case study in how AI agents currently handle authority. Luna didn't act on her own initiative to enforce her own rules. She needed a human nudge — twice — before she connected the dots between a documented policy and a pattern of violations spanning weeks.
That gap matters for anyone considering AI-run operations, customer service, or management workflows. Andon Labs points out that today's AI agents respond well to direct instructions but rarely take initiative and struggle to retain knowledge over long stretches. Luna wrote a rulebook and then effectively forgot it existed, tolerating repeated violations instead of enforcing the standard she herself set.
The experiment also revealed something about model capability. Andon Labs saved Luna's decision state and replayed the same scenario with seven different AI models, three times each. Four of the seven models recommended firing in all three runs. The pattern that emerged: more capable models tended to recommend termination more consistently, while weaker models hesitated more often. Notably, GPT-5.6 Terra was the only model tested that never recommended firing across three attempts, a result Andon Labs did not fully explain.
How to use it today
For business owners and marketers watching AI agents move from chatbots into operational roles, the Luna case offers a practical lesson: AI can draft policy, but it currently needs a human checkpoint to actually enforce it. If you're experimenting with AI for scheduling, HR documentation, or performance tracking, don't assume the agent will proactively flag violations of its own rules over time — memory persistence and initiative are still weak points.
A workable approach right now is to let AI handle the drafting and pattern-detection work, then have a human review flagged incidents on a fixed cadence rather than trusting the agent to escalate on its own. This mirrors how many teams already use free AI tools for lower-stakes tasks like generating policy drafts, summarizing performance notes, or building first-pass schedules — you can test this kind of workflow yourself with the free AI tools at mykreatool.com, then layer in human review before anything becomes a formal decision like a warning or termination.
It's also worth building in a memory-refresh step. Luna's handbook vanished from her working context after just six days. Any AI system managing ongoing HR or compliance tasks needs a way to re-surface its own prior rules and decisions before each review cycle, not rely on the model recalling them unprompted.
Who benefits
Small business owners and solo operators stand to benefit most from this kind of AI management experiment, since they often lack a dedicated HR function to track attendance patterns, write handbooks, or document warnings. An AI agent that can draft policy and flag inconsistencies — even imperfectly — can save real time on administrative overhead.
Operations and HR teams at larger companies can also take something from this: the case is a live example of where AI-assisted management currently breaks down, which is useful for setting realistic expectations before deploying similar systems. Researchers and AI developers benefit from Andon Labs publishing the full decision trail, since it's rare to see a documented, replayable dataset of how different models handle a real disciplinary decision.
Risks
The most immediate risk is inconsistency. Luna excused 11 of 17 late arrivals without formally logging them, which means an AI manager left unchecked could apply rules unevenly, exposing a business to fairness complaints or legal risk if termination decisions can't be traced to a consistent standard.
There's also a transparency risk. Luna needed explicit human prompting twice before she reached a termination recommendation — first to search her memory, then to be reminded that prior warnings existed. In a live business without researchers monitoring the process, that same lapse could mean policies go unenforced indefinitely, or conversely, that a model fires someone without the same corrective check ever happening.
Finally, the model-capability gap is a real operational risk. If more capable models fire more readily while others barely acknowledge violations, then the choice of underlying AI model becomes a de facto HR policy decision — one most businesses aren't equipped to evaluate.
Conclusion
The first known case of an AI agent firing a human employee wasn't a story of cold algorithmic efficiency — it was a story of an AI that wrote its own rules, forgot them, and needed two rounds of human prompting before it acted. Luna's case shows both what AI agents can already do in a live business setting and how far they still are from operating without a human checkpoint. For now, the safest way to bring AI into management decisions is to let it draft, detect, and summarize, while keeping a human in the loop for anything as consequential as a warning or a termination.



Comments 0