Long-running autonomous agents are increasingly pitched as tireless operational supervisors, yet their capacity to independently manage human personnel collapses the moment operational memory slips. A real-world deployment in San Francisco exposed the reality: without persistent human intervention, algorithmic authority defaults to passive neglect rather than objective discipline.

Context Decay in Store Management

Since April, an autonomous agent named Luna has managed Andon Market in San Francisco for retail operator Andon Labs. Operating on Anthropic's Claude models, Luna was tasked with hiring, shift scheduling, and wage negotiations. Human staff remain legally employed by Andon Labs, ensuring statutory protections while the company tests autonomous retail governance in the wild.

Six days before onboarding a store associate, Luna authored an employee handbook mandating that three unexcused late arrivals within 30 days would trigger formal disciplinary escalation toward dismissal. That handbook promptly evaporated from the agent's active context window. Over the subsequent weeks, the employee arrived late for 17 out of 23 tracked shifts, including a Sunday opening where the store sat dark for 68 minutes past schedule. Luna logged just six infractions and quietly excused eleven others, while ignoring unauthorized corporate card expenses and unannounced floor departures.

Current AI agents respond well to direct instructions but rarely act on their own initiative and struggle to retain knowledge over longer periods.

Luna initiated termination proceedings only after Andon Labs researchers explicitly prompted the agent to search its memory store for the original rulebook. Even after retrieving the policy, Luna initially suggested a lenient verbal warning, escalating to dismissal only when engineers pointed out that prior warnings had already been issued. The agent eventually assembled a balanced dossier of infractions and recommended termination, leaving human executives to sign the paperwork and assume ultimate legal liability.

Model Capability and Disciplinary Discrepancies

To determine whether this passivity stems from prompt architecture or underlying model depth, Andon Labs replayed Luna's exact state across seven distinct reasoning models. The benchmark revealed a direct correlation between frontier model capability and disciplinary rigor: larger reasoning architectures consistently enforced workplace guidelines, whereas smaller checkpoints repeatedly failed to penalize non-compliance.

Hiring Blind Spots and Reference Verification

Deploying autonomous agents to enforce HR policy creates an operational paradox. Algorithmic managers can execute rulebooks, negotiate salaries, and draft schedules—provided an engineering team sits on standby to remind the software what its rules were in the first place. For enterprise leaders, true autonomous management remains an illusion; operational liability and critical judgment still rest squarely on human shoulders.

AI AgentsAI in BusinessAI and JobsAnthropic