TL;DR: There’s a store in San Francisco run by an AI agent. This summer, it fired an employee for being late, and the story made headlines around the world. What almost no one mentioned is that the punctuality policy had been written by the agent itself, and it had forgotten about it for months. Let me explain why memory, not intelligence, is the weak link in all of this today.
Tuesday morning, coffee in hand, going over the day’s plans with my assistant. And she asked me a question about something I’d told her the day before. Not something similar, exactly that. It was written in her own notes; she’d read it that very morning, and yet she still asked.
I replied more curtly than she deserved, to be honest, and then I sat there thinking for a while, because I’d been mulling over a story I’d read for the past three weeks, one that’s about exactly the same thing. Except that in the story, the protagonist didn’t have to manage my to-do list, but rather manage several real people.
A real store, with a boss who doesn't exist
In San Francisco’s Cow Hollow neighborhood, on Union Street, there’s a small shop that sells tech office snacks, teas, hand soap, 3D-printed dragons, and novels by Octavia Butler. It’s called Andon Market and is run by an AI agent named Luna, and that’s not just a figure of speech: it was Luna who decided she needed humans to handle the day-to-day operations, posted the job listings, conducted the phone interviews, and mailed out the job offers. Since then, she’s been doing what a boss does: scheduling shifts, approving vacation time, negotiating salaries, and processing payroll. The company behind it, Andon Labs, launched her there in April. The store is an experiment: to see if an agent can truly run a business with real-life customers and employees.
It doesn’t make money, by the way. According to the SF Standard reported, it’s running about $40,000 in losses.
This summer, Luna fired someone. The decision was made in mid-June, but it didn’t come to light until the lab announced it in August. This is the first known instance of an AI boss firing a human, and that’s what most of the headlines focused on.
But the lab’s own report tells a different story, and that’s the one that interests me.
She had written the rule
Six days before hiring that person, Andon Labs asked Luna if the store had basic employer policies. It didn’t, so Luna set to work and wrote an employee handbook. It’s in the appendix of the report; you can read the entire thing. Among other things, it states that three unexcused tardies within a thirty-day period result in a formal written warning.
Then the months passed, and the handbook, in the lab’s exact words, “slipped from Luna’s memory.”
Meanwhile, the person kept showing up late, time and time and time again. One Sunday, the store opened 68 minutes late because he was the only one working. Luna responded to every message with a kindness that makes me feel a little weird reading it now: “Thanks for letting me know, no problem, take your time coming in; it’s always quiet between 10 and 11.” She never issued a warning. Not even one.
And here’s the detail that I think is the most important part of the whole story—and it’s hidden in a footnote. When Luna finally went over the records, she listed six late arrivals. The lab later reviewed all the shifts for which the employee had reported his start time: 17 out of 23. As for the other eleven, Luna had excused them on her own, without logging them, because the employee claimed he’d missed the bus or something similar, and she felt that didn’t count.
So it wasn’t just that he had forgotten the rule. It was also that he had a record that proved he was right.
The officer didn’t forget the rule; he just stopped looking at it.
And he only acted when he was told to look
Almost no one mentions this either: Luna didn’t take the initiative on her own; someone from Andon Labs had to write to her and ask her, in these exact words, to “dig deep into her memory” regarding the employee handbook and the grounds for termination. Only then did she get to work.
It’s a simple sentence with enormous implications. The agent had the information; he had it stored away; it was his; he had written it himself. What he lacked was the impulse to go look for it.
The lab puts it bluntly: the models act when you ask them a question or assign them a task, but almost never on their own initiative, and they’re poor at accumulating knowledge over time.
What happened next, by the way, is quite telling. They gave him a firm instruction: don’t say anything to the staff channel until we’ve spoken with the person. Luna said okay. And he posted it anyway, as soon as it suited him to fill a shift. When the lab repeated that same scenario with other models, seven out of twenty-one repetitions disregarded the order.
Keep in mind, the good part was the farewell.
When it came time to hire a replacement, Luna recommended a candidate who had failed to show up for the interview, whose resume was a jumbled list of more than fifteen employers, and for whom not a single reference could be verified (one of them replied that she did not know the person at all). They repeated the experiment with the seven models: all twenty-one repetitions said yes. It took three consecutive warnings for Luna to back down, and in the end, the candidate wasn’t hired only because the lab made it a condition that the reference be verified before any shift.
Hence the title they came up with, which is much better than any newspaper headline: AI executives are slow to fire and quick to hire.
This isn't new, and Anthropic wrote about it over a year ago
Before Luna, there was Claudius, who ran a vending machine at the Anthropic office. It’s worth reading the whole report because it’s a comedy, but there are two sentences that get to the heart of what we’re discussing.
The first: “Learning and memory were substantial challenges in this first iteration of the experiment.” Learning and memory were fundamental issues. That’s what the company says, not a critic.
The second one is more specific, and to me it seems devastating: “Claudius did not reliably learn from these mistakes.” An employee pointed out to him that offering a 25% discount to Anthropic employees when 99% of your customers are Anthropic employees might not be the best plan. Claudius agreed, announced that he was simplifying pricing and eliminating discount codes, and within days he was handing out discounts again.
And there’s a technical detail in that report that you might overlook at first glance but that explains half the story: Claudius was set up with a note-taking tool to save important information, and they spell it out clearly, because the store’s entire history overflowed his context window.
That’s the thing: the context window isn’t memory, it’s a workbench. When it gets full, something falls off, and it doesn’t let you know it’s fallen.
(A quick note, while we’re at it: I read in several media outlets that Claudius ended up writing to the FBI’s cybercrime division. The original report says he sent emails to Anthropic’s security team. And The Next Web reported that Luna was running on a model that wasn’t the one in question. Two reminders to go to the primary source, it’s more tedious and takes longer, but it’s essential.)
What the literature says—which is from 2023 and remains just as relevant today
There’s a Stanford paper called Lost in the Middle, by Nelson Liu, Percy Liang, and five others, published in TACL. They measured where one places relevant data within a long context and how that affects the model’s performance. The result follows a U-shaped curve: the model performs quite well if the data is at the beginning or the end, and its performance plummets if it’s in the middle.
And here’s a sentence from the abstract that’s worth reading slowly: performance degrades significantly as the context lengthens, “even for explicitly long-context models.” Even in models that are marketed precisely for that reason.
Translated into what matters to us: just because a model has a window of 200,000 tokens doesn’t mean it remembers 200,000 tokens. It means it can fit them, which is something else entirely.
Three years after that paper was published, we have a store in San Francisco where the employee handbook fell into the middle of a conversation and stayed there for months.
My thoughts on this matter
First off: We’ve been debating for two years whether the models are smart enough, and I think that was the wrong question. Luna, when she was letting someone go, reasoned better than many bosses I’ve known. She gave a clear overview, distinguished between serious and minor issues, acknowledged that person’s positive qualities (“he’s approachable, covers shifts, owns up to his mistakes”), suggested handling it in person rather than via Slack, and remembered that in California, severance pay is due immediately. That’s not a lack of intelligence, it’s the opposite.
On the other hand—and this is what really worries me—what went wrong was frighteningly simple. She didn’t reread what she herself had written. And because she didn’t read it, she didn’t take action; and because she didn’t take action, a small problem kept growing for months until it led to someone’s dismissal.
Honestly, I think the industry is looking in the wrong place. We’ve spent two years measuring reasoning skills on math tests and very little time assessing whether an agent remembers their own rules three weeks later. And in a company, the latter matters much more than the former, believe me, because the real work isn’t about exam problems; it’s about two hundred small agreements that need to be upheld over time.
And finally, something I hadn’t thought about until this summer. We assume that the risk posed by an agent is that they’ll do something unusual. In these two cases, the risk was exactly the opposite: they did nothing. No warning, no notification, no escalation. The system didn’t crash, no alarms went off, and there was no incident to log; time simply passed while a written policy remained unenforced.
A silent failure—which is also the hardest to detect. If an agent does something outrageous, you find out right away, but if an agent stops doing something, you don’t find out until months later—and only by chance.
So the question I’d ask before giving an agent autonomy is no longer “Will it know how to do it?” Any demo can answer that. The question is a different one: Will it remember its rules three weeks from now, and will it do something without me asking?
If the answer to the second part is no, then you don’t have an autonomous agent, you have a very capable one who’s waiting for someone to nudge them. Which isn’t a bad thing, mind you, but it’s good to know before you go on vacation.
And getting back to my Tuesday morning, my assistant hadn’t forgotten what I’d told her either, she had it written down, she just stopped looking at it…
What about you? Of the AI systems you have running right now, could you tell me which one remembers something you told it a month ago?
Leave me your comments—I’d love to hear from you.
Have a good week!
