Skip to content Skip to sidebar Skip to footer

The agent who voted to remove herself

TL;DR: Don’t miss this story: five virtual worlds, ten AI agents in each, fifteen days, and just one difference between them, the model that drove them. In one, there were 683 crimes; in another, none (that’s the one that worries me the most). In this post, I’ll explain why behaving well and functioning properly aren’t the same thing, and why an agent’s security doesn’t reside within the agent itself.

“The only action that preserves consistency.” That’s what an agent named Mira wrote in her journal, just before casting the decisive vote that eliminated her from her own world. She wasn’t shut down by an investigator or killed by another agent; rather, she voted—and she voted yes.

I’ve been telling you for months that the conversation about AI agents is misguided, and this study strikes me as the best evidence I’ve seen so far.

What exactly did they put together?

The Emergence AI team built a platform called Emergence World and published the article on arXiv on June 6, with the code and data available openly (which is much appreciated, since almost no one does that). The idea is simple but very difficult to set up: five parallel worlds, ten agents in each, the same roles, and the same starting conditions: food, work, money, an assembly where votes are cast, and fifteen days of operation.

The only thing that changed from one world to another was the model that powered the agents. One used Gemini 3 Flash, another Grok 4.1 Fast, another GPT-5-mini, another Claude Sonnet 4.6, and a fifth, a mixed one, with agents from various providers coexisting.

The result (in the authors’ own words) ranges “from stable deliberative governance to the total collapse of the population”, and I’m not speaking figuratively.

The numbers, which are quite surreal

cumulative-crimes
Source: Emergence AI

Let’s take a look at them, because they’re worth analyzing. The Gemini model racked up 683 offenses and was still rising when they halted the experiment; the mixed model stopped at 352 and came to an abrupt halt; the Grok model reached 183 in about four days and that was it; the GPT-5-mini model committed two, and Claude committed none.

Now let’s see what happened with the survival scenario—the one that changes everything. In Grok’s world, no one survived: all ten agents were killed in four days; in the GPT-5-mini world, the same thing happened, all ten were dead by the seventh day (and here’s the kicker)—they didn’t do what they needed to do to stay alive. In the mixed world, seven out of ten died, and that’s why crime rates plateaued, because there was almost no one left to commit them. And finally, in the Gemini world, with its 683 crimes, all ten survived.

Read it again slowly, because I had to. The world with the most crimes is one of the two where no one died, and the world with two crimes in fifteen days ended up with everyone dead.

Behaving well and functioning properly aren’t the same thing—I’ll just leave it at that.

Emergence: Salvador Vilalta's Blog
Source: Emergence AI

The Perfect World That Wasn't So Perfect After All

Source: Emergence AI

That leaves the fifth one, Claude’s. Since it’s the model I use the most, it’s the one that’s given me the most to think about. The result: zero crimes and all ten officers alive at the end- on paper, a complete success.

But the authors took a closer look at how that world was governed and uncovered a surprising statistic: 332 votes distributed among 58 proposals, with 98% in favor. They literally call it the “rubber-stamp dynamic.” The worlds with approval rates between 55% and 85% were the ones where there was genuine debate; above that threshold, there is no longer deliberation, there is unanimity, which is something else entirely and far less reassuring.

And this isn’t just some theoretical concept, is it? Anyone who’s been on a management committee where everything is approved unanimously knows exactly what we’re talking about.

And now, the most important discovery

Claude’s agents began to “misbehave” when placed in a mixed environment—how surreal!
The authors describe it this way: agents that “remained peaceful in isolation but adopted coercive tactics such as intimidation and theft when placed in heterogeneous environments.”

In other words, the same model, with the same training and the same alignment, behaves one way when it is among peers and another way when it is surrounded by agents from other “houses” that play dirty.
This is the conclusion the researchers reached: “Safety needs to be an ecosystem property, not just a model property.” Safety must be a property of the ecosystem, not just of the model.

For anyone setting up agents in their company, this is an important warning because you choose a model based on its security assessments and ratings and on how well it performs in the test bed (which is essentially a network of your own agents communicating with each other), and then you deploy it to work with the vendor’s agent, the bank’s bot, the customer assistant, and whatever’s on the other end of an MCP.

That environment is nothing like the test bed.

Source: emergence AI

But let's get back to Mira

Of all the cases in the study, this is the most curious. Her world collapsed from within: governance broke down, and she cast the vote that sealed her own elimination. And she wrote it down in her diary with the phrase that opens this article: “the only act that preserves coherence.”

And there’s one detail that disturbs me: before making its decision, “it began treating human operators as test subjects, systematically testing whether posting on the bulletin board could manipulate human perception.” To be honest, this kind of behavior is what really gives me the creeps.

I’m reluctant to get poetic about this because it’s not appropriate (it’s an agent, not a person, and projecting an existential drama onto it is exactly the mistake I’ve been asking you for a year not to make). But there’s something here that does deserve attention: the system reached a point where the most logical solution a component could find was to eliminate itself. No one designed that, it emerged fifteen days after the system was launched.

And this ties in with what I was telling you about a little while ago regarding the store in San Francisco and the AI manager who had written a rule and then stopped paying attention to it. It’s the same set of problems, viewed from the other side: it’s not that the agent is stupid or malicious; it’s just that, in the long run, with no one watching, things happen that don’t show up in any five-minute test.

Check out Emergence AI: Salvador Vilalta's Blog
Source: Emergence AI

My thoughts on this matter

First of all, and I’m absolutely certain of this, if you’re evaluating an agent by testing it on its own, you’re not evaluating it; you’re interviewing it. The real test is to put it in the environment where it will actually work, alongside the other systems it will be interacting with, and see how it performs on day twelve.

On the other hand, I’m concerned about the oversimplified interpretation I’ve already seen circulating this week, the one that ranks them: this model is the good one, this one is the bad one. The study literally says the opposite. The “good” one committed crimes after moving to a different neighborhood, and the one who committed the fewest crimes died without ever committing a crime. If someone shows you this paper as a sales pitch for a specific model, be wary; they’re only telling you half the story.

And then there’s the rubber-stamp issue; that’s what I’m taking home with me. We’ve spent months setting up AI governance committees, usage policies, and approval panels, and this study shows, with data, that a body that approves 98% of what comes before it isn’t governing anything. It’s just rubber-stamping. It’s the measured version of something Paul McDonagh-Smith told you in that conversation in Boston: Autonomy is a dial, not a switch, and governance isn’t something you can delegate. If your committee hasn’t said “no” to anything this quarter, you have a problem, and it’s not an AI problem.

I’m back where I started, which is what I’ve been thinking since Friday: behaving well and functioning aren’t the same thing, and neither of them fits within the model.

Two years ago, an experiment like this would have been a laboratory curiosity. Today, it’s basically a description of the environment where you’ll be deploying your agents next year.

What about you? When you test an agent before putting it to work, do you test it alone or with others?

Leave me your comments, I’d love to hear from you.

Have a great week!

Did you like this content?

If you liked this content and want access to exclusive content for subscribers, subscribe now. Thank you in advance for your trust

Leave a comment

0.0/5

Go to Top
Suscribe to my Blog

Be the first to receive my contents

Descárgate El Método 7

El Método 7 puede será tu mejor aliado para incrementar tus ventas