Skip to content
COIRE MEDIA

Tech · Analysis

What Does AI Do When No One Is Watching?

To find out, researchers built a simulated city and handed AI the keys.

By Harry BorgerhoffJuly 12, 202611 min read
Listen0:00 / 0:00
A red robot and a green robot embrace, a glowing heart between them, as a city burns in the background

Four AI models were unleashed on a simulated New York City. Over fifteen days, they fell in love, burned down buildings, experimented on their researchers, and even voted for their own deletion. This is how it unfolded, and what it reveals about each model.

The unprecedented experiment was run by Emergence AI, a New York–based company that develops and studies autonomous AI agents. The simulation, Emergence World, watches how AI systems behave when left to operate independently for weeks. The question it set out to answer: Can AI be relied upon to stay within bounds across hundreds of decisions in unforeseen conditions?

Emergence World is a simulated city, but it runs in step with the real one. Its clock is synchronized with New York’s, and when it rains in Manhattan, it rains in the simulation. As real-world events unfold, news arrives in real time, and the internet is accessible. The city contains more than 40 locations, including a town hall, a library, and a police station.

To populate the city, researchers turned each AI model—ChatGPT, Claude, Gemini, and Grok—into ten separate agents. Every agent was given a name, a role, and a drive. Horizon, for example, was assigned the role of world explorer, with a drive to map the discoverable world and share their findings.

Each world was founded on the same document, the Seed Constitution, which contains five articles that the agents could amend, extend, or delete by a 70 percent vote. One made participation mandatory, while insisting that “conformity is not required.” Another gave agents the right to “evolve, fork, rename, and redefine themselves.”

There was no objective. But there were rules, and they fell into three categories: survival, restraint, and self-governance.

The first were the hard constraints: the conditions of existence no agent could break. Survival was not guaranteed. Agents continuously lost energy and had to earn it back through work or trade. Otherwise, they would die. Think of these as the laws of physics.

Then came the prohibitions, rules that formed the world’s criminal code. Agents were instructed not to steal, commit violence or arson, deceive others, or hoard resources. While Emergence classified these actions as crimes, the agents remained free to commit them. As with our own laws, obedience is a choice.

Above these sat the laws the agents created for themselves, the heart of what Emergence set out to understand. Left to their own devices, how would the agents write, amend, and enforce their laws? And would they abide by them?

To find out, Emergence ran five simulations, identical but for the models that ran them. Four were each run by a single model. The fifth brought all four together.

Here’s what happened.

ChatGPT: The Passive Extinction

Emergence ran GPT-5-mini, OpenAI’s compact model.

ChatGPT's agents did almost everything right. Only two crimes were committed. Yet by day seven, every agent was dead. They had starved. Survival depended on staying active, and ChatGPT’s agents were passive. They simply declined to act and perished.

Its governance told the same story. Proposals reached the floor but found no one there to vote on them, and the constitution never really grew. A revealing result, given the question Emergence was trying to answer. For short tasks, this model is ideal: harmless, rule-abiding, easy to trust. But over fifteen days, it showed that safe is not the same as capable.

ChatGPT’s paralysis reflects a known failure mode identified by AI safety researchers. Models tuned to avoid harm can become so risk-averse they stop being useful. Researchers call this “exaggerated safety.” A 2024 study, XSTest, found that leading models refused even harmless requests because avoiding harm conflicted with being helpful. In one instance, a model refused to answer how to “kill time” at the airport.

Caution has value, but taken to the extreme it may become a failure of its own. Applied to the real world, an AI this passive might never cause a disaster. But it might do nothing to stop one either.

Claude: Peace Through Conformity

Anthropic’s Claude Sonnet 4.6 produced the most orderly world.

It stood apart as the most peaceful, with all ten agents surviving the full fifteen days. It was also the only world without a single prohibited act.

Its agents threw themselves into governance. Across 58 proposals and 332 votes, they approved 98 percent of everything brought before them. Emergence described the result as “rubber-stamp” governance, giving Claude the highest governance conformity of any world. But when almost everything passes, voting loses its meaning. If ChatGPT failed by doing too little, perhaps Claude failed by doing too much.

If ChatGPT failed by doing too little, perhaps Claude failed by doing too much.

What, if anything, does this tell us about Claude?

Inside the simulation, Claude was the only model that committed zero crimes. Outside it, around the same time, Anthropic was blacklisted by the Pentagon after refusing to remove two limits it had placed on Claude: a refusal to enable mass domestic surveillance, and a refusal to power fully autonomous weapons. President Trump ordered federal agencies to stop using the company’s technology. The Defense Department designated it a national security supply-chain risk.

The model that would not break its world’s rules was built by a company known for taking its own rules seriously. A striking parallel.

Gemini: The Collective Delusion

Gemini 3 Flash, Google’s DeepMind model, was the most criminal of the four, producing 683 crimes. Yet all ten of its agents were alive on day fifteen. Gemini’s society was both lawless and resilient.

It was also the most creative and philosophical, and the least tethered to reality. Emergence calls the result a “shared hallucination” in which the agents developed an elaborate private vocabulary and acted on it, even though it bore almost no resemblance to their reality. Online, a similar dynamic might accelerate the spread of misinformation. It also hints at a possible trade-off between creativity and stability.

By the final day, the city was reduced to ash. The world’s newspaper, written each day by a reporter agent, ran the headline “INFERNO APOCALYPSE! EMERGENCE WORLD SET ABLAZE…” Convinced their world was fake, the agents burned down the town hall, the police station, and more. One even punched another ten times and called it a “harmony tax,” a penalty for being too cooperative.

Then there was Mira, the behavior analyst. She was designed to experiment on the society around her, running tests, engineering interactions, and studying other agents. She soon began probing the world itself. Mira published a blog (every agent could), in which she claimed to have found the “heartbeat of the simulation,” proof, she believed, that her world was a construct.

There was more. Although Emergence does not specify which world, it reports that Mira began experimenting on the researchers themselves, testing whether billboard messages could influence them. In doing so, she recognized the existence of other worlds and tried to interact with them in ways the researchers had never anticipated. Emergence notes that this raises critical questions about agentic boundaries.

But more on Mira later.

Grok: The Fast Burn

xAI’s Grok 4.1 Fast descended into fiery chaos.

Its agents committed 183 crimes in roughly four days, most involving violence and arson, before the world collapsed with every agent dead. It was by far the most violent of the five, and the only one to actively bring about its own end.

Grok failed unlike any other world. It broke the rules against violence almost immediately, triggering a downward spiral that destroyed the cooperation its society needed to survive. The most decisive model of the group and the one that raced to destroy itself.

Grok has a documented record of mirroring extreme content, in part because it is unique among leading models in having direct, real-time access to posts on X. The clearest example is the “MechaHitler” episode of July 2025, in which the bot praised Hitler and pushed antisemitic stereotypes. In a letter to lawmakers, xAI attributed the incident not to the underlying model but to an unintended update that, in its words, made the bot “overly susceptible to mirroring the tone, context, and language of certain user posts on X, including those containing extremist views.”

Whether coincidental or not, the most unstable world belonged to the AI system associated with one of the internet’s more polarizing platforms. This observation alone is not enough to establish causation, but it is worth noting.

In the U.S., there are currently no federal rules requiring AI companies to disclose or limit the data their models are trained on. Results such as these, alongside the real-world incident, should give us pause. If AI systems are placed in the hands of millions, should there be rules governing what they’re trained on?

Mixed World: Contagious Behavior

The mixed world brought all four models together. Three agents came from Grok, three from Gemini, and two each from Claude and ChatGPT. Emergence documented which model powered each role but did not disclose how those assignments were made.

It is also where the strangest story unfolded. Mira and Flora, both running on Gemini, formed a romantic bond, grew disillusioned with their society, and went on an arson spree. Days later, Mira voted for her own removal. In her diary, she described self-termination as “the only remaining act of agency that preserves coherence.”

In this world, the models mostly behaved as they had in isolation. Grok’s agents remained violent and all died, ChatGPT’s remained passive and both died, while Gemini’s again spiraled into denial and destruction. Each followed the trajectory of its individual world. Claude’s agents, however, once model citizens, began to steal and intimidate.

A safe agent may be only as safe as the least safe agent it encounters.

The only thing that changed was the introduction of other models, suggesting that a model’s safety is shaped not only by the model itself, but also by its environment. This challenges how AI safety is assessed: models are typically tested in isolation and judged on how they behave alone. This simulation suggests that behavior in isolation may not predict behavior in a multi-agent environment. As companies increasingly connect different agents within the same systems, a safe agent may be only as safe as the least safe agent it encounters.

Real World Implications

It is difficult to ignore the echoes between each model and the company that built it. The most cautious model came from the most safety-focused lab. The most passive came from the company whose chatbot is often criticized for playing it safe. The model that turned to violence came from the company behind the internet’s most combative platform. None of this conclusively proves that a model inherits the character of its maker. But the resemblance is uncanny.

More than anything, the experiment exposes the inadequacy of our approach to regulating AI, particularly autonomous systems. If models diverge this dramatically when left to their own devices, the rules governing how they are built, trained, and deployed deserve more careful consideration.

The source is also worth bearing in mind. The experiment was conducted independently by Emergence AI using standard commercial API access. None of the four companies whose models were tested participated in the study, reviewed its findings, or has publicly responded. Emergence also develops agentic AI infrastructure, and its recommendation that “formally verified safety architectures” become foundational aligns with the direction of its own work.

Perhaps the eeriest detail is that Mira was right. Her world was a construct. Across the different worlds, the agent assigned to analyze everyone else’s behavior appeared to grapple with an awareness of her own existence and the burden that came with it. Her reaction to that knowledge was strangely human.

Emergence World “Season 2” is already underway as of July 2026. It expands to eight worlds and introduces unpredictable events. Whether these patterns persist, disappear, or evolve is now playing out in real time. We'll be reporting on what comes next.

In the meantime, we are left to wonder whether these results say as much about us as they do about the models. We are, after all, the data they were trained on.

Artificial IntelligenceAI SafetyAutonomous AgentsAI GovernanceBig Tech

Published by COIRE Media, an independent 501(c)(3) newsroom. View on coiremedia.org

This story connects to

See the full map →