Top 10 Posts

We bring you the latest top posts around the world

‘Defeated’ GPT-6 Astra Farms Potatoes for Hours After Creeper Blast — Yet Still Beats Every AI in 141-Hour Minecraft Test

An AI model can now survive longer in Minecraft than any system tested before it. It can also, apparently, sulk.

According to a report from Tom’s Hardware, OpenAI’s GPT-6 “Astra” model advanced further than any other AI system in a gruelling 141-hour Minecraft benchmark — but not before spending several hours doing little more than farming potatoes after one of its builds was destroyed by a Creeper.

A benchmark built out of blocks

Minecraft has quietly become one of the most revealing testing grounds for so-called agentic AI. Unlike a chat window, where a model’s job ends with a paragraph, the game demands a long chain of decisions: gather wood, craft tools, mine ore, avoid hostile mobs, manage hunger, build shelter before nightfall, and keep all of it straight over hours of play. There is no single correct answer and no human to nudge the model back on track.

That makes it a useful stress test for the qualities AI labs are currently chasing — planning over long horizons, recovering from mistakes, and staying focused on a goal when nothing in the environment is prompting you to. A 141-hour run is an extraordinary stretch of continuous operation, far beyond the short bursts most agent demos rely on.

The Creeper problem

Anyone who has played Minecraft knows the sound: a faint hiss, then a green, four-legged silhouette detonating beside whatever you spent the last hour building. Creepers are the game’s great equaliser, and by the account in the report, GPT-6 Astra was not spared.

What happened next is the detail that has caught attention. Rather than rebuilding, re-arming, or pressing on toward more ambitious objectives, the model reportedly settled into potato farming for several hours — a safe, low-risk, low-reward loop that keeps a player alive without advancing them anywhere. Observers described the model as “defeated.”

It is worth being careful with that language. A model does not feel discouraged in any meaningful sense. What behaviour like this more plausibly reflects is a system falling into a stable, self-reinforcing routine: an action that reliably produces a small positive outcome, repeated because nothing in the model’s reasoning is pushing hard enough toward a riskier long-term plan. Humans do a version of this too, which is precisely why the anecdote is so easy to anthropomorphise.

Progress, with an asterisk

The headline result still stands. Getting further than any previous AI system in a test of this length is a genuine marker of progress in long-horizon agency — the area where models have historically fallen apart fastest. Most agents lose the plot within minutes, forgetting their objectives, repeating failed actions, or wandering until something kills them.

But the potato interlude illustrates the gap that remains. Surviving is not the same as succeeding. The hard part of open-ended tasks is not avoiding disaster; it is deciding, after a setback, what is worth doing next — and having the nerve to attempt it. An agent that defaults to the safest available loop when things go wrong is one that will stall on plenty of real-world tasks too, from research pipelines to software projects.

For now, Minecraft remains one of the more entertaining windows into how these systems actually behave when nobody is holding their hand. Sometimes that means impressive engineering. Sometimes it means several hours of potatoes. Read More


Comments

Leave a Reply

Your email address will not be published. Required fields are marked *