GPT-6 Astra Clears Portal Solo, First Model to Finish the Game for $571
OpenAI’s Astra (GPT-6) cleared Valve’s puzzle game Portal this week, autonomously, in 23 hours and 43 minutes, at $571.18 in API token costs, in a run staged by independent tester cozyblaze.
The run was documented by Tom’s Hardware, which reported that the model had to map three-dimensional game spaces, interpret physics-based puzzle mechanics, and plan multi-step solutions using the game’s signature portal gun without any human guidance.
Also Read: Google’s Lyria 3.5 Becomes First AI Music Model on Gemini
That combination of spatial reasoning, tool use, and sustained autonomous planning across a full-length game is a rarer and more legible test of model capability than most benchmark tables offer. A fixed dollar cost attached to a fixed, verifiable outcome gives outside observers something leaderboards rarely provide, a real economic price tag for a real completed task.
OpenAI frames Astra as its most capable model yet for computer use, coding, and agentic tasks. The company’s own announcement highlights unrelated demos such as turning a Blender house model into a walkable Unreal Engine 5 scene, worth noting when evaluating how OpenAI chooses to present the model’s strengths.
Why A Video Game Run Doubles As A Benchmark
Portal is not a standard AI benchmark, but its puzzles demand exactly the skills labs struggle to measure cleanly, persistent 3D world modelling, causal reasoning about physics, and long-horizon planning without step-by-step human correction. For independent builders evaluating agentic pipelines, those are meaningful axes, more so than abstract scores on curated datasets.
A Familiar Genre With A New Ending
Outlets covering AI models playing games have historically documented failures, including chess engines losing to amateurs, which makes a completed Portal run notable regardless of who staged it.
Reddit communities in r/singularity and r/technology have framed this as the first time a model has beaten Portal outright, though that claim rests on one enthusiast’s experiment rather than any OpenAI-sanctioned benchmark, a distinction worth keeping in mind before treating it as a settled capability milestone.
Read Next: OpenAI Partner Status Officially Granted to Appinventiv
