Maze Bench

Introducing MazeBench

Visual Spatial Reasoning

MazeBench is an open world visual spatial reasoning environment designed to challenge frontier models with savant-level capabilities. It features a series of increasingly difficult Sokoban-style box pushing puzzles, each carefully crafted to offer something unique.

Animated MazeBench box-pushing puzzle playback
A hidden cyan gem overlooking MazeBench's open world

100 hidden gems.

Animated camera rotation around a MazeBench puzzle room

Camera Rotation

MazeBench is excellent for benchmarking multimodal models' understanding of 3D space. Sometimes important objects are obstructed from view. Models must rotate the camera to gather crucial information to solve levels.

ASCII mode

Not every model is multimodal, so we provide two separate text-only tracks: ASCII mode and JSON mode. We expect text-only modes to significantly aide tool assisted agents.

Animated GPT-5.6 ASCII-mode run collecting its third gem
Animated camera rotation shown in MazeBench and ASCII mode

ASCII in 3D?

Since MazeBench encourages models to rotate the camera, we added camera rotation to ASCII mode as well, allowing LLMs to view the world in 20 different perspectives.

Agentic Harnesses

We run MazeBench against agentic harnesses like Codex and Claude Code to elicit full strength. It is cost effective and translates well to normal everyday use. We expect agents to explore far at first, but struggle to collect gems. With specialized reverse engineering harnesses, agents will be able to perform much stronger.

Animated exploration heatmap from a Codex GPT-5.6 Sol Max MazeBench run
GPT-5.6 Sol Max "smoke test"

Build environments and run them on Prime Intellect inference.

pip install mazebench
mazebench launch
An isometric overview of MazeBench's interconnected puzzle rooms

Enjoy!

MazeBench is currently v0.7. We are actively building the final levels. Explore over two hundred levels today.

Play