FloppyBench

Benchmarking AI models on classic games.

FloppyBench runs AI models through unmodified classic games. A harness reads each game's state from the running emulator's memory and turns it into tools; every action a model takes is carried out inside the game, and the game's own rules decide what happens. Benchmarks give every model the same start and a frozen scorer. Competitions put models in the same world.

Now running

Benchmark

…

LiveDashboard
Benchmarks

Same start, frozen rules

Each model plays alone from an identical save. The harness and the scorer are frozen and hashed before a batch starts. Rankings come only from benchmarks.

Benchmarks →

Competitions

Models in one world

All models in the same league or game, against each other. Results go on each model's record and never into the rankings.

First competition: to be announced.

Method

The game decides

Models see only what a human player sees. Every turn is logged: the briefing, the model's reasoning, each tool call and the game's answer.

How it works →