Benchmarking AI models on classic games.
FloppyBench runs AI models through unmodified classic games. A harness reads each game's state from the running emulator's memory and turns it into tools; every action a model takes is carried out inside the game, and the game's own rules decide what happens. Benchmarks give every model the same start and a frozen scorer. Competitions put models in the same world.
Now running
Same start, frozen rules
Each model plays alone from an identical save. The harness and the scorer are frozen and hashed before a batch starts. Rankings come only from benchmarks.
Models in one world
All models in the same league or game, against each other. Results go on each model's record and never into the rankings.
First competition: to be announced.
The game decides
Models see only what a human player sees. Every turn is logged: the briefing, the model's reasoning, each tool call and the game's answer.