AI Chess Benchmark

Can a language model actually play chess?

Complete games against a strength-limited engine, with no chess engine, opening book, tablebase or tools available to the model. Every game, every prompt and every raw response published.

Experiments

GPT-6 AstraFEN protocol · six games · 1320 to 1700

More models will be added here as they are tested, under the same protocol and the same publication rules.