GPT-6 Astra vs Stockfish
A frontier language model played complete games of chess against a strength-limited engine at three limiter settings, with no chess engine, opening book, tablebase or tools of its own.
FEN protocol · Stockfish 18 · UCI_LimitStrength · 200,000 nodes · 1 thread · 16 MB hash
What happened
Astra won both games at the Stockfish 1320 and 1500 limiter settings, then lost both games at 1700. Six games, one per colour at each of three rungs.
| Limiter setting | As White | As Black | Astra ACPL |
|---|---|---|---|
| 1320 | Win | Win | 13.4 / 11.0 |
| 1500 | Win | Win | 35.0 / 30.4 |
| 1700 | Loss | Loss | 56.7 / 32.1 |
Move quality degraded as the limiter rose, and at the 1700 setting the engine posted the better average centipawn loss for the first time.
These are Stockfish limiter settings, not human Elo ratings, and none of this is a rating of the model. UCI_LimitStrength at 1500 does not play like a 1500 rated human, the fixed 200,000 node budget makes the configuration specific to this benchmark, and two games per rung cannot separate a strength boundary from the variance of two games. The result is the six game record above and nothing more.
Reliability
Every move was returned as a bare SAN token and validated against the position before it was played. An illegal move ends a game immediately as a loss, with no retry and no repair.
A separate preregistered reliability study sampled 40 positions across four conditions. It produced one illegal response in 160 and stopped under its own futility rule, which specified in advance that two or fewer failures would end the study without a comparison between conditions. That rule fired, and no condition is claimed to be more reliable than another.
The caveat that matters most. Astra received the authoritative FEN position every turn. This experiment therefore tests chess move selection given the current position, not its ability to remember the entire board unaided.
Why publish the losses
The 1700 pair are losses, and one earlier game under a stricter protocol was lost to an illegal move in a winning position. Both are published with the same records as the wins: every prompt sent, every raw response received, the engine configuration, and the post-game analysis. A benchmark that only reports the games it won is not measuring anything.