Skip to content
Chessglance
Open app
Menu
Benchmark

Exact boards,
not close ones.

Ten tools, identical crops, 2,154 boards in five suites. A board scores only when all 64 squares are right.

PreliminaryUpdated October 6, 2026
a1 — Results

Exact boards by suite, in percent

Exact boards (%): all 64 squares correct
Toollichess227 boardslichess21,000 boardsrendered600 boardsbooks107 boardsnoise220 boards
Ours
100.0 (best)
100.0 (best)
99.3 (best)
84.1 (best)
92.3 (best)
tsoj/Chess_diagram_to_FEN
94.7
94.8
95.8
81.3
44.1
2d-chess-ocr
72.7
71.7
78.5
55.1
37.7
fenshot end-to-end
94.3
93.8
91.5
25.2
0.5
chessimg2pos
40.5
40.2
53.5
5.6
1.8
tensorflow_chessbot
29.5
31.0
50.8
0.0
0.0
chess-fen-detector
24.2
23.4
39.2
0.9
0.5
linrock
14.5
13.2
39.7
0.0
0.5
board-to-fen
20.3
20.6
14.0
0.0
0.5
kings-vision
7.0
7.3
25.2
0.0
0.5

Preliminary until a third-party test set and a second held-out book set are added. Tool versions are pinned in the published scripts. Hatched bars are other tools; solid bars are ours. fenshot is run end-to-end, finding the board itself.

b1 — Reading the table

Close on screens and books, far ahead on noise.

On clean lichess renders the best other tool, tsoj/Chess_diagram_to_FEN, reads 94.7% of boards exactly and ours 100.0%. On 1920s books ours reads 84.1% and tsoj/Chess_diagram_to_FEN 81.3%, a 2.8-point lead on 107 diagrams. That is close. Every book diagram tsoj/Chess_diagram_to_FEN reads exactly, we read exactly too, and we read three more. The wide gap is the noise suite: JPEG blocks, low resolution, textures and screen photos. Ours reads 92.3% and the next best tool 44.1%. That suite overlaps our training augmentations, so weigh it less than the real book scans.

We still lose boards. Old books are our weakest suite, and failures there are mostly one or two squares.

The lichess, rendered and noise suites use renderers and piece sets that overlap our training data. The books suite uses three held-out books that are never trained on.

On books, every diagram tsoj reads exactly we also read exactly: 87 both, 3 only us, 0 only tsoj, 17 neither. The 3-board gap on 107 diagrams is small.

A third-party test set and a second held-out book set are still to be added.

  • lichessReal lichess renders: boards drawn by lichess's own renderer in many themes and piece sets, both orientations.
  • lichess2A larger, independently drawn set of lichess renders.
  • renderedBoard libraries (chessiro-canvas, react-chessboard, chessground) with arrows, highlights and move dots.
  • booksHeld-out 1920s book scans from Wikisource: hatched squares, outline pieces, page noise.
  • noise11 corruption conditions x 20 boards: heavy noise, JPEG, low resolution, textures, screen photos.
c1 — Method

How we ran it

  1. Identical crops. Every tool gets the same image per test case: an exact crop of the board, or the board image itself for lichess.
  2. Each tool's best documented mode. Tools that find the board themselves were also run on crops padded with a white margin (10%, 25%, 50%); the best setting per tool is reported.
  3. True orientation given. Every tool's output is scored as seen, and the scorer turns the truth around for boards shown from Black's side.
  4. Exact boards only. A board counts when all 64 squares match the label; one wrong square is a miss. Test positions are legal placements.
  5. Nothing is fixed up afterwards. Competitor output gets format conversion only; a tool that crashes or finds no board scores an empty board.
  6. Scripts published. The harness, per-tool settings and predictions are published with the final results so anyone can rerun them.

Spot an unfair setting for your tool? Tell us at hello@chessglance.com and we will rerun it and publish the change. The 107 book diagrams come from public-domain books transcribed on Wikisource; the books in the test set are never used for training.