Skip to content
Chessglance
Open app
Menu
Blog — October 6, 2026

Reading 100-year-old chess books with a neural network

Hatched squares, outline pieces and yellowed paper. How we built a real test set from Wikisource and took exact-board accuracy on 1920s book diagrams from 24% to 84%.

5 min readbooksaccuracytraining

Update, 6 October 2026. We added a second held-out set: 737 diagrams from eight more books in English, Spanish, Italian and Czech (1846–1927). The same model reads 62.1% of them exactly (97.7% of squares): 95–98% on printing styles close to what it has seen, 6–18% on books with unusual piece glyphs, and none of the 1846 picture-style figurines. Across both sets it is 65% of 844 diagrams. The 84.1% below is real, but it is one set of three books; the larger number is the honest one, and the next training run targets it.

Screens are the easy case. A lichess board is a clean grid with flat colours and a known piece set. A 1920s book diagram is a printed halftone: dark squares filled with diagonal hatching, white pieces drawn as outlines with a paper-coloured halo, ink that bled, a page scanned a century later. The model we had before run 4 read 24.3% of them exactly. The current one reads 84.1%. This is how we measured that, and what changed.

A real test set, not a synthetic one

You cannot judge a book reader on boards you drew yourself. We needed real scans with trusted labels.

English Wikisource transcribes public-domain books page by page. Each page pairs a scan with wikitext, and diagrams are written with a {{Chess diagram}} template whose 64 cells are the position. Volunteers proofread the text against the scan. So every page is a labelled example. Our fetcher lists the pages that use the template, parses each diagram the way Wikisource renders it, downloads the scan, and skips diagrams whose markup is unreliable.

That gave 334 scoreable diagrams from Capablanca’s Chess Fundamentals and My Chess Career, Nimzowitsch’s My System, The Chess-Player’s Text Book, Rogers and a few smaller books.

The first run was humbling. The board finder located 99.4% of the diagrams, so detection was fine. But only 30.5% were exactly right, with 95.1% of squares correct. Black pawns came out as bishops. Hatched dark squares swallowed pieces.

Three books we never train on

We hold out three books completely: My Chess Career, Rogers’ How to Play Chess and My System. Their 107 diagrams are the books suite in our benchmark. Nothing from them goes into training. The release gate rejects any model that drops more than a point on this suite or on four others.

Model Exact boards Squares correct
Previous deployed model 24.3% 97.68%
Run 4 48.6% 98.70%
Run 5 (deployed) 84.1% 99.62%

Square accuracy barely moves; exact boards move a lot. That is the point of the metric. A board with one wrong square is a wrong position, and 98.7% of squares still leaves roughly one error per board.

Step one: make the fake books look like the real ones

Most of our training boards are rendered. For books the renderer draws ink coverage rather than colours, then applies print physics: ink spread, fading, bleed-through from the reverse page, skew, binarisation, low DPI. Run 4 added what the real scans taught us: dense, fine hatching at 45 to 65% coverage, like 1920s prints, where pieces on dark squares were being missed.

Run 4 also changed how we train: a moving average of the weights, no label smoothing (it degrades the uncertain-square flag), and checkpoints picked by exact-board accuracy on held-out boards. Exact boards on the books suite doubled, from 24.3% to 48.6%.

Step two: add real scans, carefully

Run 5 did two things. It trained a larger ConvNeXt-Tiny teacher and distilled two small networks from it, the ones that ship. And it put real book diagrams into the training mix: 196 crops from Chess Fundamentals, The Chess-Player’s Text Book and other small sources, kept only where our detector already matched at least 56 of 64 squares, then repeated 40 times so they carry weight next to tens of thousands of rendered boards.

The three test books stayed out. Capablanca reuses some positions across books, so five test positions also appear in the training books. Without them the score is 83.3%, which is the number to remember if you distrust the headline.

The result is the 84.1% above, with the invented-piece rate on empty squares at 0.13%.

The scanA 1920s printed chess diagram with hatched dark squares, a white bishop on b8 and a black king on e6.
What we read
B
p
k
p
p
p
p
P
b
P
P
P
K
Held-out, exact: a white bishop on b8 and a black bishop on g4, both read correctly. My Chess Career, page 129.

What the 18 misses look like

We went through every failure square by square and looked at several of the scans. Of 18 wrong boards in our rerun for this post, 15 are off by exactly one square, one by two, one by three, and one by seven, where several pieces landed a file over. The one-square errors are the ones you would predict: a queen read in the wrong colour, a bishop read as a pawn, a black pawn on a hatched square missed.

Our rerun is a simpler single-board pass and scores 89 of 107 (83.2%); the release gate runs the full pipeline and reports 84.1%. Same set, same models.

Two things we did not expect.

Some of the misses are the label’s fault. In the failures we checked by eye, the scan shows what we read and the transcription disagrees with the printed page: a rook on h1 that is not on the diagram, a knight on h7 where the page has a pawn. Proofreading catches most errors, not all.

The scanA printed chess diagram with a black pawn on h7.
What we read
r
k
r
p
b
n
p
p
p
p
p
b
P
N
B
P
P
P
P
P
R
B
R
K
Scored as a miss. We read a pawn on h7; the label says knight; the page shows a pawn. My Chess Career, page 159.

We have not audited all 18, so the headline stays 84.1%. But the real number is probably a little higher.

The flag does not catch everything. Only 4 of the 18 wrong boards had a square the model marked as uncertain. On old print, an error can arrive confident. We have measured this for books only. For book conversion we still tell publishers the same thing: output is a draft for a person to check, and the uncertain squares are where to look first, not the only place.

What is next

More real scans of more books and printers, diagram styles beyond the 1920s, and a better uncertainty estimate for print. Every book we add to the held-out list makes the 84% more honest. To try it on your own book, paste a page into the app; Pro converts a whole PDF to PGN.