Skip to content
Devroop.More writing

Experiment · version 1

LayaChess: teaching a System 1 decision model to play chess

I gave a fast, intuitive AI model a chessboard to see how far it would get. This is version 1 of the experiment, and I'll be releasing improvements soon.

Play against Laya

· 12 min read · Devroop Saha

The LayaChess board in a browser: a game position next to panels showing Laya's win chance for its candidate moves and Stockfish's preferred move.
LayaChess running locally. Laya has just left the opening book and picked Nxe4 itself. On the right: its win chance for each candidate move, and Stockfish quietly preferring d6.

Two ways of thinking

Psychologists like to split thinking into two modes. System 1 is fast and intuitive: you glance at something and just know. System 2 is slow and deliberate: you sit down and work it out step by step.

Most of the AI everyone talks about right now leans towards System 2. Language models write out their reasoning one token at a time. But there's a quieter family of models built for the other mode, usually called decision models. You give one a situation and a question with some options, and in a single pass it gives you a probability for each option. No text, no chain of thought, just a judgement.

Two of them caught my attention:

I went with Laya, because I could train it (and because I didn't have a Jev API key).

Chess felt like the perfect test. Strong players use both systems: intuition shortlists a handful of moves almost instantly, and calculation checks them. So the experiment was simple to state. Make Laya the intuition, give it a search algorithm for the calculation, and see what kind of chess comes out.

See it play

Here's a short demo of the current version running on my laptop: a few moves from the opening book, then Laya taking over and playing on its own, with Stockfish's opinion shown alongside.

LayaChess, version 1, running locally.

How the engine works

Laya never sees the rules of chess, and it never has to. The engine around it handles the rules, and Laya only answers the one thing it's good at: how good does this option look?

  1. 1Position

    The current board, plus whose turn it is, castling rights and en passant.

  2. 2Rules: every legal move

    A chess library lists the legal moves. Laya never has to guess what's allowed, so it can't play an illegal move.

  3. 3Laya, one question per moveSystem 1 · intuition

    “White plays Nf3 (knight g1-f3). Win chance for white?” Ten answer levels, 0-10% up to 90-100%. One forward pass per question, no text generated.

  4. 4A win chance for every moveSystem 1 · intuition

    The answer distribution is averaged into a single number. Played alone, the engine would just pick the highest one.

  5. 5Search (Monte Carlo tree search)System 2 · calculation

    Laya's numbers decide which moves to explore first and how good each position looks. The search plays the most promising lines forward, sends every new position back to step 3, and keeps checkmates and draws exact from the rules.

  6. 6The move

    The move the search visited most. Its tree is reused on the next turn.

Opening book (optional)

Plays known opening moves first, then hands over to Laya once the game leaves the book.

Stockfish (comparison only)

Shown next to Laya on the board so you can see when they agree. It never picks Laya's moves.

How LayaChess turns a position into a move.

Each question looks roughly like this. The board goes in as plain text (piece lists turned out to work better than drawing an 8x8 grid), followed by the move being judged:

{"to_move": "black",
 "white": "Kg1 Qc5 Rd1 Rf1 Bc4 Pf2 Pg2 Ph2 Pe3 Pb4 Pa5",
 "black": "Kg8 Qe5 Rg6 Re8 Nd5 Pa6 Pc6 Ph6 Pb7 Pf7 Pg7",
 "castling": "-", "en_passant": "-"}

question: black plays Rg4 (rook g6-g4). Win chance for black?
options:  0-10% | 10-20% | 20-30% | ... | 90-100%

Laya spreads its confidence over the ten levels, I take the average, and that becomes the move's win chance. Doing this for every legal move gives the engine its intuition: a ranked shortlist.

The search is Monte Carlo tree search, the same family of algorithm that AlphaZero and Leela Chess Zero use. It takes Laya's shortlist, plays the most promising lines a few moves ahead, asks Laya about every new position it reaches, and settles on the move that holds up best. Checkmate, stalemate and draws by repetition always come from the rules, never from the model's opinion.

Around that sit the usual engine pieces: an optional opening book for the first few moves, the standard UCI protocol so it plugs into chess apps, and a small browser board for playing it. The board also shows what Stockfish would have played in each position. That's purely there for comparison; Stockfish never chooses Laya's moves.

How it differs from other chess engines

Every chess engine is some mix of two things: a way to judge a position (the evaluation) and a way to look ahead (the search). The big differences between engines come from how much they lean on each.

Classical engines: minimax and alpha-beta

The traditional approach, used from Deep Blue through to older versions of Stockfish, is a tree search called minimax: try every move, every reply, every reply to that, and assume both sides always pick their best option. The tree explodes quickly, so engines use alpha-beta pruning to skip branches that provably can't change the result, plus many tricks on top (move ordering, iterative deepening, caching positions they've already seen). At the bottom of the tree they score the position with a hand-written formula: material, king safety, pawn structure and so on. The evaluation is simple and very fast, and the strength comes from looking at enormous numbers of positions.

Modern Stockfish: alpha-beta plus NNUE

Today's Stockfish keeps the alpha-beta search but replaced the hand-written formula with NNUE, a small neural network designed to run extremely fast on a normal CPU. Only a few of its inputs change when a piece moves, so it can be updated incrementally instead of recomputed from scratch. The result is a learned evaluation that still lets the engine search millions of positions per second.

AlphaZero and Leela Chess Zero: a big network plus tree search

AlphaZero, and its open-source successor Leela Chess Zero (Lc0), flipped the balance. They use a large neural network that looks at the board and outputs two things at once: how promising each move is, and how good the position is. That network learned chess by playing millions of games against itself. Instead of alpha-beta, they use Monte Carlo tree search, which spends its time on the moves the network likes. They look at far fewer positions than Stockfish, thousands to tens of thousands per second on a GPU, but each one is judged much more cleverly.

Searchless models: no look-ahead at all

DeepMind's searchless chess models go to the extreme: a transformer trained on Stockfish's judgements that picks a move from the current position alone, with no search. It works surprisingly well when the model is big and the data is huge, but it can't double-check a tactic it doesn't immediately see.

Where LayaChess sits

LayaChess borrows from both of the last two. Like DeepMind's models, it judges each move by the win chance after playing it, trained on Stockfish's labels. Like Leela, it puts Monte Carlo tree search on top so it can look ahead. The difference is the brain itself:

So LayaChess isn't trying to compete with these engines. It's a test of a different question: what happens when you give a general-purpose decision model a chessboard and a bit of search?

Training it

I didn't have to label any positions myself. Google DeepMind released ChessBench with their “searchless chess” research: positions from real games where every legal move comes with Stockfish's win probability. The full set has about 15 billion of them.

For this first version I trained on a sample of about 2 million moves (2.05M to be exact), on Kaggle's free GPUs: two NVIDIA T4s, two sessions of roughly eleven hours each.

The thing that mattered most was keeping the questions short. My first format drew the board as a grid and used sixteen answer levels, which came to about 335 tokens per question and around 49,000 training examples an hour. Piece lists and ten levels brought it down to about 194 tokens, and training speed doubled to roughly 97,000 an hour. Same idea, half the words.

It wasn't all smooth. The two GPUs crashed the model on the first try, memory ran out on the second, and at one point I started the second training session before the first had finished uploading. Each run now saves a checkpoint every 45 minutes, which I strongly recommend to anyone training on free hardware.

What it learned

I measured progress on 300 positions from games the model never trained on, with two numbers: how often Laya's favourite move matches Stockfish's favourite, and how far its win chance is from Stockfish's on average.

Two line charts over 0 to 2.05 million training examples. Left: the share of positions where Laya picks Stockfish's best move rises from 6% to 27%. Right: the average win-chance error falls from 28.6 to 8.2 points.
Best-move accuracy went from 6% to 27% (a random legal move gets about 3%), and the win-chance error dropped from 28.6 to 8.2 points.

Most of the gain happens in the first hundred thousand examples, when the model learns to tell a good position from a bad one. Telling a good move from a slightly better one is much harder, and progress slows down from there. The best-move number is also noisy: with 300 test positions it moves by two or three points just by chance, which I learned the stressful way.

Is 27% any good?

On its own the number doesn't mean much, so here's the closest reference I know of. In DeepMind's searchless chess paper, measured the same way on the same kind of test positions, a small 9M-parameter transformer built for chess reached about 63%. It got there after roughly 5 billion training examples (5 million steps of 1,024), and their full-size models were trained on 128 TPUs each.

LayaChess saw about 2 million examples, roughly 2,500 times fewer, over one weekend with about 22 hours of training on free GPUs, and it reached 27%. Picking a random legal move gets about 3%. So it's nowhere near a purpose-built model trained at that scale, and it isn't meant to be. The point of this version was to check that a general-purpose decision model can learn chess at all, and the curve says it can and is still climbing.

The more satisfying test was simply looking at the moves, before and after training:

How strong is it?

Honestly, not very strong yet. In my early games against Stockfish at its weakest official setting (Elo 1320 on Stockfish's own scale), LayaChess has lost every game so far, although one of them lasted 88 moves. That setting sounds like a beginner, but it's a sneakily solid opponent that doesn't hang pieces the way beginners do.

My rough estimate is that this version plays somewhere around 1000 to 1200 against human players, on a chess.com-style scale. That's an estimate, not a measured rating. A proper rating from rated games is one of the next things on the list.

For context, DeepMind trained their searchless models on billions of examples. This one saw 2 million. It learned a surprising amount from that, and it clearly has a lot of room left.

Play it yourself

You can play against Laya on its Hugging Face Space. It's the same board as in the video, with an opening book, Laya's win chance for its candidate moves, and Stockfish's opinion alongside.

Because Laya judges one move at a time, every position means about 35 passes through a 421M-parameter model. Keeping a GPU server running around the clock for an experiment didn't make sense, so the Space runs on Hugging Face's free shared GPUs instead. Expect the first move to take a little while, and Laya's thinking time is capped at 10 seconds a move.

Prefer to run it on your own machine? The code and setup steps are on GitHub, and the model is on Hugging Face.

What's next

This is version 1, and I'll be releasing improvements soon. The main things I'm working on:

What I'd tell myself at the start

Check whether someone has already built the dataset before planning to build it yourself. Keep model inputs short, because half the tokens meant twice the training. Don't judge a run by one noisy number. And don't underestimate Stockfish at 1320.

Most of all: a general-purpose System 1 model really can pick up chess. It just needs a lot more practice than one experiment gave it.

Credits

Built on Laya by Convai Innovations, ChessBench by Google DeepMind, Stockfish and python-chess.