Two ways of thinking
Psychologists like to split thinking into two modes. System 1 is fast and intuitive: you glance at something and just know. System 2 is slow and deliberate: you sit down and work it out step by step.
Most of the AI everyone talks about right now leans towards System 2. Language models write out their reasoning one token at a time. But there's a quieter family of models built for the other mode, usually called decision models. You give one a situation and a question with some options, and in a single pass it gives you a probability for each option. No text, no chain of thought, just a judgement.
Two of them caught my attention:
- Jev by TypeSafe AI, a closed API (fittingly, its endpoint is called
systemone). - Laya by Convai Innovations, open source under Apache 2.0, built on a 421M-parameter ModernBERT-large encoder, and designed to be fine-tuned.
I went with Laya, because I could train it (and because I didn't have a Jev API key).
Chess felt like the perfect test. Strong players use both systems: intuition shortlists a handful of moves almost instantly, and calculation checks them. So the experiment was simple to state. Make Laya the intuition, give it a search algorithm for the calculation, and see what kind of chess comes out.
See it play
Here's a short demo of the current version running on my laptop: a few moves from the opening book, then Laya taking over and playing on its own, with Stockfish's opinion shown alongside.
How the engine works
Laya never sees the rules of chess, and it never has to. The engine around it handles the rules, and Laya only answers the one thing it's good at: how good does this option look?
- 1Position
The current board, plus whose turn it is, castling rights and en passant.
- 2Rules: every legal move
A chess library lists the legal moves. Laya never has to guess what's allowed, so it can't play an illegal move.
- 3Laya, one question per moveSystem 1 · intuition
“White plays Nf3 (knight g1-f3). Win chance for white?” Ten answer levels, 0-10% up to 90-100%. One forward pass per question, no text generated.
- 4A win chance for every moveSystem 1 · intuition
The answer distribution is averaged into a single number. Played alone, the engine would just pick the highest one.
- 5Search (Monte Carlo tree search)System 2 · calculation
Laya's numbers decide which moves to explore first and how good each position looks. The search plays the most promising lines forward, sends every new position back to step 3, and keeps checkmates and draws exact from the rules.
- 6The move
The move the search visited most. Its tree is reused on the next turn.
Opening book (optional)
Plays known opening moves first, then hands over to Laya once the game leaves the book.
Stockfish (comparison only)
Shown next to Laya on the board so you can see when they agree. It never picks Laya's moves.
Each question looks roughly like this. The board goes in as plain text (piece lists turned out to work better than drawing an 8x8 grid), followed by the move being judged:
{"to_move": "black",
"white": "Kg1 Qc5 Rd1 Rf1 Bc4 Pf2 Pg2 Ph2 Pe3 Pb4 Pa5",
"black": "Kg8 Qe5 Rg6 Re8 Nd5 Pa6 Pc6 Ph6 Pb7 Pf7 Pg7",
"castling": "-", "en_passant": "-"}
question: black plays Rg4 (rook g6-g4). Win chance for black?
options: 0-10% | 10-20% | 20-30% | ... | 90-100%Laya spreads its confidence over the ten levels, I take the average, and that becomes the move's win chance. Doing this for every legal move gives the engine its intuition: a ranked shortlist.
The search is Monte Carlo tree search, the same family of algorithm that AlphaZero and Leela Chess Zero use. It takes Laya's shortlist, plays the most promising lines a few moves ahead, asks Laya about every new position it reaches, and settles on the move that holds up best. Checkmate, stalemate and draws by repetition always come from the rules, never from the model's opinion.
Around that sit the usual engine pieces: an optional opening book for the first few moves, the standard UCI protocol so it plugs into chess apps, and a small browser board for playing it. The board also shows what Stockfish would have played in each position. That's purely there for comparison; Stockfish never chooses Laya's moves.
How it differs from other chess engines
Every chess engine is some mix of two things: a way to judge a position (the evaluation) and a way to look ahead (the search). The big differences between engines come from how much they lean on each.
Classical engines: minimax and alpha-beta
The traditional approach, used from Deep Blue through to older versions of Stockfish, is a tree search called minimax: try every move, every reply, every reply to that, and assume both sides always pick their best option. The tree explodes quickly, so engines use alpha-beta pruning to skip branches that provably can't change the result, plus many tricks on top (move ordering, iterative deepening, caching positions they've already seen). At the bottom of the tree they score the position with a hand-written formula: material, king safety, pawn structure and so on. The evaluation is simple and very fast, and the strength comes from looking at enormous numbers of positions.
Modern Stockfish: alpha-beta plus NNUE
Today's Stockfish keeps the alpha-beta search but replaced the hand-written formula with NNUE, a small neural network designed to run extremely fast on a normal CPU. Only a few of its inputs change when a piece moves, so it can be updated incrementally instead of recomputed from scratch. The result is a learned evaluation that still lets the engine search millions of positions per second.
AlphaZero and Leela Chess Zero: a big network plus tree search
AlphaZero, and its open-source successor Leela Chess Zero (Lc0), flipped the balance. They use a large neural network that looks at the board and outputs two things at once: how promising each move is, and how good the position is. That network learned chess by playing millions of games against itself. Instead of alpha-beta, they use Monte Carlo tree search, which spends its time on the moves the network likes. They look at far fewer positions than Stockfish, thousands to tens of thousands per second on a GPU, but each one is judged much more cleverly.
Searchless models: no look-ahead at all
DeepMind's searchless chess models go to the extreme: a transformer trained on Stockfish's judgements that picks a move from the current position alone, with no search. It works surprisingly well when the model is big and the data is huge, but it can't double-check a tactic it doesn't immediately see.
Where LayaChess sits
LayaChess borrows from both of the last two. Like DeepMind's models, it judges each move by the win chance after playing it, trained on Stockfish's labels. Like Leela, it puts Monte Carlo tree search on top so it can look ahead. The difference is the brain itself:
- It wasn't built for chess. Leela's and DeepMind's networks are designed around the board. Laya is a general System 1 decision model, and the board reaches it as a text question it answers in its usual format.
- It judges one move at a time. Leela scores every move in one pass. Laya needs a separate pass for each legal move, around 35 per position.
- So it's slow. About 2.5 positions per second on a T4 GPU, compared with thousands for Leela and millions for Stockfish. Its search can only afford to look a little way ahead.
- And it learned from very little. No self-play, no billions of games: about 2 million labelled moves.
So LayaChess isn't trying to compete with these engines. It's a test of a different question: what happens when you give a general-purpose decision model a chessboard and a bit of search?
Training it
I didn't have to label any positions myself. Google DeepMind released ChessBench with their “searchless chess” research: positions from real games where every legal move comes with Stockfish's win probability. The full set has about 15 billion of them.
For this first version I trained on a sample of about 2 million moves (2.05M to be exact), on Kaggle's free GPUs: two NVIDIA T4s, two sessions of roughly eleven hours each.
The thing that mattered most was keeping the questions short. My first format drew the board as a grid and used sixteen answer levels, which came to about 335 tokens per question and around 49,000 training examples an hour. Piece lists and ten levels brought it down to about 194 tokens, and training speed doubled to roughly 97,000 an hour. Same idea, half the words.
It wasn't all smooth. The two GPUs crashed the model on the first try, memory ran out on the second, and at one point I started the second training session before the first had finished uploading. Each run now saves a checkpoint every 45 minutes, which I strongly recommend to anyone training on free hardware.
What it learned
I measured progress on 300 positions from games the model never trained on, with two numbers: how often Laya's favourite move matches Stockfish's favourite, and how far its win chance is from Stockfish's on average.

Most of the gain happens in the first hundred thousand examples, when the model learns to tell a good position from a bad one. Telling a good move from a slightly better one is much harder, and progress slows down from there. The best-move number is also noisy: with 300 test positions it moves by two or three points just by chance, which I learned the stressful way.
Is 27% any good?
On its own the number doesn't mean much, so here's the closest reference I know of. In DeepMind's searchless chess paper, measured the same way on the same kind of test positions, a small 9M-parameter transformer built for chess reached about 63%. It got there after roughly 5 billion training examples (5 million steps of 1,024), and their full-size models were trained on 128 TPUs each.
LayaChess saw about 2 million examples, roughly 2,500 times fewer, over one weekend with about 22 hours of training on free GPUs, and it reached 27%. Picking a random legal move gets about 3%. So it's nowhere near a purpose-built model trained at that scale, and it isn't meant to be. The point of this version was to check that a general-purpose decision model can learn chess at all, and the curve says it can and is still climbing.
The more satisfying test was simply looking at the moves, before and after training:
- Opening move as White: b4 before, d4 after.
- Answer to 1.e4: h6 before. After, its top choices are d6, d5, e6, Nf6 and c6, all real openings.
- Scholar's Mate, with Qxf7 checkmate on the board: before, it preferred the quiet d3. After, it plays Qxf7 with a 94% win chance, and the next best move sits at 33%.
- A free queen left hanging: it takes it, also at 94%.
How strong is it?
Honestly, not very strong yet. In my early games against Stockfish at its weakest official setting (Elo 1320 on Stockfish's own scale), LayaChess has lost every game so far, although one of them lasted 88 moves. That setting sounds like a beginner, but it's a sneakily solid opponent that doesn't hang pieces the way beginners do.
My rough estimate is that this version plays somewhere around 1000 to 1200 against human players, on a chess.com-style scale. That's an estimate, not a measured rating. A proper rating from rated games is one of the next things on the list.
For context, DeepMind trained their searchless models on billions of examples. This one saw 2 million. It learned a surprising amount from that, and it clearly has a lot of room left.
Play it yourself
You can play against Laya on its Hugging Face Space. It's the same board as in the video, with an opening book, Laya's win chance for its candidate moves, and Stockfish's opinion alongside.
Because Laya judges one move at a time, every position means about 35 passes through a 421M-parameter model. Keeping a GPU server running around the clock for an experiment didn't make sense, so the Space runs on Hugging Face's free shared GPUs instead. Expect the first move to take a little while, and Laya's thinking time is capped at 10 seconds a move.
Prefer to run it on your own machine? The code and setup steps are on GitHub, and the model is on Hugging Face.
What's next
This is version 1, and I'll be releasing improvements soon. The main things I'm working on:
- Scoring all the moves of a position at once, so search can look much deeper in the same time.
- Training on more of the data. Two million moves is a small slice of what's available.
- A measured rating from real rated games instead of an estimate.
What I'd tell myself at the start
Check whether someone has already built the dataset before planning to build it yourself. Keep model inputs short, because half the tokens meant twice the training. Don't judge a run by one noisy number. And don't underestimate Stockfish at 1320.
Most of all: a general-purpose System 1 model really can pick up chess. It just needs a lot more practice than one experiment gave it.
Credits
Built on Laya by Convai Innovations, ChessBench by Google DeepMind, Stockfish and python-chess.
