← Learn backgammon

How backgammon engines evaluate positions

What a neural network sees in a position, what 2-ply and 3-ply mean, how a rollout plays a position out, and how precise each answer is.

#The network: a judgement at a glance

At the heart of every strong backgammon engine is a neural network: a function that reads a position and returns the chances of the six ways the game can end (win or lose a single game, a gammon or a backgammon). Those chances give the position's equity.

Nobody writes the rules it follows. It learns them from games, classically by playing itself millions of times and adjusting after each game towards what actually happened, the method Gerald Tesauro's TD-Gammon showed in the early 1990s. What comes out is closer to a strong player's intuition than to a calculation: it evaluates a position in a fraction of a millisecond, is usually very good, and is occasionally wrong in ways no strong human would be.

#Looking ahead: what a ply is

An engine does not have to trust its first impression. It can look ahead, and the depth it looks is counted in plies, one ply per turn of lookahead:

  • 1-ply: try every legal play and ask the network what each resulting position is worth. The best number wins.
  • 2-ply: for each candidate play, consider all 21 different rolls the opponent could throw next, find their best reply to each, and average what the network says about those positions, weighting each roll by how often it comes up.
  • 3-ply: one turn deeper again, averaging over your own next roll as well.

Each ply multiplies the work by roughly twenty, because every roll has to be tried against every reasonable reply. Engines keep it affordable by pruning: plays that look hopeless at a shallow depth are not searched deeper. Deeper is not automatically right, but errors in the network's judgement tend to average out over the rolls, so each ply usually improves the answer.

Programs count plies differently. HedgeHog and eXtreme Gammon call the plain network evaluation 1-ply; GNU Backgammon calls the same thing 0-ply, so its 2-ply is HedgeHog's 3-ply. Which depth HedgeHog's analysis offers, and which one to use, is in Analysing a match.

#Rollouts: playing it out

A rollout answers the question the direct way: it plays the position out, many times, and counts what happened. Each game is played by the engine against itself, choosing every move at some fixed depth, and the results are averaged. Played to the end with enough games, a rollout does not depend on the network's judgement of the starting position at all, only on the quality of the moves in the games it plays, which is why rollouts are the reference that analysis depths are measured against.

Three techniques make a rollout far more precise than the same number of games played at random:

  • Stratified dice. Instead of rolling the first roll at random, HedgeHog's rollouts give each of the 36 possible first rolls the same share of the games, and with 1296 games or more every combination of the first two rolls the same share. The luck of those rolls then drops out of the result entirely.
  • Variance reduction. On every turn the engine works out what the average roll would have been worth, before the dice are thrown, and after the throw it notes how much better or worse the actual roll was. That difference is the luck of the roll, and subtracting it at the end leaves an unbiased result with much less noise. It costs about twenty times the work per game and is worth it.
  • Truncation. A rollout does not have to play every game to the end. It can stop after a number of turns and score the position it reached with the network, which saves time when the remaining play is simple.

#How precise a rollout is

Every rollout comes with a standard error: how far its answer could plausibly be from the one an infinitely long rollout would give. Two plays are separated by a rollout only when the gap between them is larger than about two standard errors of the difference. Closer than that, they are equal as far as the rollout can tell, and should be reported that way.

The opening move pages are built this way: each candidate play rolled out 2592 times to the end of the game, with a standard error of about 0.002 each. That is precise enough to separate 3-1's 8/5 6/5 from every alternative by a wide margin, and not precise enough to separate the top two plays of 2-1, which the page therefore lists as equals.

#Exact answers in the endgame

Once both sides are bearing off, there are few enough positions left to compute exactly, at least for the ones with checkers on the lower points. HedgeHog carries bearoff tables: precomputed databases that hold the exact chances for every position in them, found by working backwards from the end of the game. There the engine does not estimate; it looks the answer up. The network, the search and the rollouts are for everything before that point.

#What this means for your analysis

An analysis is the engine's opinion at a chosen depth, not the truth, and the difference matters most where two plays are close. A play flagged as a small inaccuracy at 2-ply may be correct at 3-ply, and the reverse. The large errors are the ones to trust and to study first: when a play is worse by a tenth of a point, no depth is going to change the verdict. The error rate that sums them up, and how far to trust it, is in Backgammon error rates and PR.