#What a rollout setting trades
A rollout plays a position out many times and averages the results. Every setting in a rollout dialog trades time against one of two things: noise, the luck of the dice that a finite number of games leaves in the answer, and bias, the error that comes from playing those games with imperfect moves or stopping them early. More games only cure the first. Better moves only cure the second.
This article measures both on a common rollout shape: 360 games, each stopped after seven moves and scored by the network, with variance reduction. That is the shape of eXtreme Gammon's XGRoller+ level as its manual describes it. Every number below comes from HedgeHog's engine and HedgeHog's networks.
#How precise a short rollout is
We rolled out six middle-game positions forty times each, every time with fresh dice, and measured how much the answer moved from one rollout to the next.
- With variance reduction, the equity of a 360-game, seven-move rollout typically moves by about 0.0045 (one standard deviation), between 0.0035 and 0.0061 depending on the position. Two plays closer than about 0.012 are not separated by one such rollout.
- Without it, the same rollout moves by about 0.022. Variance reduction removes about 24 times the variance, which is what 24 times as many games would buy.
The point of a short rollout is speed. Even so, its answer is precise enough to separate most plays a 2-ply or 3-ply search cannot.
#The reported error has to come from each game
A rollout reports its standard error along with the answer, and that error should be computed from the equity of each game, not from the error of each outcome separately. The six outcome chances are not independent: a game won with a gammon is also a game won, so the chance of winning and the chance of a gammon rise and fall together. Adding up their separate errors as if they were unrelated leaves out exactly that, and every piece it leaves out has the same sign. On contact positions the true spread can be up to twice what the outcome-by-outcome error reports.
HedgeHog computes it from each game's own equity, and the reported error matches the spread measured above. A standard error that is too small is worse than none: it makes two plays look separated when they are not.
#Deeper moves where they matter
The moves played inside a rollout are usually chosen at the quickest depth, the plain network evaluation. Its mistakes are a source of bias that no number of games removes. They matter most at the start: a misplay on the first move shifts every game that follows it, while one twenty moves in mostly washes out.
So engines let the first moves be searched deeper. XGRoller+ plays the first two moves at 2-ply and the rest at 1-ply, and GNU Backgammon offers the same choice.
The catch is variance reduction. It needs the value of the average roll on every turn, which means choosing a move for all 21 rolls at the same depth as the move actually played. With 2-ply first moves that is 21 two-ply searches per turn, per game. In our test, adding 2-ply first moves made a rollout of the opening position 120 times slower: 388 seconds instead of 3.1. Two things bring that back down.
#Pruning: not searching every play
A typical roll has around twenty legal plays, and most of them are obviously bad. Engines rank every play with the quick network evaluation first and search only the promising ones at the full depth: GNU Backgammon calls the rule a move filter, eXtreme Gammon a search interval. With a strong network that saves most of the work for almost no equity. At 2-ply, eXtreme Gammon's first step (up to 8 plays within 0.16 of the best) searches about a fifth of the plays and costs about 0.00002 per decision with HedgeHog's strongest networks.
How the filters work, why the same filter costs five times as much on one network as on another, and how far each network can be pruned before it costs more than eXtreme Gammon's own setting does are in Move filters.
#The first-move cache
The last saving is the largest, and it costs nothing in accuracy. Every game of a rollout starts from the same position, so the first move only has 21 different rolls to answer, however many games are played. eXtreme Gammon's study puts it plainly: the first move "need[s] to be calculated only 21 times (the rest of the time the program will get the result from the cache)". The same holds for the average-roll value variance reduction needs at the start, and nearly as well for the second move.
Remembering those answers instead of recomputing them in every game changes no result, only the time. For the opening position, 360 games stopped after seven moves, with variance reduction:
| setting | without the cache | with the cache |
|---|---|---|
| every move at network depth | 3.1 s | 2.4 s |
| first two moves at 2-ply | 388 s | 16.4 s |
| first two moves at 2-ply, search interval pruning | 109 s | 6.3 s |
With pruning and the cache together, searching the first two moves at 2-ply costs about two and a half times a plain rollout instead of 120 times.
#What to take from it
- Read a rollout's standard error before its ranking. A short rollout with variance reduction is good to about 0.005 per play. Plays closer than about 0.012 are equal as far as it can tell.
- Spend extra time on the first moves, not on wider pruning. That is where a move error biases every game, and with a cache it is cheap. eXtreme Gammon's own advice is the same.
- A filter is a bet on the network. With a strong network the published defaults cost almost nothing and can be tightened; with a weaker one they need widening. Move filters has the numbers.
How deep HedgeHog's analysis searches and what each preset costs is in Analysing a match.