Benchmarks
A head-to-head round-robin across HedgeHog's engines (the free default FOX plus the premium Xerxes and Colossus) and the leading reference engines (XG, bgsage, GnuBG). Every engine runs under HedgeHog's own search, so the numbers compare the networks, not the software. Points-per-game come from 10-million-game matches; hover any cell for its win % and 95% confidence interval. Cells marked TBD are still computing.
Benchmarks
2026-07-200-ply · Cubeful · 10M games · 15 / 15 matchups complete
Each cell is the row engine's PPG vs the column engine (positive = row ahead). Hover for win % and 95% CI. TBD = still computing.
| FOX | Xerxes | XG | Colossus | bgsage | GnuBG | |
|---|---|---|---|---|---|---|
| FOX | · | +0.0060 | +0.0049 | -0.0397 | +0.0178 | +0.0277 |
| Xerxes | -0.0060 | · | -0.0001 | -0.0466 | +0.0140 | +0.0227 |
| XG | -0.0049 | +0.0001 | · | -0.0468 | +0.0136 | +0.0206 |
| Colossus | +0.0397 | +0.0466 | +0.0468 | · | +0.0582 | +0.0673 |
| bgsage | -0.0178 | -0.0140 | -0.0136 | -0.0582 | · | +0.0130 |
| GnuBG | -0.0277 | -0.0227 | -0.0206 | -0.0673 | -0.0130 | · |
- FOX: HedgeHog's house engine: the prod default (fox-v0.2)
- Xerxes: HedgeHog engine: fast, strong
- XG: eXtreme Gammon (full import)
- Colossus: HedgeHog engine: slow, stronger
- bgsage: bgsage Stage 9 (19-NN backgame-aware)
- GnuBG: GNU Backgammon
Evaluation speed (positions/sec)
| Engine | Raw (0-ply) | 1-ply | 2-ply | 3-ply |
|---|---|---|---|---|
| FOX | 195K | 687 | 4.1 | 0.43 |
| Colossus | 40K | 105 | 1.4 | 0.24 |
| GnuBG | 197K | 615 | 3.8 | 0.38 |
| bgsage | 133K | 365 | 2.6 | 0.37 |
| Xerxes | 237K | 752 | 4.6 | 0.52 |
| XG | 198K | 525 | 3.5 | 0.38 |
Single-thread evaluation throughput. 0-ply is direct NN evaluation; 1/2/3-ply are expectiminimax searches using the production move filter at 2+ ply. All engines measured back-to-back on one pinned core. The 3-ply figures are high-variance (few completed searches). · single AMD Ryzen 5 3600 core
Methodology
Each matchup is decided by 10 million head-to-head games of self-play. The matrix shows cubeful results (with the doubling cube and its take/drop decisions), scored as cube-weighted money points; a cubeless (checker play only) round-robin is computed separately. Each cell is the row engine's points-per-game margin against the column engine, so the number flips sign when you swap the two. Every figure carries a 95% confidence interval. Hover any cell to see it.
Why we measure at 0-ply
0-ply means each engine plays straight from its neural network's evaluation, with no lookahead search. That isolates the quality of the network itself, which is what these benchmarks are meant to compare. Add search (1-ply, 2-ply, 3-ply) and every strong engine's lookahead corrects most of the same mistakes, so their results converge and the gap between networks shrinks toward zero. Differences that are clearly visible at 0-ply become hard to measure once deep search papers over them, so 0-ply is the most sensitive and honest way to tell two networks apart.
Why cubeful matters most
Cube decisions are the only decisions where a network's absolute probabilities matter. To double, take, or drop correctly, the engine has to know how likely each outcome actually is. Checker play asks far less: among the candidate moves for a roll, only their relative ordering matters: the engine just needs to rank the best move first, and a network can do that even if its raw win probabilities are miscalibrated. So the cubeful results stress a strictly harder skill, and a network that is merely good at ranking moves can still lose ground here.
Same search, different network
All competing engines are imported into the OGXF format and run inside HedgeHog's own engine. The move generation, search, and cube logic are therefore identical across every model; the only thing that changes is the neural network doing the evaluation. These numbers reflect differences between the networks, not between separate software implementations, and may not match results published by the original programs.