← Back to home

Benchmarks

A head-to-head round-robin across HedgeHog's engines (the free default FOX plus the premium Xerxes and Colossus) and the leading reference engines (XG, bgsage, GnuBG). Every engine runs under HedgeHog's own search, so the numbers compare the networks, not the software. Points-per-game come from 10-million-game matches; hover any cell for its win % and 95% confidence interval. Cells marked TBD are still computing.

Benchmarks

2026-07-20

0-ply · Cubeful · 10M games · 15 / 15 matchups complete

Each cell is the row engine's PPG vs the column engine (positive = row ahead). Hover for win % and 95% CI. TBD = still computing.

FOXXerxesXGColossusbgsageGnuBG
FOX·+0.0060+0.0049-0.0397+0.0178+0.0277
Xerxes-0.0060·-0.0001-0.0466+0.0140+0.0227
XG-0.0049+0.0001·-0.0468+0.0136+0.0206
Colossus+0.0397+0.0466+0.0468·+0.0582+0.0673
bgsage-0.0178-0.0140-0.0136-0.0582·+0.0130
GnuBG-0.0277-0.0227-0.0206-0.0673-0.0130·
  • FOX: HedgeHog's house engine: the prod default (fox-v0.2)
  • Xerxes: HedgeHog engine: fast, strong
  • XG: eXtreme Gammon (full import)
  • Colossus: HedgeHog engine: slow, stronger
  • bgsage: bgsage Stage 9 (19-NN backgame-aware)
  • GnuBG: GNU Backgammon

Evaluation speed (positions/sec)

EngineRaw (0-ply)1-ply2-ply3-ply
FOX195K6874.10.43
Colossus40K1051.40.24
GnuBG197K6153.80.38
bgsage133K3652.60.37
Xerxes237K7524.60.52
XG198K5253.50.38

Single-thread evaluation throughput. 0-ply is direct NN evaluation; 1/2/3-ply are expectiminimax searches using the production move filter at 2+ ply. All engines measured back-to-back on one pinned core. The 3-ply figures are high-variance (few completed searches). · single AMD Ryzen 5 3600 core

Methodology

Each matchup is decided by 10 million head-to-head games of self-play. The matrix shows cubeful results (with the doubling cube and its take/drop decisions), scored as cube-weighted money points; a cubeless (checker play only) round-robin is computed separately. Each cell is the row engine's points-per-game margin against the column engine, so the number flips sign when you swap the two. Every figure carries a 95% confidence interval. Hover any cell to see it.

Why we measure at 0-ply

0-ply means each engine plays straight from its neural network's evaluation, with no lookahead search. That isolates the quality of the network itself, which is what these benchmarks are meant to compare. Add search (1-ply, 2-ply, 3-ply) and every strong engine's lookahead corrects most of the same mistakes, so their results converge and the gap between networks shrinks toward zero. Differences that are clearly visible at 0-ply become hard to measure once deep search papers over them, so 0-ply is the most sensitive and honest way to tell two networks apart.

Why cubeful matters most

Cube decisions are the only decisions where a network's absolute probabilities matter. To double, take, or drop correctly, the engine has to know how likely each outcome actually is. Checker play asks far less: among the candidate moves for a roll, only their relative ordering matters: the engine just needs to rank the best move first, and a network can do that even if its raw win probabilities are miscalibrated. So the cubeful results stress a strictly harder skill, and a network that is merely good at ranking moves can still lose ground here.

Same search, different network

All competing engines are imported into the OGXF format and run inside HedgeHog's own engine. The move generation, search, and cube logic are therefore identical across every model; the only thing that changes is the neural network doing the evaluation. These numbers reflect differences between the networks, not between separate software implementations, and may not match results published by the original programs.