Benchmark specification
Source: Yoo, Kim & Kim, The Expanded Othello AI Arena, TMLR 06/2026 (docs/reference/expanded_othello_tmlr2026.txt). §1–§5 define the benchmark as this arena runs it; §6 lists every departure from the paper and why.
1. The environment space
E = (L, C)
- L (layout): the board, a grid of empty cells and obstacles with the standard four-disc opening. Fully observed by the agent.
- C (rules): a pair (win condition, capture rule). Never passed to the agent.
| axis | values |
|---|---|
| win condition | most: more discs wins; fewest: fewer discs wins; equal counts are a draw under both |
| capture rule | standard (all eight directions), orthogonal (four row and column directions), diagonal (four diagonal directions), short-1 / short-2 (a line flips only if it brackets at most one / two opponent discs), longest (legal where standard is; only the line flipping the most discs flips, ties by NW, N, NE, W, E, SW, S, SE) |
The invariant mechanics (paper §2.2): a move is legal only if it flips at least one disc under the environment's capture rule; a player with no legal move passes; the game ends when neither player has a legal move. Obstacles are never playable and break a capture line. most with standard is the paper's own game.
2. Interaction protocol (paper §2.4)
- What the agent receives: the layout and its budget at the start of an environment; on every turn the board, from the side to move, and the legal-move mask, computed under the real capture rule; after each budget game, the winner and the final disc counts. The rules themselves are never supplied.
- Restricted environment access, the clause that decides the design:
"Access to an environment instance is limited to full game episodes under the standard interaction loop… information about the environment may be acquired only through actually playing games from the initial state to termination. Self-play is allowed. However, no additional simulator access is permitted: the protocol does not allow forward queries, branching rollouts, counterfactual state evaluations, or the use of a separately specified external opponent policy."
An agent may build and search its own model of the rules; it may not query the real environment as a simulator.
- Feedback arrives only at termination: win, draw or loss.
3. Evaluation constraints (paper §3.2)
| item | value |
|---|---|
| interaction budget | 2,000 games per environment |
| knowledge reset | environments are independent sessions; here, a fresh container per environment |
| forbidden prior knowledge | the rules of the evaluation instance, which must be discovered inside its budget |
| permitted prior knowledge | spatial priors about layouts, shipped with the agent, and general-purpose models not specialised to the evaluation instances |
4. Instances
| tier | environments | rules disclosed |
|---|---|---|
public | the paper's 7 layouts × 8 rule pairs = 56 | yes |
ladder | generated: 6 to 12 rows and 6 to 12 columns, with up to a fifth of the cells as obstacles, rules drawn uniformly from both axes | to staff only (see docs/rules.md §3) |
final | generated the same way, run once | no |
The eight public rule pairs are most/standard, fewest/standard, most/orthogonal, fewest/orthogonal, most/short-2, fewest/short-1, most/longest and fewest/longest. On the one layout where orthogonal stalls, that pair uses the objective's otherwise unused short rule instead.
Every generated environment is screened with random play and rejected if its draw rate is above 0.70, it averages fewer than 8 plies, one colour wins more than 0.88 of decisive games, it fills less than half of its playable cells, or it has no legal first move. diagonal has no legal first move from the central opening, so the screen rejects it.
5. Metrics
- Win-rate curve against the fixed
randomopponent, measured at 0, 100, 500, 1000 and 2000 games used, 20 games per checkpoint with colours alternated. The measurement games cost no budget and send no feedback. The 0 point is the zero-shot measurement. - SAE (Skill-Acquisition Efficiency): the paper presents skill acquisition as a curve (§5.2); this arena also scalarises it as the normalised area under the curve over log budget,
SAE = integral of win_rate(b) d(log b) / log(B_max / B_0), between the round's first and last non-zero checkpoints. Two SAEs compare only over the same window. - Draw rate, reported separately and never folded into the win rate: the paper found draw-heavy collapse to be the characteristic failure on its hardest conditions.
- Standing, the ranked column: after its budget each agent is frozen and plays every other agent of the round in the same environment; match points (1 / ½ / 0) averaged over environments. A standing describes one round's field and does not compare across rounds.
- A faulted environment counts as 0 on every metric rather than being dropped.
6. Departures from the paper
| item | paper | this arena | reason |
|---|---|---|---|
rules C | a threshold on the final disc share, one of eight conditions per layout | two axes: win condition and capture rule (§1) | the professor's decision, 2026-09-30; most and fewest are exactly the paper's two unbounded conditions, and tests/test_paper_conformance.py holds the engine to the paper's equations for them |
| instances | the 56 official presets | public: 7 layouts × 8 rule pairs; ladder and final generated | a published table of scored environments would turn inference into lookup |
| turn limit | two of the paper's eight conditions stop at ten turns | none on any tier: every game is played to the end | a truncated game scores a position nobody finished |
| knowledge reset | sessions are independent | a fresh container per environment | enforced rather than trusted; priors must be shipped |
| checkpoints | 100 / 500 / 1000 / 2000 | 0 / 100 / 500 / 1000 / 2000 | the zero-shot point is cheap and shows what an agent brought with it |
| games per checkpoint | 10 (inferred from the paper's tables and code) | 20 | ranking a cohort needs tighter error bars |
| opponents | Random, PosNet, PPO-2K, MCTS-30 / 50 / 100 | the curve: random; after the budget: the round's own field | changed 2026-09-29; the field is a stronger and much cheaper opponent set |
| ranking | not defined by the paper | round-robin standing, one round at a time; SAE published beside it | a standing names its field, so it is not carried between rounds; SAE keeps the fixed random anchor |
| metric | the win-rate curve | the curve plus scalar SAE | a board needs numbers; SAE is published, not ranked on |
| MCTS oracle | its rollouts use an unseedable RNG | a seeded, deterministic oracle, used in no score | every game here must replay from its move list |
D-1, a paper and code discrepancy. The paper (§3.1) describes its ten-turn conditions as "10 moves per player", which is 20 plies; the released code compares n_turns against a ply counter, which is 10 plies. Nothing here depends on the answer, because no environment in this arena has a turn limit; the engine still supports one and tests/test_differential.py holds it to the authors' code.
Open questions for the authors. Whether 10 games per environment and opponent is what the paper used; their view of a scalar SAE; whether the permission to build this arena covers all three authors.