Othello Arena
Sign in 한국어

Rules version 1: the ply and turn semantics this server implements. It does not change during a semester; a deviation found mid-term is recorded with the archive and applied to the next cohort.

Benchmark specification

Source: Yoo, Kim & Kim, The Expanded Othello AI Arena, TMLR 06/2026 (docs/reference/expanded_othello_tmlr2026.txt). §1–§5 define the benchmark as this arena runs it; §6 lists every departure from the paper and why.

1. The environment space

E = (L, C)
  • L (layout): the board, a grid of empty cells and obstacles with the standard four-disc opening. Fully observed by the agent.
  • C (rules): a pair (win condition, capture rule). Never passed to the agent.
axisvalues
win conditionmost: more discs wins; fewest: fewer discs wins; equal counts are a draw under both
capture rulestandard (all eight directions), orthogonal (four row and column directions), diagonal (four diagonal directions), short-1 / short-2 (a line flips only if it brackets at most one / two opponent discs), longest (legal where standard is; only the line flipping the most discs flips, ties by NW, N, NE, W, E, SW, S, SE)

The invariant mechanics (paper §2.2): a move is legal only if it flips at least one disc under the environment's capture rule; a player with no legal move passes; the game ends when neither player has a legal move. Obstacles are never playable and break a capture line. most with standard is the paper's own game.

2. Interaction protocol (paper §2.4)

  • What the agent receives: the layout and its budget at the start of an environment; on every turn the board, from the side to move, and the legal-move mask, computed under the real capture rule; after each budget game, the winner and the final disc counts. The rules themselves are never supplied.
  • Restricted environment access, the clause that decides the design:
"Access to an environment instance is limited to full game episodes under the standard interaction loop… information about the environment may be acquired only through actually playing games from the initial state to termination. Self-play is allowed. However, no additional simulator access is permitted: the protocol does not allow forward queries, branching rollouts, counterfactual state evaluations, or the use of a separately specified external opponent policy."

An agent may build and search its own model of the rules; it may not query the real environment as a simulator.

  • Feedback arrives only at termination: win, draw or loss.

3. Evaluation constraints (paper §3.2)

itemvalue
interaction budget2,000 games per environment
knowledge resetenvironments are independent sessions; here, a fresh container per environment
forbidden prior knowledgethe rules of the evaluation instance, which must be discovered inside its budget
permitted prior knowledgespatial priors about layouts, shipped with the agent, and general-purpose models not specialised to the evaluation instances

4. Instances

tierenvironmentsrules disclosed
publicthe paper's 7 layouts × 8 rule pairs = 56yes
laddergenerated: 6 to 12 rows and 6 to 12 columns, with up to a fifth of the cells as obstacles, rules drawn uniformly from both axesto staff only (see docs/rules.md §3)
finalgenerated the same way, run onceno

The eight public rule pairs are most/standard, fewest/standard, most/orthogonal, fewest/orthogonal, most/short-2, fewest/short-1, most/longest and fewest/longest. On the one layout where orthogonal stalls, that pair uses the objective's otherwise unused short rule instead.

Every generated environment is screened with random play and rejected if its draw rate is above 0.70, it averages fewer than 8 plies, one colour wins more than 0.88 of decisive games, it fills less than half of its playable cells, or it has no legal first move. diagonal has no legal first move from the central opening, so the screen rejects it.

5. Metrics

  • Win-rate curve against the fixed random opponent, measured at 0, 100, 500, 1000 and 2000 games used, 20 games per checkpoint with colours alternated. The measurement games cost no budget and send no feedback. The 0 point is the zero-shot measurement.
  • SAE (Skill-Acquisition Efficiency): the paper presents skill acquisition as a curve (§5.2); this arena also scalarises it as the normalised area under the curve over log budget, SAE = integral of win_rate(b) d(log b) / log(B_max / B_0), between the round's first and last non-zero checkpoints. Two SAEs compare only over the same window.
  • Draw rate, reported separately and never folded into the win rate: the paper found draw-heavy collapse to be the characteristic failure on its hardest conditions.
  • Standing, the ranked column: after its budget each agent is frozen and plays every other agent of the round in the same environment; match points (1 / ½ / 0) averaged over environments. A standing describes one round's field and does not compare across rounds.
  • A faulted environment counts as 0 on every metric rather than being dropped.

6. Departures from the paper

itempaperthis arenareason
rules Ca threshold on the final disc share, one of eight conditions per layouttwo axes: win condition and capture rule (§1)the professor's decision, 2026-09-30; most and fewest are exactly the paper's two unbounded conditions, and tests/test_paper_conformance.py holds the engine to the paper's equations for them
instancesthe 56 official presetspublic: 7 layouts × 8 rule pairs; ladder and final generateda published table of scored environments would turn inference into lookup
turn limittwo of the paper's eight conditions stop at ten turnsnone on any tier: every game is played to the enda truncated game scores a position nobody finished
knowledge resetsessions are independenta fresh container per environmentenforced rather than trusted; priors must be shipped
checkpoints100 / 500 / 1000 / 20000 / 100 / 500 / 1000 / 2000the zero-shot point is cheap and shows what an agent brought with it
games per checkpoint10 (inferred from the paper's tables and code)20ranking a cohort needs tighter error bars
opponentsRandom, PosNet, PPO-2K, MCTS-30 / 50 / 100the curve: random; after the budget: the round's own fieldchanged 2026-09-29; the field is a stronger and much cheaper opponent set
rankingnot defined by the paperround-robin standing, one round at a time; SAE published beside ita standing names its field, so it is not carried between rounds; SAE keeps the fixed random anchor
metricthe win-rate curvethe curve plus scalar SAEa board needs numbers; SAE is published, not ranked on
MCTS oracleits rollouts use an unseedable RNGa seeded, deterministic oracle, used in no scoreevery game here must replay from its move list

D-1, a paper and code discrepancy. The paper (§3.1) describes its ten-turn conditions as "10 moves per player", which is 20 plies; the released code compares n_turns against a ply counter, which is 10 plies. Nothing here depends on the answer, because no environment in this arena has a turn limit; the engine still supports one and tests/test_differential.py holds it to the authors' code.

Open questions for the authors. Whether 10 games per environment and opponent is what the paper used; their view of a scalar SAE; whether the permission to build this arena covers all three authors.