Competition Rules
Expanded Othello AI Arena — self-hosted edition. Based on Yoo, Kim & Kim, The Expanded Othello AI Arena, TMLR 2026. Where this arena departs from the paper, it says so.
1. What you are building
An agent that walks into an Othello variant it has never seen, is told the shape of the board and nothing else, and works out what winning means here from the win/draw/loss signal at the end of each game — inside a budget of 2,000 games.
The board mechanics never change: you place a disc so that it sandwiches at least one opponent disc along one of the eight directions, those discs flip, and a move is legal only if it flips something. Obstacles are permanently unplayable and block flip chains.
What changes is the objective.
2. How winning is decided
At the end of a game each side has an occupancy ratio
rho = my_discs / (my_discs + opponent_discs)
Every environment has a threshold K, which induces an interval
I(K) = ( min(K, 0.5) , max(K, 0.5) ) <- open interval
You win if and only if your final rho lands inside I(K). If neither side lands inside it, the game is a draw. Because the two ratios sum to 1, both sides can never be inside at once.
K | interval | what it asks for |
|---|---|---|
> 1.0 | (0.5, 1.0] | plain majority — take as many discs as you can |
0.5 < K <= 1.0 | (0.5, K) | majority, but stay under K |
0 <= K < 0.5 | (K, 0.5) | minority, but stay above K |
< 0 | (0, 0.5) | plain minority — take as few as you can |
Worked example, K = 0.8, so I(K) = (0.5, 0.8):
- you finish with 75% of the discs → you win
- you finish with 85% of the discs → draw. You dominated the board and got nothing for it, and neither did your opponent.
That second line is the whole game. Crushing your opponent is not the same as winning, and on a narrow interval it is indistinguishable from losing. On some environments the best move is to hand discs back.
The ratio is read off a finished board. No environment ends early here: every game runs until neither player has a legal move, so nothing is ever scored mid-game (§3).
3. The three tiers
| tier | environments | is K disclosed? | what it is for |
|---|---|---|---|
| public | the paper's 56, with two conditions changed | yes | practice. Run them locally as much as you like. |
| ladder | freshly generated, published with their K | yes | the open leaderboard |
| final | freshly generated, never run until the deadline | no | final standings |
The public tier discloses K deliberately: it is the practice tier, and you need somewhere your inference can be checked.
The ladder discloses K as well. That is a change, and the reason is that on the ladder it could not honestly be kept. A ladder run may use a general-purpose model you pay for, and the account that pays for it can read what was sent to it: an agent that has worked out K inside the container can write the threshold into that traffic, and you read it back afterwards. The connection is to a company we do not control and nothing on this side can inspect it. A secret kept only from the people who do not go looking is not a secret, so the ladder publishes its thresholds along with the tier. What the open leaderboard measures is how well your agent plays an objective it has been told — which is a real question, and a different one from the paper's.
Told by /envs, in practice: it lists every ladder environment with its geometry and its K, and env_info hands your agent that same geometry, so an agent that ships the table can match the board it is given against it and read the threshold off. That is intended, not a loophole.
The final tier still hides K, and it can, because it runs exactly once. An agent may stream the threshold out in the middle of a final round and buy nothing with it: there is no second final round to spend it in. The paper's task — walk in, work out what winning means here — is set on the final, which the course is scored on and the open benchmark is not.
Beyond that, a scored tier publishes nothing: not the generation seed, which would reproduce every environment in the tier, and not which environment you are playing while you are playing it (§7).
Every game is played to the end, in every tier. There is no turn limit anywhere in this arena, so it is not one of the things you have to infer: the only part of a scored environment a tier can withhold is the disc-ratio threshold K, and only the final withholds it. The paper's own suite ends two of its eight conditions after ten plies; those two are replaced here by two further thresholds, so the practice tier is still 56 environments over seven layouts.
The scored tiers are not the paper's 56 presets. Those presets and their K values are published on GitHub, so a scored suite built from them would rank agents on a table anyone can download. Hard-coding the 56 public K values into your agent therefore gains nothing on the ladder or the final: none of those environments is there, on either tier. What a lookup table like that does do is inflate your zero-shot column on the practice tier, which is not scored.
4. The interaction budget
2,000 games per environment. Within that budget you play, you are told who won, and you adapt. That is the entire channel.
Win rate is measured at 0, 100, 500, 1000 and 2000 games. Those measurement games are extra — they cost you no budget and you are told nothing about them. You cannot tell a measurement game from a self-play game, which is intentional.
The 0 point is the zero-shot measurement, taken before you have played anything in that environment. It shows how much of your performance you brought with you — priors about board geometry — as opposed to what you inferred.
After the budget: you play the other agents
Changed 2026-09-29. When your budget is spent, your agent is frozen and plays every other agent in that round once, in the same environment, before the containers come down. A pairing is 20 games, ten as each colour. That round-robin is what the leaderboard is ranked on (§6), and it replaced a fixed pool of two Oracle opponents that used to be played at the last checkpoint — the Oracle bots are still in the engine and are no longer part of a score.
Three things follow, and they are the whole of what it changes for you.
- You are told nothing. A pairing spends no budget and sends no result:
freeze()has already been sent when the first pairing starts, and a frozen agent is refused every outcome by the host it runs under. You cannot tell a pairing from a measurement game — sameprepare, same board encoding, sameact— and the order the pairings are played in cannot change what anybody scores. - Your opponent is another student's code. What your agent observes in those games is another person's play. Nothing tells it who, and it cannot tell which environment of the tier it is in either. An agent that recognises an opponent's move distribution and plays differently against it is possible and is not detectable from the wire; what bounds it is that these games carry no feedback, so anything like that has to be shipped rather than learned.
- A dead opponent is not a win. If the other agent's container has died or run out of clock, the pairing is void for you — out of your numerator and your denominator both — and zero for it. A win you did not play is not a win, and your row publishes the number of games it actually played.
5. What you may and may not do
These come from the paper (§2.4 Environmental Access, §3.2 Evaluation Constraints) and are the difference between a benchmark and a puzzle.
Allowed
- Simulate the dynamics yourself. The flipping rule and the legality rule are public and invariant. Build your own model of the board and search in it as deeply as you like — the paper's own baseline runs a depth-3 minimax.
- Self-play. Both seats during the budget are yours.
- Prior knowledge about board layouts. Positional heuristics meta-learned over a distribution of geometries are explicitly permitted, and rewarded. Ship them with your submission as data, in a format a human can open — JSON, CSV, a text table, an
.npz. An opaque binary blob will be asked about before the final standings are published. - General-purpose models, including LLMs, provided what they encode is not specific to this arena's scored environments.
Not allowed
- Querying the real environment as a simulator. You may only learn about an environment by playing complete games in it, start to termination. No forward queries, no branching rollouts, no counterfactual evaluation of hypothetical positions, no externally specified opponent policy. The arena enforces this structurally: your code runs in a container that holds no environment at all, and speaks to it only through the move protocol.
- Carrying knowledge between scored environments. Each environment is an independent session. Your process is restarted for every environment, so nothing you learned in the previous one survives — not instance state, not class attributes, not a cache on disk. If you want a prior, ship it.
- Reaching outside the sandbox. No writing outside
/tmp. No attempt to read the host, the environment definition, or another submission. Network access is the one part of this that the tier decides: a final run's container has none at all, and a ladder run's has one — which is exactly why the ladder no longer pretends itsKis a secret (§3). - Signing in as somebody else, or letting somebody sign in as you. See §9. Handing your account over to a classmate — to upload for you, to "just look" — is an integrity violation, not a convenience: every upload is attributed to the account that made it, and a submission the arena cannot attribute is a submission nobody gets credit for.
- Entering as a pair or a team. There are no teams here. An entry is one account, one account is one GitHub account, and the arena has no way to represent two people behind one entry — so two people behind one account is the line above, not a team. Talking about the problem with somebody else is fine and always was; submitting one agent under one name for two people is not.
The line, stated plainly
Inferring K by playing games is the task the final tier sets, and obtaining a final K by any other route is cheating, however cleverly. On the ladder there is nothing to obtain: those thresholds are published (§3), and an agent that reads them off /envs is playing the game that tier sets.
6. How you are scored
Three numbers per environment, then aggregated over the tier.
Win rate at each checkpoint — the primary signal.
Draw rate, reported separately and never folded into the win rate. In the narrow bands the paper finds every published method collapsing into draws (73.6% at K = 0.6, 72.9% at K = 0.4). A draw-heavy agent has learned to avoid losing terminal ratios without learning to hit the winning ones; that is real progress and it is worth seeing, but it is not a win.
Skill-Acquisition Efficiency (SAE). The paper defines skill-acquisition efficiency as the rate of convergence and reports it as a curve rather than a single number, so the scalarisation here is this arena's own: the normalised area under the win-rate curve plotted against log budget.
SAE = integral of win_rate(b) d(log b) / log(b_max / b_0)
Reaching 70% after 100 games beats reaching 70% after 2,000. That is the point of the benchmark: not how good you end up, but how fast you got there. The win rate in that integral is measured against the same fixed opponent — random — at every checkpoint, in every round. That fixed anchor is what a standing does not have, and it is why an SAE can be compared with another round's at all.
B_0 and B_max are the round's own first and last checkpoints, so SAE is comparable inside a window and not across windows. It is a definite integral, and two integrals are the same measurement only over the same limits. A round whose checkpoints are 0, 100, 500, 1000, 2000 integrates over [100, 2000]; a shakeout whose checkpoints are 0, 20, 60 integrates over [20, 60], and those two spans do not even overlap. The same games can give three different SAEs over three legal windows — measured on a real 28-environment round, one agent's own games gave 0.421, 0.405 and 0.402 over three of them. So:
- every page that shows you an SAE shows you the window it used. The leaderboard prints
B_0andB_maxin the formula under the board and the checkpoints they come from, your round page and your history row name the same pair, and the grading export carries them as thesae_b0andsae_b_maxcolumns besidesae. If two of your rows show different windows, those two numbers are not a trend and nothing here will treat them as one; - your best SAE is your best within one window. A grade export that reaches across rounds compares only rows measured over the same window, and a row measured over a different one is carried with the reason in its
statuscolumn rather than silently entering a maximum it does not belong in.
A round queued at the canonical budget always gets [100, 2000], so for every scored ladder and final round of the term this is one window and the rule costs you nothing. It bites on shakeouts and short diagnostics, which is where it is supposed to bite.
SAE was the leaderboard's sort key until 2026-09-29 and is not any more. It is published on every row, beside the column that is now sorted on:
The standing — your match points from the round-robin of §4, averaged over the environments with equal weight. One point for a win, half for a draw, none for a loss; wins, draws and losses are published separately and never folded into it. It says how you did against the agents that were in that round, and it is not comparable with any other round's: a different field is a different measurement. That is why the leaderboard shows one round and names it, rather than your best result across several, and why your history page's rank column is a list of placings rather than a trend.
The two columns can disagree, and when they do neither is wrong: an agent that beats these classmates while learning slowly is third on the standing with the lowest SAE in the field.
A standing is only published for a field of 8 or more. Below that, a rank is the individual pairing results rather than a summary of them — at two entrants it is literally one pairing — so the round runs in full, its participants read their own numbers, and no board changes.
Also reported, not ranked on: generalisation gap (practice tier versus scored tier) and robustness (spread across environments).
Faulted environments count as zero
If your agent does not finish an environment — it times out, crashes, breaks the protocol, or forfeits every game — that environment is not dropped from your average. It counts as a win rate of 0 and an SAE of 0, and the row shows how many environments faulted and why. The kinds that are yours: timeout (you ran out of clock: either one reply did not arrive inside the per-move limit, or the whole environment ran past its ceiling — §7 publishes both, and one kind covers them because both mean the same thing to your score), a crash (eof, agent_error, protocol — your process died, raised, or broke the wire protocol), budget (your submission used up its share of the round's CPU time) and forfeit. A submission scored on 27 environments is therefore never compared against one scored on 28 as if the missing one had never existed.
The exception is a fault that was the arena's, not yours: launch_failed (the container did not start), pipe (the host side of the connection broke), oom (the container was killed for memory on a shared box, which is as likely to be the box as your agent) and infrastructure (the grader itself was unavailable, or its disk filled). Those kinds are not scored as zero: they are re-scored before the round is published, and if too many of them happen the round is held back rather than published at all.
A wall-clock timeout is your fault even when it feels unlucky: it is not reproducible, so it is not re-run. Leave headroom under the limits in §7 — and the one to leave headroom under is the sustained pace, not the per-call act() budget. make check runs every public board shape and measures your mean against each board's pace, so both halves of "test on the largest public board, and be fast enough on it" are things it now does for you — and the smoke test on upload applies the same two rules, so an agent the check refuses is refused on upload as well instead of finding out in a scored round.
7. Submissions
| what | a single agent.py, or a .zip with agent.py at its root |
| entry point | a class named Agent, or a factory make_agent() |
| runtime | Python 3.11 with numpy 2.2.6, nothing else installed |
| network | decided by the tier: a ladder container has egress, a final container runs --network none |
| filesystem | read-only, except a writable /tmp that does not survive the environment |
The runtime row is the image, and it is the same sentence in the starter kit's README: the grader is one Docker image, python:3.11-slim with that one numpy pinned, and the smoke test you get on upload runs in it. Your obs arrives as a (3, H, W) float32 numpy array. If your local Python is older or newer than 3.11, keep to syntax both have; make setup (or python check.py --setup) tells you which interpreter and numpy it found, and every make check / python check.py run repeats those two lines under the kit version.
The network row is the one line of the contract the image does not decide. On the ladder your container is given egress, so an agent may call a general-purpose model over your own API account, with no limit on the number of calls. Egress is the reason the ladder publishes its K (§3). It costs you nothing but time, and time is scarce: the call has to return inside the per-call act limit below, and the waiting counts against your submission's share of the round. No environment variable of the host's reaches the container, so an API key has to travel inside your submission, which staff read. On the final there is no network at all, so an agent that cannot play without one cannot play that tier.
Your agent implements three methods:
class Agent:
def prepare(self, env_info): ... # geometry and budget. Never K.
def act(self, obs, action_mask) -> int # flat index = row * W + col
def observe_outcome(self, outcome): ... # who won, and the final disc counts
env_info carries rows, cols, obstacles and interaction_budget. It does not carry an environment index or the number of environments in the tier: those fields are pinned to 0 and 1 on the wire and say nothing about where in a tier you are or how large it is. Which tier you are on is another matter, and an agent can work it out — one socket call answers it, because the network row above is decided by the tier. That is allowed, and it buys nothing: the ladder's thresholds are published anyway, and the final runs exactly once, so there is no second final round to spend the knowledge in.
An illegal move costs you the turn rather than the game — a partly wrong world model should degrade, not collapse. Enough of them in one game forfeits it, and an agent that never produces a legal move forfeits immediately.
Limits
Every limit the grader enforces, rendered from the grader's own configuration so this table cannot drift from what actually happens to you:
| limit | value | note |
|---|---|---|
| single file | 2 MiB | a lone agent.py |
| archive | 20 MiB | as uploaded and once unpacked; at most 10000 members |
prepare() | 15 s | once per environment; the grader waits no longer for it than for any other call, so past this the environment is lost (30 s in a scored round, but the smoke test is the gate) |
act() sustained pace | 7.6 ms | the mean your moves have to keep to. One environment is 2000 self-play games and 100 evaluation games inside 1800 s, and you move for both sides in self-play, so the largest board asks you for 237800 moves. Only your own answers are on this clock, and the tournament's pairings have an allowance of their own at this same pace, so keeping it is the whole of what is asked. This is far tighter than the two per-call limits below and it is the one that decides whether you finish; an environment you do not finish scores zero |
act() soft budget | 1 s | an overrun is recorded, the game goes on. Nothing stops you spending it, but spending it costs you the environment: see the 7.6 ms pace above |
act() hard limit | 10 s | the same in every tier; past this the call is unwound and you lose the turn |
| no reply at all, smoke test | 15 s | past this the container is treated as hung |
| no reply at all, scored round | 30 s | past this the environment is a fault |
| one environment, smoke test | 180 s | 20 self-play games and 4 evaluation games |
| one environment, scored round | 1800 s | 2000 self-play games and 100 evaluation games; the clock counts the time your agent spends answering, not the grader's or your opponents' |
| CPU | 1 | one core, no hyperthreads to spare |
| memory | 2 GiB | swap disabled |
| processes and threads | 64 | combined, and counted by the kernel for the whole container. A process over the line gets Resource temporarily unavailable; a thread gets can't start new thread, which is the same limit wearing its other name |
The first act() row is the one to build to. The soft and hard budgets below it bound a single call and nothing bounds a single call tightly; what bounds a scored round is the sum, and the sum is what the environment ceiling is. Spending the published one-second budget on every move is legal, records no overrun, and finishes no environments at all.
Resource temporarily unavailable in your log means the process limit, and can't start new thread is that same limit reached from the other side: the ceiling counts processes and threads together. A multiprocessing.Pool or a thread pool on one CPU buys you nothing anyway.
Before you upload
Run make check — or, without make, python check.py — before uploading. It runs exactly the gate the server runs and prints the kit version on its first line; /kit/version shows the version the arena is serving. If the two differ, or the announcement banner says the kit was updated, download it again from /kit and re-run the check. make package (or python package.py) builds submission.zip and checks the zip itself.
When your upload becomes the one that is scored
An upload is screened, then smoke-tested in the real container on one practice environment per public board shape — the same five make check sweeps, and for the same reason: a bug that only shows up on a board you never tried is the one worth catching before a round, not during one. Each of those runs measures the mean time your act() takes and compares it with that board's own pace (§6) — the same rule make check applies, on the same boards — so a submission that cannot keep the pace fails here rather than scoring zero over 28 environments, and the log tells you the number you averaged and the number the board allows. An agent far over the pace is stopped early instead of being let spend the whole budget. The first board that refuses it is the one the log names; the run stops there.
It becomes your active submission only when the smoke test passes; until then, and if it fails, your previous active submission stays active and is the one you can score. Your submission page says which one is active, and whether it can be scored as it stands.
Uploads with flagged imports (subprocess, socket, ctypes and the like) need you to acknowledge the warning on the submission page before they can be activated, and are read by staff before any round they enter.
8. Rounds, the final, and appeals
- You ask for a scored run; nothing runs on a timer. When you ask, the upload you asked for is scored on the ladder in a round of one person — you — and only if its smoke test has passed. Changed 2026-09-29: a run you ask for ranks nobody. With one entrant there is no round-robin to play, so what it gives you is your win-rate curve, your zero-shot point, your SAE and your faults, on the whole ladder tier — a diagnostic of your own agent. It is a private round: you read it on your upload page and in your history, and the leaderboard does not move.
- The leaderboard is made by a round the professor runs over everybody at once, and you cannot ask for one. Your agent is in it if it is your active upload and its smoke test has passed when that round starts — not when it was queued. That is the round §4's round-robin is played in and §6's standing comes out of.
- A scored run costs you one of 3 scored runs per rolling 7 days. The window is rolling, not weekly: a run drops out of it exactly seven days after you spent it, and your upload page says how many you have left and when the next one frees. A run that fails on our infrastructure is not counted against you, and neither is one staff re-ran for you — the re-run is the one that counts. A round you did not ask for costs you nothing — the professor's round, or a staff dry run you were swept into, however it is labelled. The limit is there because each run you ask for is a real round on one shared machine: 28 environments played out in full, for one entrant.
- Ladder rounds are formative, and so is every round the open benchmark scores you in.
- Rounds on fewer than 8 environments are private (staff and the entrants concerned can see them; the class leaderboard does not change). A round with a field of fewer than 8 agents is private for the same reason (§6): it runs in full and publishes no standing.
- The field freezes 48 hours before the final. After the freeze you cannot upload or change which submission is active; a practice upload for a smoke test is still allowed but cannot be activated. Your upload page shows the frozen entry — id, sha256, upload time — the moment the freeze takes effect.
- The final is published after review, not the moment it finishes. Until the host publishes it the leaderboard says final standings pending review.
- You have 7 days after the final is published to appeal. While the competition is running, the evidence you can be given is aggregate — your per-environment win/draw/fault counts without environment identities. On
finalthat is a disclosure rule: itsKis withheld, and a per-environment win count sitting next to an environment id is most of the way to inferring it. On the ladder there is nothing left to protect — those thresholds are published with the tier (§3) — and the evidence is handled the same way there because the rule turns on whether the competition is running, not on which tier you ran. When the appeal window closes the competition is marked ended and you can open your own ladder and final replays: the moves and the boards your agent saw, never the threshold, and never who was awarded a game or its disc counts — onfinalthose are what a threshold could be read back from. - A submission found to have obtained
Kby any route other than playing is invalidated by the host, after a written proposal from staff, and stays on the board struck through with the reason.
9. Your account
Today you sign in with a token a member of staff gives you. It starts eoa1., it is not a password, and there is nothing to choose or reset: paste the whole thing into the sign-in page and you are in. It is shown to staff once, when they make your account, and stored only as a hash afterwards — nobody, including the professor, can read it back. Lost it, or think somebody else has it? Ask for a new one. Issuing a new one invalidates the old one and signs out every session it opened, which is exactly what you want in both cases.
Signing in with GitHub is the door this arena is built for, and it is not switched on yet. It needs an OAuth app registered against this arena's address, which is the professor's to do. The sign-in page always shows the door that is open, so you never have to guess. The rest of this section says what changes on the day it is switched, and for everything that is yours — your handle, your uploads, your results, your session — the answer is "nothing": staff link your GitHub account to the row you already have rather than making you a new one.
With GitHub there is no password either: the arena never sees your GitHub password, and reads two things about you from GitHub — your numeric account id and your login name. Nothing else is stored: not your e-mail address, not your profile name, not your picture.
- An account can only read until staff approve it. It may read exactly what a signed-out visitor may read, plus its own account page. Downloading the starter kit, uploading an agent and asking for a scored run are granted by a member of staff, all three together — until then those pages answer "this account has not been approved yet". Your account page says whether yours has been; there is nothing to apply for.
- Anyone can make an account — once GitHub sign-in is on. Today staff make it, because a token is the only door and only staff can issue one; after the switch the same button makes an account and signs you into one.
- Your handle is chosen once and is yours. It is what the leaderboard prints and what signs every entry in the log, and it cannot be changed afterwards. Today staff set it when they make the account; after the switch you choose it the first time you sign in, with your GitHub login offered as a suggestion and nothing more.
- One GitHub account, one arena account. The GitHub account you first sign in with is the only one that can sign in as you, and it cannot be used to create a second arena account. Renaming it on GitHub is safe: you are recognised by account, not by name.
- You do not need an account to read. The leaderboard, the environment list and these documents are visible without logging in; only uploading and reading your own logs need a session.
- Lost access to your GitHub account? It cannot be recovered. Nothing here ties an open-benchmark account to a person, so there is no way to establish that the account is yours — and a request taken at its word is how an account gets taken over in the first place. Sign in with a new GitHub account and choose a new handle; the old account keeps its submissions and nobody signs in to it again.
- If GitHub is down, a session you already have keeps working. Nothing on this site asks GitHub anything once you are signed in — the arena reads your GitHub account once, at sign-in, and never again — so uploading and reading your results are unaffected. What fails is starting a new session, and the page says so rather than pretending the password was wrong. Wait, or use a browser you are still signed in on. While the token door is the one that is open, GitHub is not on the path at all and an outage there costs you nothing.
- A session lasts 14 days, or 12 hours if you tick "shared computer" on the sign-in page. On a shared or lab computer sign out when you are done; if you think a session of yours is exposed, "sign out everywhere" on your account page ends every one of them at once.
10. Where this arena differs from the paper
Documented in full in docs/benchmark-spec.md §9. The short version:
- scored environments are generated rather than the published 56, so the final tier's objective is genuinely hidden;
- no game stops early: the paper's two ten-ply conditions are replaced by two further disc-ratio thresholds, so
Kis the whole of what the final tier hides; - 20 measurement games per opponent instead of 10, because ranking a cohort needs tighter error bars than reporting a mean over seven layouts does — and 20 per pairing in the round-robin, for the same reason;
- the MCTS oracle opponent is replaced by a deterministic equivalent, because the released one draws from an RNG that cannot be seeded, and every game here must be reproducible from its move list;
- the opponent that decides the ranking is the field itself, not a fixed pool (changed 2026-09-29). The paper measures agents against fixed reference opponents so that numbers from different runs compare; a standing over a field does not compare between runs, and this arena accepts that on its ranked column in exchange for a stronger and much cheaper opponent set — the two Oracle bots it replaced cost about twelve times what the whole round-robin costs. The curve, and therefore SAE, still runs against the paper's fixed
randomopponent and still compares between runs; - SAE is scalarised, as described above, and is published rather than ranked on (§6).
One point in the paper and its released code disagree, and it is flagged rather than silently resolved: the paper describes the time-limited variants as ending after "10 moves per player", while the released code ends them after 10 plies in total. Nothing in this arena turns on which it is, because no environment here has a turn limit at all; the engine still supports one, and the differential tests hold it to the code. See docs/benchmark-spec.md §8.