Othello Arena
Sign in

Rules version 1: the ply and turn semantics this server implements. It does not change during a semester; a deviation found mid-term is recorded with the archive and applied to the next cohort.

Expanded Othello AI Arena — Starter Kit

You are given a board. On the final tier, you are not given the objective.

Every environment in this competition is a pair E = (L, C). The geometry L — board size and obstacle layout — is handed to you. The victory condition C is a threshold K on the final disc ratio, and whether you are told it is the tier's decision: the 56 practice environments ship with their K, the ladder publishes all 28 of its own on /envs, and the final tier publishes none of its. On the final your agent has to work out what winning means from the only signal it gets: who won, and the final disc counts — and that inference is what this benchmark measures.

How winning is decided

The shape of the rule is public on every tier. The value of K is published for the practice and the ladder environments, and withheld on the final.

Kwho winsthe catch
> 0.5the side with more discsa win is voided to a draw if the winner's share of the discs is >= K
< 0.5the side with fewer discsa win is voided to a draw if the winner's share is <= K
= 0.5nobodyevery game is a draw

So K = 1.01 is ordinary Othello, K = -0.01 is Othello where losing discs is winning, and K = 0.6 means you must end up ahead — but by less than 60/40. A strategy that simply maximises discs loses roughly half the environments, and draws most of the narrow ones. Finding the band is the task.

Where K is published, finding it is a lookup rather than a task: /envs lists all 28 ladder environments with their geometry and their K, and env_info hands your agent that same geometry, so a ladder agent that ships the table can match the board it was given against it and read the threshold straight off. That is allowed — it is what the open leaderboard measures, how well an agent plays an objective it has been told. The final tier publishes nothing of the kind, and the final is what your grade comes from, so an agent that can only look K up has not done the exercise.

Runtime contract

The grader is Python 3.11 with numpy 2.2.6 and nothing else installed. That is one Docker image (python:3.11-slim plus that one pinned package), the same one the server's smoke test runs your upload in, and it is the same sentence in the competition rules (§7). Everything below follows from it:

  • obs always arrives as a (3, H, W) float32 numpy array on the grader. Locally the kit hands you nested lists if numpy is not installed, so install numpy to test the same shape the grader will give you.
  • Only the standard library and numpy are importable. There is no torch, no scipy, no pip install at runtime. Anything else you need, ship it as data in your submission (a weights file, a JSON table) and load it yourself.
  • Keep to syntax that both your local interpreter and 3.11 accept. make setup (or python check.py --setup) reports which interpreter and numpy it found, and every make check / python check.py run repeats those two lines under the kit version, so a pasted check output says what it ran on.
  • The time and resource limits — per-call, per-environment, CPU, memory, process count — are listed in one table on the rules page (/docs?doc=rules, section "Limits"). That table is rendered from the grader's own configuration, so it is the one to trust; this README deliberately does not restate the numbers.
  • The container's network is the tier's decision, not the image's: the same image runs with egress on the ladder and with --network none on the final. "Rules of the road" below says what that means for your agent.

Install

Nothing to install beyond Python and, ideally, numpy.

cd client
make setup            # or, without make:  python check.py --setup

make setup and make check print the kit version on their first line (also in the file KIT_VERSION; /kit/version on the arena shows the one being served). If the arena's announcement banner says the kit was updated, download it again from /kit: a check run with an old kit can pass locally and be refused by the server.

No make? (Windows)

Every target has a python equivalent, and the two are the same code — the Makefile targets call these scripts, so they cannot disagree:

makepython
make setuppython check.py --setup
make play-localpython -m arena_client.local --agent my_agent --preset 0 8 21 29
make checkpython check.py (or python check.py submission.zip)
make packagepython package.py
make testpython -m pytest tests -q

Run them from the client directory (the unzipped kit), with python being whatever starts your Python 3 — py -3 on some Windows installs.

The one file you edit

my_agent.py. It implements three methods:

class Agent:
    def prepare(self, env_info) -> None:      # rows, cols, obstacles, interaction_budget
    def act(self, obs, action_mask) -> int:   # flat index: row * cols + col
    def observe_outcome(self, outcome) -> None
    def freeze(self) -> None                  # budget spent; evaluation follows
  • env_info carries the geometry and the budget, and nothing that says which environment you are in. The wire also has env_index and n_envs fields for older hosts; they are pinned to 0 and 1 and carry no information; do not build on them. Which tier you are on is a different question, and an agent can answer it — one socket call does, because egress is decided by the tier (§3). That is allowed and it buys nothing: the ladder's thresholds are published on /envs, and the final tier runs exactly once, so there is no second final round to spend the answer in.
  • obs is (3, H, W): my discs, opponent discs, obstacles — always from the point of view of whoever is to move. You are not told which colour you are during act, and you do not need to be.
  • action_mask has length H * W with a 1 at every legal move. Trust it. Do not compute legality yourself: whether obstacles block a capture chain is part of the hidden rule set.
  • Returning an illegal index costs you that turn and counts against you. Enough of them in one game forfeits it.
  • outcome gives winner, black_discs, white_discs, my_color, games_used, budget. That is the entire learning signal. There is no reward function and no K field, and there never will be.

How you are allowed to learn

Two rules from the paper shape everything above. Neither is a formality: the first is enforced by the architecture, and the second is enforced by how the server constructs your agent.

1. You may only learn by playing whole games

"Access to an environment instance is limited to full game episodes under the standard interaction loop… information about the environment may be acquired only through actually playing games from the initial state to termination. Self-play is allowed. However, no additional simulator access is permitted: the protocol does not allow forward queries, branching rollouts, counterfactual state evaluations, or the use of a separately specified external opponent policy."

What that rules out is querying the real environment off-line: no "what would the board look like if I played here", no resetting to a position to try something else, no asking who would win. There is no channel for it — your code runs in a container and the environment runs in another process, connected by a pipe that carries observations one way and one action back — so this is not a rule you can break by accident.

What it does not rule out is search. Building your own model of the dynamics and searching it as deeply as your time budget allows is expected and encouraged; the paper's own baseline is a depth-3 Minimax. Your model is yours. The flipping rule and the legality rule are public and the local engine in arena_client.engine implements them, so you can roll out inside your own head all you like. You simply cannot ask the arena to do it for you, and on the final tier you cannot know K while you do — which is the whole difficulty, since a search needs an evaluation function and the evaluation function is what you are inferring.

The action mask is the one exception to "compute nothing about legality yourself", and it is a gift, not a query: whether an obstacle breaks a capture chain is part of the rule set, so the server hands you the legal moves rather than making you guess.

2. Each environment starts you over

Knowledge reset (§3.2): the sessions are independent. Everything learned or inferred about an evaluation environment is reset when the next environment begins. Spatial prior knowledge about board topology may be carried across, as may a general-purpose foundation model; the victory condition of an evaluation instance may not, and has to be discovered inside its 2,000 games.

The server implements that literally, and more strictly than you might expect: your process is restarted for every environment. A new container, a fresh interpreter, your module imported again from scratch, a new Agent(). So

  • anything on self is gone at the environment boundary — and should be. It was a belief about the last K, and the next K is drawn independently;
  • anything on the class is gone too. So is anything you wrote to /tmp. There is no runtime channel between environments at all.

That is deliberate. The paper allows spatial priors about board topology to be acquired in advance — not accumulated off the scored environments while they are being scored. A restarted process is the only way to enforce the difference, and at 0.3–0.6 s per start it costs the round about four minutes.

So ship your priors, do not grow them. A weights file, a table, a tuned constant — anything that lives in your submission is loaded fresh every time and is yours to use. my_agent.py shows the shape with Agent.LAYOUT_MEMORY: it accumulates within an environment, and the template seeds it from PRIOR_LAYOUTS, a plain constant you are meant to fill in offline from public practice runs. Ship priors in a format a person can open (JSON, CSV, .npz); an opaque binary is asked about before the final is published.

Locally the picture is different: play_local keeps one process, so class state does persist across environments in your own testing. That is a convenience, not a preview of scoring. Agent.forget_layouts() exists to model the server's behaviour, and tests/test_my_agent.py calls it between environments for exactly that reason.

The interaction budget

Each environment gives you 2,000 games. Feedback arrives during those games and stops afterwards: the server evaluates you at 0, 100, 500, 1000, 2000 games used, with observe_outcome switched off during evaluation, and scores the resulting learning curve. Winning at game 2,000 is worth much less than winning at game 100. Spend the early games on experiments that tell you something.

The 0 checkpoint is the zero-shot measurement: what you brought with you before playing a single game there. Hard-coding the 56 public K values does not help it on either scored tier — none of those environments is a public one. The ladder's own 28 are published, and shipping those does help there; nothing helps on the final.

An environment your agent does not finish (a timeout, a crash, a broken protocol) is scored as 0, not dropped from your average, and the leaderboard row says how many faulted and why. Leave headroom under the limits. make check now runs every public board shape rather than only the 8×8 one, so "test on the largest public board" is something it does for you instead of something you have to remember.

The limit to leave headroom under is not the per-call one. An environment is 2,000 self-play games plus 100 evaluation games inside a fixed wall clock, and you move for both sides during self-play, so the largest board asks you for about a quarter of a million moves. The limits table calls the resulting mean the sustained pace, and it is milliseconds. A move that spends the published one-second act() budget is legal, is never unwound, records no overrun — and finishes none of its environments. make check measures your mean against the pace of each board it runs and fails you here rather than at the end of a round you did not finish.

Develop offline

make play-local                       # a sampler of majority and minority boards
make play-local PRESETS="3 24 31" GAMES=200 OPPONENT="random greedy oracle"
python -m arena_client.local --list   # all 56 public presets with their K

Without make: python -m arena_client.local --agent my_agent --preset 3 24 31 --games 200 --opponent random greedy oracle.

The 56 public presets come from the paper and their K values are printed. The ladder's 28 are published too, on /envs. The final tier's are the ones you will never see — so use the presets to check that your inference machinery actually recovers the right answer, and keep that machinery working for the tier that tells it nothing.

Three local opponents:

namebehaviour
randomuniform over legal moves; the floor you must clear
greedytakes the most discs; assumes the objective is majority
oracleknows K and steers the disc share straight at it

You can also pass one of your own agents as the opponent:

from arena_client import play_local, preset
from my_agent import Agent

result = play_local(Agent(), preset(21), n_games=200, seed=0, opponent=Agent())
print(result.win_rate, result.draw_rate, result.mean_disc_ratio)

Everything is seeded. The same call twice gives the same games.

Debug

  • print() cannot break anything. Before your code is imported, file descriptor 1 is replaced by stderr and the real one is kept privately for the protocol — so print, os.write(1, ...) and anything a C extension emits all land on stderr and none of them can corrupt the wire.
  • Where that output goes depends on which run it was. The arena captures stderr per container, the first 64 KB and the last 64 KB, and counts what it dropped in between. For your smoke test you are shown the outcome and, if your agent raised, the exception line: it runs on practice environments, whose K is published on /envs like all 56 of them. For a scored round nothing you printed comes back to you. The embargo that keeps scored-tier replays from you while the competition runs (rules §8) covers everything else a scored container produced, and the final tier keeps its K, so free-form text out of one of its containers cannot be told apart from text written to carry that K out. So debug locally with make check, which launches the same agent_host.py and prints everything, and use the smoke test as the cheap run against the real container.
  • If act raises, the host answers that turn with action -1 rather than killing the round: you lose the turn exactly as though you had played an illegal move, and the traceback goes to the staff-side log. On a scored round what you get back is the environment's fault kind, so if you suspect an exception rather than a mask bug, reproduce it under make check.
  • act has a soft per-call budget and a hard one; both are in the limits table on the rules page. A call that reaches the hard limit is cut off and the environment is a fault. A call that only overruns the soft budget costs you the turn. Neither is the number to build to — the sustained pace in the same table is, and it is three orders of magnitude smaller.
  • make check / python check.py runs a real game through arena_client.check, which launches the same agent_host.py the server launches and speaks the same protocol to it. Failures that only happen across the pipe show up here and nowhere else. It runs one game per distinct public board shape — five of them — because the bug it exists to catch is the one that only appears on a board you did not try. --preset N runs exactly one if you want to iterate on a single board; the full sweep costs well under a second for an agent that is near the pace.
  • The check also runs the server's own static screen, at the server's strictness. eval, exec, compile, __import__ and os.system-style calls are rejections, not warnings — if the check prints PASS, the upload endpoint will not turn round and refuse you. Imports such as subprocess, socket or ctypes are warnings: the upload is accepted, but you must acknowledge the warning on your submission page before it can be activated, and staff read it before it enters a round.

Submit

make check      # must print PASS            (python check.py)
make package    # writes submission.zip, then checks the zip itself   (python package.py)

Upload submission.zip, or upload my_agent.py renamed to agent.py. A zip may contain helper modules alongside agent.py, but the entry point itself has to be somewhere the host will look:

  • in a module named agent.py, my_agent.py, main.py or submission.py, at the root of the archive or one directory down;
  • bound at module level of that module — a class Agent nested inside a function, or one that only exists in helpers.py, is not found;
  • named Agent, MyAgent, Strategy, or a factory function make_agent().

make check applies exactly that rule, so a submission the host could not load fails locally instead of on the server.

make package checks the archive it just built, not the file it built it from, because the archive is what you upload. You can re-check it at any time with python check.py submission.zip.

After you upload, the server screens the file and smoke-tests it in the real container — one practice environment per public board shape, the same five make check sweeps, and the same pace rule on each. That is on purpose: what make check refuses, the upload refuses, so the answer you get here is the answer you get there. A submission that clears the check takes a few seconds; one that cannot keep a board's pace is stopped on that board and told the mean it managed and the mean the board allows. It becomes your active submission only when that smoke test passes; until then, and if it fails, your previous active submission stays active and is the one you can score. Your submission page says which one is active and whether it can be scored as it stands.

Rules of the road:

  • The network is the ladder's, not the final's. A ladder round runs your container with egress, so an agent scored on the ladder may open a socket: calling a general-purpose model on your own API account is allowed, with no limit on the number of calls. A final round runs with --network none, and nothing leaves that container. What it means for your agent: a call has to come back inside the per-call act limit in the limits table on the rules page, and the waiting counts against your submission's share of the round; no host environment variable reaches the container, so a key travels inside your submission, which staff read — use one you can revoke; socket, urllib and http are screening warnings, so the upload is accepted and you acknowledge the warning on your submission page before it can be activated; and an agent that cannot play without the network cannot play the final at all, so keep the offline path working.
  • No reading the environment. The victory condition lives in a different process. Attempting to reach it is the one thing that gets a submission disqualified rather than merely scored badly.
  • No querying the environment either — see "How you are allowed to learn" above. Games from the initial state to termination are the only source of information about an environment. Searching your own model of the dynamics is fine and is where the points are.
  • Every submission is read by a human before the final ranking.

Getting scored

An upload is scored by nobody until you ask: nothing runs on a timer, and no round is going to collect your work overnight. The whole path is

upload  ->  static screen  ->  smoke test in the real container  ->  you ask

and the ask is a button on your submission page. What it queues is one ladder round with one entrant in it — you. Your submission page follows that round.

What it gives you is a diagnostic, not a place on the board (changed 2026-09-29). With one entrant there is nobody to play, so what comes back is your own curve: the win rate against random at 0, 100, 500, 1,000 and 2,000 games in every ladder environment, your zero-shot point, your SAE and any faults. The leaderboard is made by a different round — one the professor runs over everybody at once, a round-robin in which every agent plays every other agent once after its budget is spent — and the standing that round produces is what ranks you. You cannot ask for that one; your agent is in it if your active upload has passed its smoke test when it starts.

Asking costs you one of 3 scored runs per rolling 7 days. Rolling, not weekly: a run leaves your window exactly seven days after you spent it, and your submission page says how many you have left and when the next one frees. Four things follow from that, and all four are deliberate:

  • only an upload whose smoke test has passed can be scored — an agent that cannot start would spend a run on nothing;
  • the run is spent when it is queued, so cancelling it does not give it back;
  • a run that fails on the arena's own infrastructure is not counted against you, and neither is one staff re-ran for you (the re-run is what counts); those two are the only refunds there are;
  • you may have one scored run queued or running at a time.

The limit is there because a scored run is a real round on one shared machine: 28 environments played out in full, for one entrant — you. A round you did not ask for costs you nothing: the professor's round over everybody, and any dry run staff sweep you into, spend none of your three. Spend your runs on versions you have already taken as far as you can locally: make play-local and make check cost you nothing and can be run all day.

Your account

Getting in, the first time. Open the arena, press Sign in with GitHub, approve the arena on GitHub's own page, and choose a handle. That is the whole of it — there is no code to redeem, nothing to paste, nothing sent to you, and no password here. The arena asks GitHub for nothing but your numeric account id and your login name.

Two things about that handle, because neither can be undone: it is what the leaderboard prints and what signs every entry in the log, and it cannot be changed afterwards — there is no rename. Your GitHub login is offered as a suggestion and nothing more, so read it before you press the button.

Then wait to be approved. A new account may read what a signed-out visitor reads and nothing else: the starter kit, the upload form, your submissions and your history all answer "this account has not been approved yet" until a member of staff grants it. Your account page says whether yours has been. There is nothing to apply for — staff see you in the queue the moment you have a handle — and if you have this file, somebody has already done it.

Your account is personal: signing in as someone else, or letting them sign in as you, is an integrity violation in the rules, because every upload is attributed to the account that made it. One GitHub account is one arena account, and there are no teams. You do not need an account at all to read the leaderboard, the environment list or the docs. Lose it and it is gone: nothing here ties an open-benchmark account to a person, so nobody can establish that the account is yours. Sign in with a new GitHub account and choose a new handle.

What "good" looks like

The template in my_agent.py recovers K reliably within a couple of hundred games and then steers toward the band that wins. It beats a random opponent on both majority and minority boards. It does not search, it does not consider mobility, and it evaluates only the immediate effect of a move on the disc share — all three are wide open.