Othello Arena
Sign in

You are given a board. On the final tier, you are not given the objective.

Every environment here is a pair E = (L, C): the geometry L — board size and obstacles — is handed to your agent, and the victory condition C is a threshold K on the final disc ratio. K is the part this benchmark hides, and the final tier hides it — working out what winning means there is the task, and your agent has 2,000 games per environment and one signal to do it with: who won, and the final disc counts. The ladder publishes its K with the tier, so the open leaderboard measures how well an agent plays an objective it has been told — a different question from the paper's, and a real one.

Sign in with GitHub Read the rules

deadlinenot yet announced
practice environments56K disclosed
interaction budget2,000games per environment, enforced by the server
measured at0, 100, 500, 1000, 2000games
isolationone containerper environment, no environment definition inside it; network on the ladder, none on the final

1. Get the kit

One file to edit: my_agent.py. The practice tier ships with it, with K disclosed, so your inference has somewhere to be checked.

Download the starter kit Read the agent API

2. Check it locally

make check runs exactly the gate this server runs, and make play-local plays the 56 practice environments.

Browse the environments

3. Upload it

A single agent.py, or a .zip with agent.py at its root. Screening is immediate; a smoke test follows within a minute.

Sign in with GitHub

4. Ask for a diagnostic, and be ranked by a round

Nothing runs on a timer. A run you ask for from your submission page is a private diagnostic — your win-rate curve, your SAE, ranked against nobody — and it spends one of your three per rolling seven days. The board is made by a round the professor runs over the whole field at once: after every agent has spent its interaction budget it plays every other agent once, and that standing is the ranking. SAE is published beside it and is the number that means the same thing in every round; draw rate is its own column and is never folded into a score.

Open the leaderboard

Last ladder round

#agentstandingSAEdraw
1 deepening-dave 0.592 0.360 46.4%
2 netty-nadia 0.584 0.624 11.6%
3 minimax-mina 0.559 0.432 31.2%
4 mcts-milo 0.553 0.500 20.5%
5 prober-pia 0.549 0.573 16.1%

See the full leaderboard

Why the scored tiers are fresh

The scored tiers are freshly generated, not the paper's published 56 presets, because those presets and their K values are on GitHub. Ladder and final replays stay host-only while the competition is live — one embargo for both tiers, because on the final a move sequence plus its result narrows K.