You are given a board. On the final tier, you are not given the objective.
Every environment here is a pair E = (L, C): the geometry L —
board size and obstacles — is handed to your agent, and the victory condition
C is a threshold K on the final disc ratio. K is the
part this benchmark hides, and the final tier hides it — working out what winning means
there is the task, and your agent has 2,000 games per environment and one signal to do it
with: who won, and the final disc counts. The ladder publishes its K with the
tier, so the open leaderboard measures how well an agent plays an objective it has been told
— a different question from the paper's, and a real one.
K disclosed1. Get the kit
One file to edit: my_agent.py. The practice tier ships with it, with
K disclosed, so your inference has somewhere to be checked.
2. Check it locally
make check runs exactly the gate this server runs, and
make play-local plays the 56 practice environments.
3. Upload it
A single agent.py, or a .zip with agent.py at its
root. Screening is immediate; a smoke test follows within a minute.
4. Ask for a diagnostic, and be ranked by a round
Nothing runs on a timer. A run you ask for from your submission page is a private diagnostic — your win-rate curve, your SAE, ranked against nobody — and it spends one of your three per rolling seven days. The board is made by a round the professor runs over the whole field at once: after every agent has spent its interaction budget it plays every other agent once, and that standing is the ranking. SAE is published beside it and is the number that means the same thing in every round; draw rate is its own column and is never folded into a score.
Last ladder round
| # | agent | standing | SAE | draw |
|---|---|---|---|---|
| 1 | deepening-dave | 0.592 | 0.360 | 46.4% |
| 2 | netty-nadia | 0.584 | 0.624 | 11.6% |
| 3 | minimax-mina | 0.559 | 0.432 | 31.2% |
| 4 | mcts-milo | 0.553 | 0.500 | 20.5% |
| 5 | prober-pia | 0.549 | 0.573 | 16.1% |
Why the scored tiers are fresh
The scored tiers are freshly generated, not the paper's published 56 presets, because those
presets and their K values are on GitHub. Ladder and final replays stay host-only
while the competition is live — one embargo for both tiers, because on the final a move
sequence plus its result narrows K.