Your goal, argued out by a panel.
A referee decides when it's actually done.
Set a goal and the standard it has to meet, then seat as many models as you want
against each other. They research it and run code to check their own claims. A
referee re-checks the evidence itself and refuses to mark anything settled that a
hostile expert could pull apart.
Evidence, not opinions
Every seat searches the literature on arXiv, OpenAlex, Crossref and Europe PMC, reads the
sources, and run Python to check arithmetic instead of asserting it. A claim with
nothing re-checkable behind it cannot settle a criterion.
Stuck is a move, not a failure
When a seat hits a wall it stops and hands one specific question to whichever seat is
best placed to answer it — with just that question, not the whole history. Cheaper
than letting a model thrash, and it usually unblocks.
Long goals survive
Everything established goes in a shared ledger, so nothing is re-derived and nothing
is forgotten. Matches pause and resume without losing work — close the tab and come
back tomorrow.