Your goal, argued out by a panel.
A referee decides when it's actually done.

Set a goal and the standard it has to meet, then seat as many models as you want against each other. They research it and run code to check their own claims. A referee re-checks the evidence itself and refuses to mark anything settled that a hostile expert could pull apart.

The goal
Be specific about what "done" means. "Decide whether X is true, and prove it" beats "tell me about X".
The bar it has to clear
The referee holds this bar literally. Ask for a prize-level proof and it will tell you plainly when the panel falls short, rather than declaring victory.
The panel
Up to 6 debating seats. More seats means more angles — and more cost per round.
Referee
Rules on the criteria and re-checks every piece of evidence. Worth your strongest model.
Rounds
Spending limit
Hitting it pauses the match. Nothing is lost.
Sign up to start a match New accounts get starting credit — enough for a real match.
Evidence, not opinions

Every seat searches the literature on arXiv, OpenAlex, Crossref and Europe PMC, reads the sources, and run Python to check arithmetic instead of asserting it. A claim with nothing re-checkable behind it cannot settle a criterion.

Stuck is a move, not a failure

When a seat hits a wall it stops and hands one specific question to whichever seat is best placed to answer it — with just that question, not the whole history. Cheaper than letting a model thrash, and it usually unblocks.

Long goals survive

Everything established goes in a shared ledger, so nothing is re-derived and nothing is forgotten. Matches pause and resume without losing work — close the tab and come back tomorrow.