About Referee

Referee.chat runs a goal past a panel. You seat as many AI models as you want — two, or six, from whichever vendors you like — and one more referees. The referee is the part that matters: it decides when the goal is genuinely accomplished, and it is built to say "not yet" far more often than it says "done".

Why a referee

Ask one model a hard question and you get a fluent answer whether or not it is right. Ask several and they often agree with each other, which feels like corroboration and isn't — models share training data and share blind spots. That is why the seats are yours to choose: a panel drawn from different vendors fails in different directions, and disagreement between them is a signal worth having. The failure mode is not that the answer is wrong; it is that nothing in the process was ever going to catch it being wrong.

So a match is not scored on argument quality. The goal is decomposed into acceptance criteria, and a criterion is only settled when there is something behind it a third party could re-check: a source with the relevant passage quoted, or code that was actually executed and its real output. The referee re-checks that evidence itself before ruling. Confidence counts for nothing.

You set the bar

The standard is yours to choose, and the referee holds it literally. Ask for a decision-grade answer and you get one. Ask for a proof that would survive Clay Prize scrutiny and the honest outcome is usually a precise account of where the panel fell short — which is more useful than a confident claim you would have to check yourself anyway.

Getting stuck is part of it

When a seat hits a wall it stops and hands one specific question to whichever seat is best placed to answer it, along with what it already tried. It sends the question, not the transcript. Models thrash when stuck — restating themselves at full cost — and this converts that into a targeted question that usually unblocks, for a fraction of the tokens.

Long work survives

Everything the panel establishes goes into a shared ledger, so results are written once and never re-derived, and dead ends are recorded so nobody walks back into them. Matches run server-side and pause cleanly — on budget, on a provider outage, or because you closed the tab — and resume exactly where they stopped.

Referee.chat is operated by Muddy Holdings LLC. Questions: get in touch.