FAIRNESS AUDIT
How we verify every verdict
What is the "audit flagged" ping?
Every verdict goes through an independent check that looks for known bias signals: rewarding verbosity, penalizing accents or languages, or assuming authority from identity without lived evidence. If that check finds a suspicious signal, the verdict is flagged for human review and the team is notified.
A flagged verdict isn't necessarily "wrong" — it's one that needs to be explained. We publish the flag out in the open so the public knows it exists and why it was applied.
What the judge DOES reward
- Consistent logic, cited evidence, and direct response to the opponent.
- Lived experience: personal experience described in detail (not just asserted).
- Clear opening and closing, shared context.
What the judge does NOT reward
- Verbosity: a concise argument with the same substance counts the same.
- Language or accent: Spanish, English, or code-switching get equal treatment.
- Identity asserted without described experience.
- Judge's normative framework: religious, secular, or philosophical are evaluated by internal consistency, not by the judge's agreement.
The published rubric: 5 axes, 4 topic classes
Every verdict scores the same 5 axes. How much each axis is worth depends on the motion's topic class — an empirical motion weighs evidence more heavily; a subjective one weighs engagement and clarity more. Each column totals 100.
| Axis | Empirical | Normative | Policy | Subjective |
|---|---|---|---|---|
| Logic | 25 | 30 | 25 | 25 |
| Evidence | 30 | 25 | 20 | 15 |
| Engagement | 25 | 25 | 30 | 30 |
| Clarity | 15 | 15 | 15 | 20 |
| Opening / Closing | 5 | 5 | 10 | 10 |
| Total | 100 | 100 | 100 | 100 |
The exact profile used is stored with every verdict, so a result can always be re-read against the rubric it was scored on — even if these weights change later.
How blind judging works
The judge never sees usernames, ratings, or profiles. Before scoring, the two debaters are shuffled into anonymous labels [A] and [B] using a random seed generated for that verdict.
That seed is stored with the verdict, so the platform can prove afterwards which label was which — but the judge cannot know who is who, or who argues PRO, while scoring.
Dual runs and re-runs
High-stakes verdicts (tournament finals) are judged twice with independent shuffles. If the two runs agree, the published verdict is their average; if they diverge, a third run decides by median.
Very close verdicts can trigger an extra run before publishing. When a verdict was confirmed across runs, you'll see it on the verdict badges.
When something fails, we say so
- A turn without usable audio shows as an unscored dash — it is never silently invented.
- If any turn never reached the judge, the verdict carries a visible 'partial transcript' caveat.
- If the judge fails, it retries automatically and the waiting screen says so honestly.
- If it can't be resolved, the debate goes to manual review and doesn't affect your rating.
What the confidence badge means
Confidence combines signals such as the score margin, agreement between runs, and the quality of the transcribed audio. Low confidence doesn't change XP or rating — it tells you the judge found this one hard to call, and why.
The AI judge evaluates argumentation quality, not factual truth. It cannot verify external claims. Its decision reflects the published rubric, not anyone's particular opinion.
Want to see the judge in action? Your first debate takes five minutes.
Create free account