Can you still out-think AI?
One question from a subject you pick. Fifty words each — you and AI. Then a judge, also AI, scores both answers without being told which is which.
Free · One round a day · Account needed
one of them is yours ↑Five steps, then you’re done
- 01Pick a subject and a difficulty.
- 02You get one question — a real one, asked by a real person on Stack Exchange.
- 03Write your answer. Fifty words, hard cap.
- 04AI gets the same question and the same fifty words.
- 05A judge scores both against a fixed rubric, without being told whose is whose.
Pick your ground
- Logic & Philosophy
- Conceptual CS
- Finance
- Legal
- Medicine
- Biology
- Engineering
- JavaScript
- Python
Each one has its own rubric. Finance rewards spotting the real problem; Logic rewards the reasoning holding up. Your rating is per subject, so being strong in one doesn’t flatter the rest.
Legal and Medicine run on rubrics we haven’t signed off yet. They’re playable — treat their scores as work in progress.
What the judge can’t do
It can’t see whose answer is whose.
Both answers go in blind, in a random order, re-shuffled on every re-run.
It can’t mark its own homework.
The model you played is barred from refereeing your round.
It can’t reward you for writing more.
Fifty words each. You’re stopped at fifty; AI gets cut off at fifty, mid-sentence if it has to. And the judge is told not to score length.
It can’t invent a score.
Every dimension is checked in code against its own ceiling. A verdict that doesn’t hold up is re-run, then set aside for a human.
It can’t wave through a made-up fact.
A flagged hallucination forces that answer’s factual score to zero. That happens in code, not by the judge’s good manners. Vague, could-apply-to-anything claims lose eight points.
Same rubric for both of you. Four scored dimensions, five in Medicine, weighted per subject and always adding to one hundred.
There’s one way to find out.
It doesn’t go easy on you
- Your rating moves both ways. It’s Glicko-2, and a bad round costs you.
- Picking Beginner doesn’t get you an easy opponent. Difficulty chooses the question, never who you’re up against.
- Your opponent is a random draw. You can’t hunt for a weak one.
- Close isn’t a win. Inside two and a half points either way it’s a tie, and the question can come back.
One round a day
That’s the whole allowance, and it’s deliberate. One question, one answer, then it’s out of your hands until midnight your time. No second go today. If a round breaks or the judge can’t call it, the day is handed back to you.
Playing is free. We won’t sell you extra rounds, and there’s no tier that plays more. Five rated rounds and you’re eligible for the leaderboards, which are only visible to players who are signed in.
There’s no losing screen
You either out-thought it, scored less, or it was too close to call. That isn’t softness. A one-word verdict on a fifty-word answer teaches you nothing. What’s worth having is the gap, and where it came from.
out-thought, or scored less. never beatenSo — can you?
One question. Fifty words. One round a day.