Skip to content
outthinkAILog in
You vs AI · 50 words each

Can you still out-think AI?

One question from a subject you pick. Fifty words each — you and AI. Then a judge, also AI, scores both answers without being told which is which.

Free · One round a day · Account needed

one of them is yours ↑
The whole game

Five steps, then you’re done

  1. 01Pick a subject and a difficulty.
  2. 02You get one question — a real one, asked by a real person on Stack Exchange.
  3. 03Write your answer. Fifty words, hard cap.
  4. 04AI gets the same question and the same fifty words.
  5. 05A judge scores both against a fixed rubric, without being told whose is whose.
9 subjects

Pick your ground

  • Logic & Philosophy
  • Conceptual CS
  • Finance
  • Legal
  • Medicine
  • Biology
  • Engineering
  • JavaScript
  • Python

Each one has its own rubric. Finance rewards spotting the real problem; Logic rewards the reasoning holding up. Your rating is per subject, so being strong in one doesn’t flatter the rest.

Legal and Medicine run on rubrics we haven’t signed off yet. They’re playable — treat their scores as work in progress.

The judging

What the judge can’t do

  • It can’t see whose answer is whose.

    Both answers go in blind, in a random order, re-shuffled on every re-run.

  • It can’t mark its own homework.

    The model you played is barred from refereeing your round.

  • It can’t reward you for writing more.

    Fifty words each. You’re stopped at fifty; AI gets cut off at fifty, mid-sentence if it has to. And the judge is told not to score length.

  • It can’t invent a score.

    Every dimension is checked in code against its own ceiling. A verdict that doesn’t hold up is re-run, then set aside for a human.

  • It can’t wave through a made-up fact.

    A flagged hallucination forces that answer’s factual score to zero. That happens in code, not by the judge’s good manners. Vague, could-apply-to-anything claims lose eight points.

Same rubric for both of you. Four scored dimensions, five in Medicine, weighted per subject and always adding to one hundred.

There’s one way to find out.

The catch

It doesn’t go easy on you

  • Your rating moves both ways. It’s Glicko-2, and a bad round costs you.
  • Picking Beginner doesn’t get you an easy opponent. Difficulty chooses the question, never who you’re up against.
  • Your opponent is a random draw. You can’t hunt for a weak one.
  • Close isn’t a win. Inside two and a half points either way it’s a tie, and the question can come back.
The pace

One round a day

That’s the whole allowance, and it’s deliberate. One question, one answer, then it’s out of your hands until midnight your time. No second go today. If a round breaks or the judge can’t call it, the day is handed back to you.

Playing is free. We won’t sell you extra rounds, and there’s no tier that plays more. Five rated rounds and you’re eligible for the leaderboards, which are only visible to players who are signed in.

One more thing

There’s no losing screen

You either out-thought it, scored less, or it was too close to call. That isn’t softness. A one-word verdict on a fifty-word answer teaches you nothing. What’s worth having is the gap, and where it came from.

out-thought, or scored less. never beaten

So — can you?

One question. Fifty words. One round a day.

Or help us decide what to build next.