An assessment an AI can pass — but can't defend.
Candidates solve a real engineering ticket, in any AI tool they like, then defend their fix out loud. Every point on the score traces back to evidence you can read or listen to yourself.
Follow-up answers were thinner than the PR itself.
Five stages, including the one that can't be faked
Get a real ticket
An ambiguous ticket from a real company codebase — a symptom to chase, not a puzzle to solve.
Fix it, then write it up
Use any AI tool you like. We don't measure typing — we measure judgment. Open a PR and describe your fix.
Three independent AI passes review it
The diff, the write-up, and the reasoning are each scored separately, then reconciled.
Answer two follow-up questions
Fifteen minutes, no notice of what's coming. This tests whether you understand your own fix.
Defend it out loud
A spoken defence of your own change — the one part of the process an AI answer can't fake.
Execution is only 10%
Shipping code that works is table stakes. Understanding why the bug existed is what the score actually weighs.
Strong code and weak defence look different — on purpose.
When the PR review score and the final score diverge, that gap is surfaced as its own signal — not buried in the total. A wide gap usually means the code was solid but the candidate couldn't explain it under follow-up. The transcript is one click away.
Every flag is labelled advisory. Nothing is auto-rejected — the final decision rests with the hiring team, with the evidence attached.
Amber gap chip · “strong code, weak defence — review the transcript”
What we promise every candidate
- —A tooling failure never costs you points.
- —Every deduction is evidence-linked — nothing is subtracted without a reason you can see.
- —Your AI-tool declaration never changes your score.
- —Flags are advisory. A human always makes the final call.
- —Audio is never stored — only the transcript is kept.