Start free

how scoring works

A score with receipts.

Three signals build your number: what you said, the code you wrote, and how you said it. Two are measured, one is judged, and every point traces to a moment in your transcript.

40%
35%
25%
What you said
The code you wrote
How you said it

readiness = 0.40·interview + 0.35·coding + 0.25·speaking − staleness

Only surfaces you have actually tried count. The weight redistributes across them, so an untouched tab never pads your number, and going quiet for weeks pulls it back down. A formula, not a mood.

Measured, or judged. Never guessed.

Two of the three signals are computed from what you did. One is an expert judgment against your exact role. Every score names which it is, and where it came from.

SubstanceJudged

Graded against what a strong answer for your exact role must hit, only on what you were asked.

  • ownership, specifics, numbers
  • dodged questions get named

source  role rubric · graded, then validated

CodeMeasured

Run for real against hidden tests in a sandbox, while you explain your approach out loud.

  • correctness is pass or fail
  • approach probed mid-solve

source  sandbox execution · hidden tests

DeliveryMeasured

Pace, fillers, pauses, and talk time, computed straight from your speech timestamps.

  • pace · 140 to 160 wpm ideal
  • talk time · 75 to 80% you

source  speech timestamps · no model

Larpy● live · 03:12
larpy

Tell me about a time you led a project.

you · transcribing

Recording · Larpy reads as you speak
live read · this answer
filler words7 · um×2 · uh · like
pace184 wpm · rushed
hedging3 · "kind of", "i think", "i guess"
specificitylow · no metric, no "i did"
Clarity·
Structure·
·/10

Larpy: “Lead with what you owned, name the decision, and end on a measurable result.”

The number becomes a tier.

Hard bands, identical for everyone, no curve. Scores within a few points are noise; the tier is the signal. Click the legend and play the whole range.

Larpy · verdictYou vs the role
S-TIER
swe @ stripe

Larpy: “No notes. You owned every call and put a number on it. They’re drafting the offer.

94/100
substance 9.5
delivery 9.2
Roll a verdict

You can't smooth-talk it.

Our standing test: take a strong answer, strip the substance, keep the exact same confident voice. The score has to drop. If fluent nothing still scores well, our grader failed, not you.

Strong answer

We cut checkout latency 30%. I isolated the N+1 query, added an index, and measured it against a control.

88

substance.score = 88

same voice, no substance

Substance stripped, same fluency

We really improved checkout. I worked closely with the team and it went great, honestly a big win for us.

< 80

must drop ≥ 8, or the test fails

Held to the exam-industry bar.

The hardest test of any grader: does it match a human expert? We validate the same way the SAT and TOEFL are validated, and hold one bar: agree with an expert as often as two experts agree with each other.

grader agreement · 0 is a coin flip, 1 is identical

two human experts land here
0.00.600.801.0

target qwk ≥ 0.70

study in progress

What we don't claim.

Tools that claim perfection are the ones you should not trust. The honest edges:

One interview is noisy, for humans too. The readiness number across sessions is the real signal, not any single card.

Your accent is not a metric. Delivery scores pace and structure, never how you sound. A rough mic will not tank you either.

"Ahead of X% for this role" is a calibrated estimate. A model's read of where you sit, meant as a guide, not a live database of every candidate.

Code is execution-verified when the sandbox supports it. When it cannot run, we fall back to expert review, and we say so instead of dressing it up as a real run.

The fastest way to trust it is to feel it.

Run one interview and read the verdict. Every number on your card traces back to something on this page.