J! Royal Rumble · The Handbook

Thirty players. Three start. One survives.

Jeopardy! clues, Royal Rumble entrances, and a scoreboard where every dollar you win comes out of somebody else's pocket. The last player standing wins, and nothing else counts as a result. Live at j-royal-rumble.net.

Part I is the rules. Parts II–IV are why the rules are shaped this way — the format was tuned by simulation, thousands of matches per question, and this handbook keeps its wrong answers on the page, marked as overturned. A design document that quietly deletes its mistakes is less useful than one that shows the correction.

Part I

The Rules

1 · The shape of a match

The board holds six categories, five clues each, worth $100–$500 by row. The host decides where the clues come from — the televised archive, house-written boards, or a mix — and when a category runs dry, a fresh one takes its place. The board never empties. The host can also reroll a category that isn't landing; three rerolls retires it for good.

Everyone draws an entry number before the match. Three players start in the ring; everyone else waits in the queue and enters one at a time, on a fixed clock. You'll know your moment is coming: your buzzer counts down your final three clues in the queue — and your entrance music plays as you walk in.

The stake is the same for everybody at face value, but it rides the overtime multiplier: walk in at 4× and you walk in with four times the money. This was not the original design and it was not a guess. In a live match one player entered at clue 150 into 2× overtime, lasted six clues without winning a single race, was revived at half stake into 4× where the top row paid $2,000, and was gone after one more. He never had a hand to play. Across 3,000 simulated matches, entrants arriving at 2× or above who died within three clues fell from 58% to 9% once the stake scaled.

The match ends when the queue is empty and one player is left standing. There is no second place.

2 · Scoring

Correct answer. Every other player in the ring pays you the clue's value. Four opponents on a $400 clue means you gain $1,600 and each of them drops $400. Heads-up, that same clue is worth exactly $400. The early game is where the big money is — that's deliberate.

Wrong answer. You lose the clue's value and you're locked out of that clue for good.

The re-toss. A missed clue goes straight back up as a fresh buzzer race for everyone still eligible. If somebody else converts it, you pay them like everyone else — so one bad guess can cost you twice: once for the miss, once for their payday.

Nobody answers. The whole ring loses half the clue's value. Passing is not free.

3 · Elimination

Drop below zero and you're out. Exactly zero survives. If one answer knocks out several players at once, they all go together, and the answerer is credited a toss out for each. If a clue somehow eliminates the entire ring, the highest score survives.

Arcade and Tournament. A match is one shape or the other, named on every screen throughout. Arcade has the rule below switched on; because a player on the way back is ranked with time taken off their buzz, buzz order stops being buzz speed, so the room is shown places rather than times. Each player still sees their own reaction time on their own device and nobody else's. Tournament has it off: the fastest press wins every race and the times are public, because there they mean exactly what they look like. The match record keeps every real time in both.

One foot on the floor. With one exception: if you would go out having taken fewer than three clues, you don't go out at all. You stay up on half a starting stake, and for the next 40 races your press is ranked at the bottom of the 60 ms band it lands in (comebackBand; 155 ms counts as 120), so the edge is worth one band at most and a press under 60 ms is left alone — it was a flat 70% off until 0.101.0, and the reason it changed is in Part IV. Once each, automatic, nothing to declare. The gate is the design — it catches somebody flattened before they ever got going, and a player who was in the match and lost it goes out properly. Because no elimination is recorded, a bounty on that player doesn't pay and a revival life isn't spent. The edge is counted in races rather than clues so a run of stumpers doesn't burn it, and the recorded reaction time is always the real one; the discount decides who takes the clue, it does not rewrite what happened.

4 · The ceiling

There's a maximum score, sized to the field — bigger matches get more headroom. Anything you win above it is clipped, and it never sits below what a new entrant starts with, so nobody walks in already capped.

During regulation the ceiling holds still. A big early lead buys you room, not immunity — and once overtime opens, the ceiling starts to fall (see rule 6). An earlier version of the game had it falling all match; Part II records why that was wrong.

5 · Clearing the field

Knock out every other player in the ring while the queue still has people in it, and you've cleared the field. You collect a bonus equal to the value of every clue left on the board, the board is replaced entirely, and two fresh players enter instead of one. It's rare. It should be.

6 · Overtime

Overtime opens when the queue empties — or, if the match stalls badly enough, even before that (a stall on the order of 48 clues will trigger it with players still queued). Since 0.102.0 it waits one entry interval after the last arrival first (overtimeEntryGrace, on by default): the final entrant is the one person in the room who has played nothing, and the logs say everybody plays below themselves on entry — a 244 ms median press on the first live clue against 204 ms in warm-up, about ten clues to come back down — so opening overtime on their heels charged them double for it. The stall path ignores the grace. The analysis chat measured it at 1,500 matches per field size: back-half win rate 57.8% to 53.5% at six players and 63.5% to 57.3% at ten, median length 69 to 80 clues at six, and inside the noise above sixteen players, where the interval is short. Then the stakes start climbing:

7 · Buzzing, and warming up

The host reads the clue, then arms the buzzers. Your buzzer lights up, the clock starts, fastest reaction wins. One buzz per player per race. Jump the lights and you're locked out for a quarter of a second — even if the buzzers go live during your lockout. There is no edge in guessing early.

Your buzzer works before you're in. Queued or eliminated, you can buzz every clue. Warm-up buzzes never touch the match record — no phantom attempts on your stat line — but they clock your reaction time against the live field, so by the time your number is called you already know the rhythm.

Timing is compensated for audio lag so everyone races on equal footing, measured on your own device to a tenth of a millisecond. Times well under 150ms are normal here, not a glitch: in recorded live matches the human median is about 150ms, and nearly half of all winning buzzes come in under it. Good players don't react to the lights — they learn the rhythm. A perfectly timed buzz reads 0.0.

8 · If your connection drops

The match doesn't wait. Clues keep coming, your score keeps moving, and you can be eliminated while offline. Harsh — but the alternative is that pulling your ethernet cable becomes a strategy. When you reconnect, you're restored exactly as you were: same score, same entry number, same place in the queue.

9 · What gets tracked

At the end of the match, everyone gets sortable, shareable standings: clues survived, toss outs, points drained, peak score, correct and missed answers, buzzer attempts, races won, win rate, average and best reaction times, and the match's single fastest buzz. The winner is the only result. The rest is an argument for the group chat.

10 · Advanced mechanics

These are off by default. The host turns them on before the match and announces which are live. They need a keyboard, so expect them disabled for a mostly-mobile field.

🪜 TOP ROPE — Press R between clues (never mid-clue, no peeking). The next clue is worth double to you, in both directions: twice the gain if you take it, twice the loss if you miss or somebody else takes it. Everyone else plays it at face value. Top-rope winnings ignore the ceiling — the one way to score past the cap. Then five clues before you can climb again, or anyone who decides doubling is worth it declares every clue and it stops being a decision. A climb you take back before the clue is read serves no cooldown — the wait is the price of riding a clue at double, not of declaring one, and a withdrawn declaration rode nothing. It cannot be used to peek: declarations are only accepted between clues, so there is no clue on the board to look at.

🎯 TARGETING — standard, on by default. — Press T and pick a player. If you take the clue, the entire pot comes out of them alone and everyone else is spared. If they take it, you pay them the entire pot and the ring goes free. A finishing move and a kamikaze on the same button. Your target is visible to everyone, they get an alert, and 0 cancels. What an aimed miss costs on top is a host setting — nothing extra in Arcade and Chaos, half the pot in Tournament — because the old whole-pot price made ganging up on a leader self-harm.

💀 BOUNTIES — While you're still in the queue, press B to stake up to half your starting score on someone's head. Whoever eliminates them collects it — you enter that much lighter. If they survive the match, they keep it. If they eliminate you, they keep it anyway. Choose carefully.

💎 STABLES — Teams, named for you from a list of gemstones: Diamond, Ruby, Emerald, Sapphire, Onyx, Topaz. Each carries a color and a badge that tints its members' rows on every scoreboard, so the sides read at a glance without anybody reading a name. Six because that is the most a room can tell apart at once. Press J to join one, B to betray. When you win a clue, your stable pays nothing and the whole pot lands on everybody outside it, so the bigger your side the harder each outsider is hit — and the pot is then split evenly across your stable. A stable can hold at most half the ring. Leaving means betrayal: half your stack is split among the people you walk out on, and you pick a new side or go it alone. If the stable eliminates everybody else, it dissolves and they settle it between themselves.

🔁 REVIVAL — Your first elimination sends you back to the queue instead of out the door: reduced starting stake, new entry number, one life only. Expect the match to run about half again as long.

11 · When the computer hosts ALPHA

Not yet played in a real room. Everything below is built and tested; none of it has met an actual Saturday night. Treat the section as a description of intent until a match has run on it, and keep a console open when one does.

Some matches run with no human host. Two synthesized voices share the job, and the split is deliberate: Mike is the game-show host and says everything that touches the rules — the handover of the board, the category and value, the clue, the ruling, the correct response on a stumper — in a straight register, because a ruling delivered in character sounds negotiable. Gene is the ring announcer and calls the events in the wrestling register: entrances, eliminations, the field clearing, overtime. They never speak in the same step.

The voice plays out of every player's own buzzer window rather than over the call. That is a timing decision, not a convenience: the server owns each clip and knows its length, so it arms the buzzers at the end of the read through the same message a human host triggers, and the call's audio delay drops out of the problem. The cost is that a player with sound off has no host, so the buzzer says so until the browser will play, and every client reports when a clip actually started — the spread of those reports across a room is the number that sets how much settle to add, and it is recorded per clip.

Whoever has the board calls the next clue — draw 1 at the bell, then whoever last answered correctly, staying put on a stumper — by saying it aloud, or by clicking it on the board in their buzzer window. Twenty seconds without a call — counted down on the buzzer — and Mike picks one; it was twelve until the first live room asked for room to breathe on the one clock the game can afford it. The category and value are shown on every screen rather than read out, by the same room's request. The spoken call is matched on sound rather than meaning: the dollar value and the category are found by completely separate machinery, so a garbled category cannot cost you the value, and a call that could be either of two categories is refused rather than guessed — putting the wrong clue up and unpicking it in front of the room is worse than asking again.

The ears are the player's own browser. Speech recognition runs on their machine and the server receives words, never audio, which removes a streaming service, an API key, a per-minute bill and a third party holding a room's recordings from the design in one step. It costs Chrome and Edge, and a player who refuses the microphone still clicks their clues and can still be ruled on from a console. Mike judges what he was given through a ladder of deterministic checks before any model is asked — the question form and the thinking out loud are stripped, a misheard "What is" is repaired, a bare "what…" is not an answer — so most rulings never reach a model at all. A player who names only part of the answer is told to be more specific once and keeps the rest of their clock, which is a host's prompt rather than a ruling. With no model available the judge returns "I could not tell" and never "no": a failed call must not be able to cost somebody money.

It will still be wrong sometimes, and that is what the objection is for. Any player — in the ring, queued or eliminated — presses O from the moment a ruling is spoken until the next one is, and the ruling is reversed on a majority of everybody in the match or two thirds of the ring, whichever arrives first. Two thresholds because the roster is thirty and the ring is three: a majority of the room lets the crowd correct a host that is plainly wrong, and two thirds of the ring lets the people with money on the clue correct it without needing twenty-five spectators to look up from their drinks. The answering player's own O counts.

The cost of a rare event belongs on the rare event. So a ruling takes effect at once and the board goes straight back out; objections are collected through the following clue while play carries on. A reversal is then a walk-back rather than a second ruling: the engine keeps a snapshot before every scored clue, so the server restores it and re-runs the same scoring the other way round, and a clue that was in progress simply goes back on the board unrevealed to be picked again. Where the walk-back cannot express what the room wants — two players were both ruled wrong, or a correct answer was overruled and the race that should have followed cannot be run minutes later — every buzzer gets fifteen seconds to vote on who had it, with what the room heard each of them say. Most votes wins; a tie throws the clue out. A thrown-out clue pays nobody, charges nobody and does not advance the clue counter, which matters more than it sounds: the entry clock, the ceiling decay and the overtime trigger all run on that counter, so a voided clue that advanced it would bring the next player in a clue early — which is the one thing the room certainly did not vote for.

Every objection is in the match record with its count, both thresholds and what it did, so the rate at which a room disagrees with the judge is measurable rather than anecdotal. That rate is the number that decides when the judge is good enough.

How to join

Open the buzzer link on your phone or a second tab, enter the four-letter room code the host announces, add your name and (if you like) a photo — your token is one of twenty-four line-art weapons, and if someone already took your crowbar, yours comes in a different color. Pick your entrance music while you wait: one of the built-in 8-bit themes, or bring your own link.

Leave the buzzer page open — it's your buzzer for the whole match — and tap the screen once so your browser allows audio. You'll hear a horn when you enter, a rising countdown over your last three queued clues, and an alert if you're eliminated. On a desktop you can optionally show the board beside your buzzer; on a phone, watch the shared screen.

And start buzzing right away, even from the queue. It doesn't count for anything — which is exactly why it's the right time to practice.

Part II

Why the rules are shaped this way

Every number in the rules was chosen by measurement — a Monte Carlo harness plays hundreds to thousands of simulated matches per question, and two recorded live matches check the model against reality. This part keeps the wrong answers as well as the right ones.

The one number that matters

Under straight televised scoring — answer right, collect the clue's value once — the format simply doesn't work, and it fails by the widest margin a design can fail by: in 4,000 simulated matches, the player with the last entry number won essentially every time.

Figure 1 — The draw-position problem under TV-style scoring

Share of 4,000 simulated match wins vs each group's fair share

0%25%50%75%100%Draws 1–21Draws 1–21 — Measured share of wins: 0%0%Draws 1–21 — Fair share: 70%70%Draws 26–30Draws 26–30 — Measured share of wins: 96.9%96.9%Draws 26–30 — Fair share: 16.7%16.7%Draw 30 aloneDraw 30 alone — Measured share of wins: 46.7%46.7%Draw 30 alone — Fair share: 3.3%3.3%Measured share of winsFair share

Draws 1–21 won nothing at all. The five last entrants took 96.9% of all matches; the final draw alone took 46.7% — fourteen times its fair share.

The cause is arithmetic, not luck. With flat scoring, a player's expected score change per clue is V(2−P)/P — clue value V, ring size P — which is negative for every ring bigger than two. Everyone drifts downward together; the best accumulated score barely cleared the starting stake. A late entrant walks in fresh against a field that has spent half an hour being ground down.

Figure 2 — Expected score change per clue under flat scoring

E[Δ] = V(2−P)/P for a $300 clue, by ring size

$50$0$-50$-100$-150$-200$-2502345678players in the ring (P)$300 clue — 2: $0$300 clue — 3: $-100$300 clue — 4: $-150$300 clue — 5: $-180$300 clue — 6: $-200$300 clue — 7: $-214$300 clue — 8: $-225$300 cluenegative for every P > 2 — the whole field drifts down

Only heads-up play breaks even. In a full ring, even a good player bleeds — which is why nobody could bank a defense against the late numbers.

The fix is the game's signature rule: the answerer collects the clue's value from every opponent, not once. Strong early players can now accumulate a real defense before the late numbers arrive, and the elimination pace is unchanged. Every other economic rule in Part I exists to balance the consequences of this one.

Interlude: trust nothing you didn't shuffle properly

One measurement bug shaped this document more than any design decision, so it's recorded here. Entry order was shuffled with sort(() => rng() − 0.5) — which is not a shuffle. It leaves a heavy bias toward the original order, and it quietly poisoned every draw-fairness number measured with it.

Figure 3 — The shuffle that wasn't

Win rates for draws 1–3 — players who start together and should be statistically identical

0%10%20%30%identical players should tie ≈16.8%Draw 1Draw 1 — Reported win rate: 25.8%25.8%Draw 2Draw 2 — Reported win rate: 14.2%14.2%Draw 3Draw 3 — Reported win rate: 10.4%10.4%

Three identical starting positions 'won' at 25.8%, 14.2% and 10.4%. Any asymmetry between draws 1–3 is a bug by definition, because they all start in the ring together. Fisher–Yates everywhere, ever since.

The rule that came out of it: if a simulation shows draws 1–3 behaving differently from each other, the simulation is broken — and one ceiling recommendation below was reversed when this was fixed.

The ceiling, and why it (no longer) falls

Letting the answerer collect from everyone reintroduced the runaway-leader problem, so scores are capped. The question was what the cap should do over time.

Overturned

The original analysis compared fixed, rising and falling ceilings and concluded that a falling ceiling was the fairness winner — reaching a 53% back-half win rate in 76 minutes, against 118 minutes for the best fixed ceiling. The falling ceiling shipped, and the argument looked airtight.

Figure 4 — The original ceiling experiments (measured with the biased shuffle)

Each point is one ceiling policy: match length vs back-half draw win rate

45%50%55%60%65%70%75%80%607590105120perfect fairness — 50%match length (minutes)back-half draw win rateFixed 6,000: 67 min, 72% back-half winsFixed 6,000Fixed 10,000: 118 min, 55% back-half winsFixed 10,000Rising +40/clue: 100 min, 73% back-half winsRising +40/clueFalling −50/clue: 76 min, 53% back-half winsFalling −50/clue

The comparison that picked the falling ceiling. The shape of the trade-off was real; the verdict wasn't — these runs were contaminated by the shuffle bug above, and the recommendation reversed when it was fixed.

The correction. Measured with a proper shuffle, a falling ceiling favors late draws: it clips whoever is ahead, and whoever is ahead is nearly always an early entrant who has been accumulating. A latecomer arrives at a fixed stake, untouched. So in the current game the ceiling holds still for the whole main match — decay is zero during regulation.

But removing decay entirely broke something else: the ceiling's slow leak was the only drain in the system. A perfectly symmetric exchange between two even players resolves nothing — doubling both sides of an even trade leaves it even — and evenly matched robots once played 400 clues without an elimination. Hence the current rule: the ceiling falls only once overtime opens, where it does its draining work exactly when the match needs an ending and never before.

Small fields needed the most room, not the least. A live 53-clue six-player match had the winner pinned at the 6,000 cap for 20 of them, with 11,930 points swallowed: half the match, they answered correctly and gained nothing. The old ladder assumed bigger fields accumulate faster, which is true per clue and wrong per match — a lone strong player in a six-hander faces the fewest opponents so climbs slowly, but the match runs long enough for that slow climb to reach the cap and sit there. Measured over 2,000 six-player matches, the cap was raised from 6,000 to 10,500: clues spent pinned fell from 22% to 8%, and points swallowed from 5,361 to 2,061. Ceiling decay was tried again as a remedy and rejected again — a falling cap is one the leader meets sooner, so it made pinning worse (29% at −40, 36% at −80) and brought back the late-draw bias it caused the first time.

Two more findings round out the ceiling's story. It — not the entry interval — is the dominant fairness lever, a fact obscured for months by a hidden step in the old code (the cap jumped 7,500 → 11,000 at 25 players, which made 24-player matches look far less fair than 30-player ones until the confound was found; it is now a measured ladder rather than a slope, and deliberately not monotonic: small fields get the most headroom, because a lone strong player there faces the fewest opponents per clue, climbs slowly, and the match runs long enough for that slow climb to reach the cap and sit at it. A live 53-clue six-player match had the winner pinned for 20 of them). And no formula predicted fairness well across configurations (best fit r = 0.71), so the engine ships a measured lookup — 3,000 simulated matches per field-size-and-interval pair — instead of a rule someone guessed.

Figure 5 — What the tuning produced

Current presets by field size, measured by simulation

Match length

0 min20 min40 min60 min10 players10 players — Match length: 35 min35 min16 players16 players — Match length: 46 min46 min20 players20 players — Match length: 48 min48 min30 players30 players — Match length: 60 min60 min

Back-half draw win rate (fair share: 50%)

0%25%50%fair10 players10 players — Back-half win rate: 57%57%16 players16 players — Back-half win rate: 62%62%20 players20 players — Back-half win rate: 64%64%30 players30 players — Back-half win rate: 63%63%

Late numbers keep a real advantage — true to the Royal Rumble tradition — but nothing like the 100% guarantee the format started with. Longer matches trend less fair; shorter ones reward skill less. The presets sit at the compromise.

Overtime took three attempts

The failure modes are all recorded. Escalating stakes alone can't end an even match (doubling both sides of an even trade leaves it even) — that's why overtime also drops the ceiling. Opening overtime too eagerly was worse: at three stall-windows it fired during the entry phase and cost 11 minutes of length and 20 points of fairness, so it now opens when the queue empties, or only after a very long stall (eight escalation windows, about 48 clues) with players still queued. And the payout rule matters: the old behavior paid a raised clue's winner $500 on a $2,000 clue while charging the loser the full $2,000 — winners now bank the whole amount and are clipped on the next clue instead.

Figure 6 — Why overtime exists

Clues needed to resolve a perfectly even heads-up endgame

0 clues100 clues200 clues300 clues400 cluesEvenly matched, no overtimeEvenly matched, no overtime — Clues to resolution: 400 clues400 cluesSame pair, with overtimeSame pair, with overtime — Clues to resolution: 22 clues22 clues

An evenly matched pair stalled for 400 clues without overtime; with the ratcheting multiplier and the falling ceiling it resolved in 22. The multiplier ratchets — an elimination stops the climb, nothing resets it.

The longevity bonus: getting paid to still be here

There is a mechanic in this game that has paid out roughly eight starting stacks per match since the day it shipped, and this handbook has never said one word about it. Time to fix that.

Every player in the ring carries a personal clock: the number of clues resolved since you entered — not since the match started, so a latecomer isn't behind schedule, just on a later clock. Every 10th tick, the game hands you $500. That's it. No buzz required, no answer required, no announcement. It lands quietly inside the clue resolve, stacked on top of whatever else the clue did to you — which is exactly why you've never noticed it. On your 10th clue you didn't gain $500; you gained $500 minus your stake in the pot you just lost, and the scoreboard showed you one number. The money was real the whole time.

Three fine points, because the details are where mechanics live. First, the bonus is flat. The category sweep bonus scales with the overtime multiplier; the longevity bonus does not — survival money doesn't inflate, at ×8 it is still $500. Second, the roof still applies. The bonus lands and then the ceiling clips every live score the way it always does, so longevity can never float you above the roof — reaching past the roof is the top rope's job, and you pay for that one. Third, the host owns the dials: on or off, the interval (default every 10 clues), and the amount (default $500).

Does it matter? The data say yes, more than anyone would guess. In our two biggest recorded matches the clock quietly paid out about $24,500 and $27,000 in total — against a $3,000 starting stack. A player who goes wire to wire in a 98-clue match collects $4,500 in pure survival money, which is more than one entire starting stack for doing nothing but refusing to leave. That's not a rounding error; that's the drain's counterweight. The pot punishes wrong answers and the drain grinds down the quiet — and this little clock pushes back, $500 at a time, for everyone still standing.

The pot pays the fast. The clock pays the stubborn.

Stables, and what a team is worth

A stable protects its members from each other and lands the whole pot on whoever is outside. The first version simply let the pot shrink — teammates paid nothing and the winner collected less — which protected the stable and did nothing to anybody else. Loading the teammates' share onto the outsiders instead is what turned it into a mechanic rather than a no-op.

A stable is only a partial answer to a strong player. One elite in a field of normies wins 73% of twelve-player matches and 94% of six-player ones. A pack helps, monotonically, but does not level it: at twelve players the largest legal pack takes the elite from 73.1% to 68.2%. The pack can only use its advantage on clues it wins, and against somebody taking most of them there are not many. Ganging up directly on one player is what targeting is for.

Loyalty is expensive early and nearly free late. A stable protects you from the people in the ring, so as they leave there is less to be protected from — and the toll on your stack is a one-off while the protection is a stream that dries up. With ten outsiders left, walking out costs an elite 13.7 points of win rate; with three, 1.7; with two, nothing measurable. It never becomes actively better, only cheaper, so a late defection is a read on the room rather than an edge the numbers hand you.

The pot is split evenly across the stable. That taxes a strong member — an elite drops from 88.9% to 78.6% — which is the largest dent in elite dominance anything here achieves. It also more than doubles their reason to stay: the gap between staying and crossing widens from 1.4 points to 8.6. Protection turns out to be worth more than the tax.

The optional mechanics, measured

Figure 7 — What each toggle does to the match

400 simulated 30-player matches per row

Match length

0 min25 min50 min75 min100 minNone (baseline)None (baseline) — Match length: 64 min64 minTop ropeTop rope — Match length: 62 min62 minTargetingTargeting — Match length: 63 min63 minBountiesBounties — Match length: 64 min64 minRevivalRevival — Match length: 95 min95 minAll fourAll four — Match length: 88 min88 min

Back-half draw win rate (fair share: 50%)

0%25%50%75%100%fairNone (baseline)None (baseline) — Back-half win rate: 53%53%Top ropeTop rope — Back-half win rate: 54%54%TargetingTargeting — Back-half win rate: 63%63%BountiesBounties — Back-half win rate: 48%48%RevivalRevival — Back-half win rate: 52%52%All fourAll four — Back-half win rate: 63%63%

Targeting is the fairness-swingiest toggle (+10 points to the late half). Bounties are the only rule that favors the early half. Revival adds half again to the length — nearly everyone uses the second life.

Measurement trap

Revival first measured at a staggering 94% back-half win rate. Artifact: revived players draw fresh entry numbers, so every revived winner looked like a late draw by definition. Scored against the number they originally drew, revival is nearly neutral at 52% — and it's quietly the kindest rule for early draws, who have the most match left to use a second life in.

Part III

Robots, and what real players taught them

Since 0.102.0 there is a second robot set, built from this game's own recorded play rather than from broadcast. The archetypes set deals five types measured from 42 players with 40 or more presses across 46 matches, on two axes that the standards ladder below cannot separate — the typical press, and how far the slow tail sits above it: rhythm regular (29% of players, 127 ms median, 2.9× tail), gambler (7%, 101 ms, 6.9×), reactor (43%, 196 ms, 3.6×), metronome (14%, 292 ms, 1.7×) and straggler (7%, 330 ms, 5.4×). Each carries its own per-row attempt rate and press histogram. Two findings travel with it. Accuracy is flat — 81% to 87% across all five, uncorrelated with speed — so the ladder's accuracy-by-standard models something the data do not contain. And the five boundaries are a design choice, not discovered clusters: k-means on the two axes gave 140 distinct solutions across 400 seeds and 45–53% pair stability at every k, because the field is a continuum, so the cuts are stated thresholds anybody can re-run. Gambler and straggler each rest on three players; their 7% shares could plausibly be 4% or 12%. The analysis chat's measurement and method are in its buzzer-archetypes-packet; the set is reachable through the robots API and not yet from the setup page.

The tuning above needed thousands of matches, which needed believable players. The robot model is ported from Matt Schiffler's real-data generator and validated against J!ometry box scores — 3,339 real player-games — plus 2,784 more in the local archive.

The press times are this game's own since 0.104.0. The robots sampled another game's recordings for their buzz timing — 2,493 presses against a clock on which the human medianed 43 ms — and the analysis chat's recalibration note showed what that cost: two of the five tiers buzzed at the median of the single quickest player ever recorded here, and no bucket ran past 500 ms although one real press in seven is slower than that. The histograms are now built from 33 recorded matches, 64 players and 4,509 live presses, five equal tiers by each player's own median — elite 117 ms, superchamp 144, champ 183, normie 216, rookie 308 — with buckets to 4,000 ms, and the field-matching offset that used to drag the robots 190 ms onto our clock is zero, because they start there. The tier signature is anticipation rather than consistency: elite players press under 150 ms 59% of the time, rookies 17%. Early presses are the one thing the logs could not time, so each tier's early rate is measured here and its timing shape borrowed from the old recordings until the record, which now carries how early, has enough to replace it. The recalibration note's own numbers — 46 matches, 5,191 presses, medians 116/144/190/252/319 — come from a larger bundle than this repository holds and are quoted as theirs.

Figure 8 — Five robot standards vs the real population

Share of the field at each standard: robot design vs 3,339 real player-games

0%20%40%60%RookieRookie — Robot field design: 5%5%Rookie — Real players (J!ometry, 3,339 games): 3.1%3.1%NormieNormie — Robot field design: 60%60%Normie — Real players (J!ometry, 3,339 games): 55.8%55.8%ChampChamp — Robot field design: 23%23%Champ — Real players (J!ometry, 3,339 games): 17.5%17.5%SuperchampSuperchamp — Robot field design: 11%11%Superchamp — Real players (J!ometry, 3,339 games): 12.8%12.8%EliteElite — Robot field design: 1%1%Elite — Real players (J!ometry, 3,339 games): 0.1%0.1%Robot field designReal players (3,339 games)

Attempt rates per standard: Rookie 23–35%, Normie 37–61%, Champ 63–70%, Superchamp 72–88%, Elite 89–96%; accuracy runs 70–80% up to 87–95%. Mean attempt rate: robots 56%, real players 57%. An earlier uniform mix (20% elites) invalidated a round of testing — the field you test against has to look like the field that shows up.

Figure 9 — Difficulty lives in the attempt, not the answer

Modeled attempt rate by board row — each row raises the base rate to a power (exponents 0.48 → 1.40)

0%25%50%75%100%row 1row 2row 3row 4row 5board row (1 = cheapest, 5 = dearest)Casual player (25% base) — row 1: 51.4%Casual player (25% base) — row 2: 39.5%Casual player (25% base) — row 3: 30.8%Casual player (25% base) — row 4: 21.8%Casual player (25% base) — row 5: 14.4%CasualStrong player (90% base) — row 1: 95.1%Strong player (90% base) — row 2: 93.2%Strong player (90% base) — row 3: 91.4%Strong player (90% base) — row 4: 89.1%Strong player (90% base) — row 5: 86.3%StrongCasual player (25% base attempt rate)Strong player (90% base)

A casual player attempts the cheapest row 3.6× more often than the dearest; a strong player, 1.1×. Real data agree: accuracy falls only 1.9 points from top row to bottom while attempts fall 15. Hard clues don't make people wrong — they make them quiet. Across 1,772 real contestants, buzzer win rate separates players far more than accuracy (46–56% vs 82–89%).

Figure 10 — Why robot calibration waits for 16 buzzes

Estimated player reaction speed after 6 buzzes vs the settled figure, two recorded matches

0 ms100 ms200 ms300 msRecorded match ARecorded match A — After 6 buzzes: 302 msRecorded match A — Settled figure (16 buzzes): 85 ms302 ms85 msRecorded match BRecorded match B — After 6 buzzes: 242 msRecorded match B — Settled figure (16 buzzes): 60 ms242 ms60 msAfter 6 buzzesSettled figure (16 buzzes)

People start slowly. Calibrating robots to a player's first six buzzes read 302ms and 242ms where the settled figures were 85ms and 60ms — so robots now calibrate from a 190ms default, settle after sixteen human buzzes, and freeze. (Recomputing every clue made robots chase early sluggishness forever; freezing too early left them uncalibrated.)

How fast should people be fed in?

The auto interval capped at fifteen clues, and with four to six players the arithmetic always hit that cap — three people spread across a half-hour match is twenty-odd clues apart, clamped. Every small game therefore had identical pacing, and the queue took most of the night to empty. The question was what it would cost to feed people in faster.

Figure 11 — Entry interval against draw fairness

Share of wins taken by the back half of the draw; 2,500 simulated matches per point

45%50%55%60%65%what auto picks3456810121520clues between entries4 players, every 3 — 49.4% back half, 22 clues4 players, every 4 — 49.3% back half, 22 clues4 players, every 5 — 51.3% back half, 23 clues4 players, every 6 — 51.4% back half, 24 clues4 players, every 8 — 50.1% back half, 26 clues4 players, every 10 — 53% back half, 28 clues4 players, every 12 — 52.9% back half, 30 clues4 players, every 15 — 50.5% back half, 33 clues4 players, every 20 — 50.6% back half, 37 clues46 players, every 3 — 50.2% back half, 34 clues6 players, every 4 — 49.5% back half, 37 clues6 players, every 5 — 49.2% back half, 40 clues6 players, every 6 — 49% back half, 43 clues6 players, every 8 — 48.8% back half, 48 clues6 players, every 10 — 48.8% back half, 53 clues6 players, every 12 — 48.2% back half, 59 clues6 players, every 15 — 45.5% back half, 67 clues6 players, every 20 — 47.4% back half, 80 clues68 players, every 3 — 47.4% back half, 45 clues8 players, every 4 — 47.5% back half, 49 clues8 players, every 5 — 47.8% back half, 54 clues8 players, every 6 — 48.5% back half, 58 clues8 players, every 8 — 48.6% back half, 67 clues8 players, every 10 — 49.4% back half, 76 clues8 players, every 12 — 47% back half, 85 clues8 players, every 15 — 49.8% back half, 99 clues8 players, every 20 — 53% back half, 121 clues812 players, every 3 — 49.9% back half, 60 clues12 players, every 4 — 50.4% back half, 68 clues12 players, every 5 — 48.3% back half, 76 clues12 players, every 6 — 48.1% back half, 84 clues12 players, every 8 — 49.2% back half, 101 clues12 players, every 10 — 51.7% back half, 117 clues12 players, every 12 — 52.1% back half, 135 clues12 players, every 15 — 56.5% back half, 160 clues12 players, every 20 — 65.8% back half, 196 clues12an even draw

Below ten clues every field size sits within a few points of an even split. Above it the big fields drift: twelve players on a twenty-clue interval hand the back half 66% of the wins.

The interval is nearly free at small sizes and dangerous at scale — which is the opposite of what I expected before measuring. I had assumed feeding people in quickly would punish late draws, because they would arrive into a fuller ring with less time to recover. It does not. What punishes a late draw is arriving late in the match, against opponents who have been compounding since the first clue. A twelve-player field on the old cap put the last entrant 86% of the way through the night; they either walked into a fortune or a graveyard, and the draw decided which.

What the interval really controls is length. Six players at a five-clue gap play 40 clues; at twenty they play 80. A host asking for a faster game is asking for exactly this dial, and it turns out they can have it:

Old cap of 15, against the new cap of 10 — back-half share of wins, and the median clues played:

4 players: 50.5% over 33 clues → 53.0% over 28.
6 players: 45.5% over 67 clues → 48.8% over 53.
8 players: 49.8% over 99 clues → 49.4% over 76.
12 players: 56.5% over 160 clues → 51.7% over 117.

Fairness is unchanged or better at every size, and matches get materially shorter. The twelve-player case improves most, from 56.5% back-half wins to 51.7%. The cap is now ten.

Eight measured just as fairly and was tried first, but it broke the common case: six players wanting a fifteen-minute game could no longer reach it on auto, so the estimator warned every single time. A correct warning that fires constantly is a warning nobody reads. Ten keeps that game reachable and still takes a third off the old cap.

One thing this does not fix: the interval also decides when overtime can open, because overtime waits for the queue to empty. On the old cap a long match could reach its natural end before anybody was left in the queue to trigger it. That is a separate problem and it is still open.

Targeting as a standard rule

Targeting became standard because it is the only mechanic in the box that actually catches somebody running away with a match. Stables, measured earlier, turn out to do almost nothing against a strong player: a pack of normies moves an elite's win rate by a point or two. Targeting moves it properly, because it puts the whole pot on one head instead of spreading it.

One champ in a field of normies, 2,000 matches per setting. Their win rate, and the share of wins taken by the back half of the draw:

6 players — off: 52.5% champ, 46.7% back half · ganging up: 50.6% champ, 48.1% back half.
10 players — off: 39.9%, 49.6% · ganging up: 37.8%, 57.0%.
16 players — off: 31.9%, 48.4% · ganging up: 30.7%, 66.3%.

It works, and it has a cost that grows with the field. At six players — the size most of these matches actually are — the draw barely moves: 46.7% to 48.1%, which is if anything closer to even. At sixteen it moves a long way, from 48.4% to 66.3% in favor of the back half of the draw.

The reason is worth stating plainly, because it is not obvious. A room ganging up aims at whoever is leading, and the leader is nearly always somebody who entered early and has had time to accumulate. A late entrant is never the leader and so is never the target. Targeting therefore protects late draws as a side-effect of doing its job. In a small field there is not enough match left for that to compound; in a large one there is.

It is on by default anyway. A room that can see one player running away with the night and has no way to respond is a worse problem than a draw that leans late, and the host can switch it off for a big field.

Part IV

Reality check

Updated September 14, 2026 — twenty-four recorded human matches across seven player groups, through the 180-clue QPAL marathon, the arrival of two brand-new rooms, the first Tournament matches on record, an eleven-player field, and the first time anybody used an optional mechanic. The human buzz median holds at roughly 150ms under a typical host setup (in QPAL, 303 of 613 warm-up buzzes — 49% — came in under it), though the new rooms buzz on a slower clock entirely — and, as of 0.96.1, we can finally see that a good share of the fast presses in the experienced rooms are not reactions at all. Players are anonymized as P1–P45 throughout; the mapping lives outside this document.

The estimator, humbled — and diagnosed

Two matches in, the estimator was two-for-two. Easy game.

Eleven matches in, the record is more honest — and the failure mode finally has a name.

Figure 12 — The length estimator against live play

Predicted vs actual match length in clues, all eleven recorded matches

0 clues50 clues100 clues150 cluesTQUV · Aug 12TQUV · Aug 12 — Estimator predicted: 100 cluesTQUV · Aug 12 — Actual: 84 clues100 clues84 cluesRYRP · Aug 12RYRP · Aug 12 — Estimator predicted: 75 cluesRYRP · Aug 12 — Actual: 69 clues75 clues69 cluesUVQC · Aug 14UVQC · Aug 14 — Estimator predicted: 54 cluesUVQC · Aug 14 — Actual: 50 clues54 clues50 cluesTAQU · Aug 14TAQU · Aug 14 — Estimator predicted: 52 cluesTAQU · Aug 14 — Actual: 48 clues52 clues48 cluesMBTR · Aug 15MBTR · Aug 15 — Estimator predicted: 75 cluesMBTR · Aug 15 — Actual: 53 clues75 clues53 cluesGETM · Aug 15 †GETM · Aug 15 † — Estimator predicted: 75 cluesGETM · Aug 15 † — Actual: 31 clues75 clues31 cluesFTGL · Aug 15FTGL · Aug 15 — Estimator predicted: 37 cluesFTGL · Aug 15 — Actual: 84 clues84 clues37 cluesQPAL · Aug 15QPAL · Aug 15 — Estimator predicted: 75 cluesQPAL · Aug 15 — Actual: 180 clues180 clues75 cluesJHLZ · Aug 15JHLZ · Aug 15 — Estimator predicted: 37 cluesJHLZ · Aug 15 — Actual: 32 clues37 clues32 cluesVWQW · Aug 16VWQW · Aug 16 — Estimator predicted: 54 cluesVWQW · Aug 16 — Actual: 139 clues139 clues54 cluesLNET · Aug 16LNET · Aug 16 — Estimator predicted: 107 cluesLNET · Aug 16 — Actual: 93 clues107 clues93 cluesEstimator predictedActual

† GETM was ended early by the host, so its miss is administrative. The three real misses share one cause: FTGL (four players, 2.3× over — tiny fields trade points without eliminating anyone), QPAL (75 predicted, 180 played), and VWQW (54 predicted, 139 played) — the latter two both featuring latecomers, second lives, and deep overtime. LNET, played the same night as VWQW with no latecomers and no returns firing, came in at 93 against a predicted 107. The estimator is accurate exactly when nothing brings players back; return mechanics and latecomers are what it needs to learn.

Pace keeps beating the model. It assumes 17.5 seconds per clue. The first night's median came in at 19.3, and every match since the rooms learned the game has run 12.9–16.6.

Rooms speed up. The model does not.

Figure 13 — Median pace by match, in order played

Median seconds per clue; the pace model assumes 17.5s

0s5s10s15s20smodel assumes 17.5sTQUVTQUV — Median seconds per clue: 19.3s19.3sRYRPRYRP — Median seconds per clue: 16.8s16.8sUVQCUVQC — Median seconds per clue: 16.4s16.4sTAQUTAQU — Median seconds per clue: 14.9s14.9sMBTRMBTR — Median seconds per clue: 17.1s17.1sGETMGETM — Median seconds per clue: 19.6s19.6sFTGLFTGL — Median seconds per clue: 19.8s19.8sQPALQPAL — Median seconds per clue: 14.1s14.1sJHLZJHLZ — Median seconds per clue: 12.9s12.9sVWQWVWQW — Median seconds per clue: 16.6s16.6sLNETLNET — Median seconds per clue: 14.8s14.8s

Experienced rooms are markedly faster than the model. The slow outliers (GETM, FTGL) are the same night a brand-new group learned the game — pace is a property of the room, not the format, and the estimator could take a per-room prior.

Figure 14 — The prediction the record kept, against what happened

Recorded prediction versus actual clues. Pace was close to model (18.7–18.9 s/clue against 17.5 assumed), so the error is almost entirely clue count

0255075100Match 12 (GMRF) — 10 lives from 5 playersMatch 12 (GMRF) — 10 lives from 5 players — Predicted clues: 5252Match 12 (GMRF) — 10 lives from 5 players — Actual clues: 101101Match 13 (SLDV) — 10 lives from 5 playersMatch 13 (SLDV) — 10 lives from 5 players — Predicted clues: 3232Match 13 (SLDV) — 10 lives from 5 players — Actual clues: 5656Predicted cluesActual clues

This pair diagnosed a bug rather than a modeling gap. The length estimator existed twice: public/estimate.js applied a multiplier for revival, the engine’s expectedClues() did not, and the server recorded the engine’s. So these bars are not what the host was shown — the setup page predicted 89 and 55 against actuals of 101 and 56, one of them off by a single clue. Fixed in 0.95.2: one function, in the engine, with the page delegating to it. What remains genuinely unmodelled is the tail. With revival on and revivalLimit 1, lives are bounded at roster × (1 + revivalLimit), and both matches hit that ceiling exactly — all five players revived once, ten lives each. A small room then empties its queue early, overtime opens (clue 33 of 101, clue 13 of 56), and the slow drain that follows is the part no term covers. Read the bound from the settings; an earlier draft proposed a fitted ×1.7, which was an average of a miscount.

Then five matches arrived from two brand-new rooms and the miss changed shape entirely. The clue counts were close. The clock was not.

Figure 15 — The estimator's remaining miss is pace — and it's a new-room thing

Predicted versus actual minutes; clue-count predictions were close everywhere except the quick-start match (M18)

010203040M14 (Tournament, Host D)M14 (Tournament, Host D) — Predicted minutes: 1515M14 (Tournament, Host D) — Actual minutes: 3939M15 (Arcade, Host D)M15 (Arcade, Host D) — Predicted minutes: 1515M15 (Arcade, Host D) — Actual minutes: 4141M16 (Arcade)M16 (Arcade) — Predicted minutes: 1515M16 (Arcade) — Actual minutes: 4444M17 (Arcade, Host C)M17 (Arcade, Host C) — Predicted minutes: 2929M17 (Arcade, Host C) — Actual minutes: 2626M18 (Arcade, quick start)M18 (Arcade, quick start) — Predicted minutes: 66M18 (Arcade, quick start) — Actual minutes: 1616Predicted minutesActual minutes

First, the good news: the lives fix works. M17 — eight players, revival off — predicted 100 clues and 29 minutes against 78 and 26 actual, the best big-room prediction on record. Now the bad news: the new rooms ran 31.8–40.3 seconds per clue against the model's 17.5, so near-correct clue counts still turned into 2.6× time misses. Experienced rooms really do run 16–18 seconds. First-night rooms run about double. The cheap fix is already sitting in the record — every match logs its own pace median, so re-estimate from it after ten clues or so, and start an unknown room at a ~30-second prior. M18's miss is a different animal: a quick-start bug (a three-player lobby locks the entry interval at 1, and latecomers flood in one per clue), covered in the dev notes.

And then, once the lives fix landed, the miss flipped sign. Same rooms, same pace, predictions now running long instead of short.

Figure 16 — The estimator now runs about 20% long in experienced rooms

Predicted versus actual clues for the three recent 7–8 player matches with revival off; pace was on model (16.8–20.3 s/clue)

0255075100M17 (8 players, Host C)M17 (8 players, Host C) — Predicted clues: 100100M17 (8 players, Host C) — Actual clues: 7878M19 (7 players, Host B)M19 (7 players, Host B) — Predicted clues: 100100M19 (7 players, Host B) — Actual clues: 8585M20 (7 players, Host B)M20 (7 players, Host B) — Predicted clues: 100100M20 (7 players, Host B) — Actual clues: 7676Predicted cluesActual clues

After the lives fix, the miss changed sign: 100 clues predicted against 78, 85 and 76 played. Pace isn't the culprit — these rooms ran 15–18 second medians, right on the 17.5 assumption. What's left is the overtime tail. Seven players at interval 15 empties the queue around clue 60, and the drain then finished in 16–25 clues, shorter than the tail the estimator budgets. A revival-off tail of roughly 20 overtime clues would have put all three within a handful. The new-room pace problem from matches 14–18 (31–40 s/clue) is a separate animal and still needs the live re-estimate.

The host is part of the clock

Play two matches back to back with different hosts and something strange falls out of the logs: every player's reaction time moves together when the host changes, at an identical delay setting.

The compensation is a fixed number. The audio path and the reading cadence are not — those belong to whoever is holding the clues. Which means a personal best is only a personal best against the same host.

Figure 17 — Same players, different host, same delay setting

Median contested reaction for the three players who played both Aug 16 matches

0 ms100 ms200 ms300 ms400 msP11P11 — Median under Host A: 117 msP11 — Same player, Host B: 162 ms162 ms117 msP9P9 — Median under Host A: 156 msP9 — Same player, Host B: 246 ms246 ms156 msP13P13 — Median under Host A: 218 msP13 — Same player, Host B: 388 ms388 ms218 msMedian under Host ASame player, Host B

Everyone shifted slower together under Host B — +45ms, +90ms, +170ms — which is a calibration offset, not a skill change. Within a match the race stays fair (the shift is shared); across matches, personal bests are host-relative. The fix suggests itself: the same calibrate-and-freeze machinery the robots use against players, applied to the host — and until then, the match log should record who hosted.

Figure 18 — A third host, and this time no shift

Settled live buzz medians (first six buzzes per player excluded) for the four players who played both Aug 22 matches, by host

0ms100ms200ms300msP12P12 — Host C (SLDV): 129msP12 — Host A (GMRF): 140ms140ms129msP13P13 — Host C (SLDV): 268msP13 — Host A (GMRF): 269ms269ms268msP14P14 — Host C (SLDV): 179msP14 — Host A (GMRF): 168ms179ms168msP9P9 — Host C (SLDV): 235msP9 — Host A (GMRF): 160ms235ms160msHost C (SLDV)Host A (GMRF)

Three of the four shared players sit within ±11ms across the two hosts — no uniform shift, unlike the +45/+90/+170ms measured between Host A and Host B at identical delay settings (Figure 17). P9 moves +75ms alone, and a host offset that moves one player is not a host offset. So the calibration finding is real but not universal: it appeared between one pair of hosts and not another. Caveat: match 13 was the night’s first, so its early buzzes lean cold even with the warm-up exclusion.

The shark problem

Before the individual findings, the map. Every configuration this project has ever measured sits on it, which makes it the index to everything that follows.

Watch the vertical axis. That is a casual player's real chance of winning a match, and almost nothing moves it.

The two panels are left unnumbered on purpose. They are the contents page for this section, not another result in it.

The map — every measured configuration

Top shark's win probability against the casuals' combined win probability; same six-player field throughout (two sharks, one mid, three casuals)

0%5%10%15%45556575top shark's match-win probability (favors high skill →)casual players' match-win probabilityBaseline: 79.5 min, 0.3% back-half winsBaselineKickout on 2: 80.6 min, 0.1% back-half winsKickout on 2Photo-finish 100ms: 65.5 min, 0.1% back-half winsPhoto-finish 100msTrailing pick: 76.6 min, 0.1% back-half winsTrailing pickTenure gate: 76.4 min, 1.9% back-half winsTenure gateStables alone: 76.1 min, 3% back-half winsStables aloneWinner cooldown: 49.2 min, 1.7% back-half winsWinner cooldownPermanent 80% boost: 47.3 min, 4.5% back-half winsPermanent 80% boostComeback, gate wins<1: 74.3 min, 4.6% back-half winsComeback, gate wins<1gate wins<2: 68.5 min, 9% back-half winsgate wins<2gate wins<3, flat stake: 61.9 min, 11% back-half winsgate wins<3, flat stakegate wins<5: 50.6 min, 11.7% back-half winsgate wins<5Comeback ungated: 48.7 min, 2.9% back-half winsComeback ungatedShipped 0.89: 56.9 min, 14.3% back-half winsShipped 0.89+ arrival exemption: 54.3 min, 16.2% back-half wins+ arrival exemptionComeback + stables: 55.3 min, 14.2% back-half winsComeback + stables+ photo-finish 50ms: 57.5 min, 12.1% back-half wins+ photo-finish 50ms+ winner cooldown: 48.2 min, 12.2% back-half wins+ winner cooldown
no comebackcomeback familycomeback + second system

Reading the corners: the lower-right is the sharks' home — baseline, kickout on 2, trailing pick, and the tenure gate all leave the strongest buzzer near 80% with the casuals at zero. The lower-left is the trap corner: winner cooldown, the ungated comeback, and even a permanent 80% buzz boost knock the top shark down without giving the casuals anything — they just promote the second shark. The upper-left is where design effort paid: the gated comeback and everything built on it. The gate threshold walks the curve (wins<1 → 2 → 3 → 5), the OT-scaled stake and the arrival exemption push it further, and the best measured point — comeback + arrival exemption at 16.2% casual, 54.3% shark — leaves the favorite favored but beatable. Photo-finish sits alone mid-bottom: it flattens shark-versus-shark and helps nobody else.

The short list — best measured configurations for casual players

Casual players' combined match-win probability, top five against baseline

0%5%10%15%Comeback + arrival exemptionComeback + arrival exemption — Casual players' win probability: 16.2%16.2%Shipped 0.89 (gated, OT-scaled)Shipped 0.89 (gated, OT-scaled) — Casual players' win probability: 14.3%14.3%Comeback + stablesComeback + stables — Casual players' win probability: 14.2%14.2%Comeback + winner cooldownComeback + winner cooldown — Casual players' win probability: 12.2%12.2%Gate loosened to wins<5Gate loosened to wins<5 — Casual players' win probability: 11.7%11.7%Baseline (no help)Baseline (no help) — Casual players' win probability: 0.3%0.3%

Everything above 10% shares one ingredient: the gated comeback. What varies is the second system — arrival-exemption and stables add to it; race-structure levers (cooldown, photo-finish) subtract from it, because race allocation is one budget.

The pressing open question from live play isn't length — it's whether anyone but the sharks can win. Across the eleven recorded matches (ten with a single winner), buzzer speed decides races, races decide points, and points decide survival.

Figure 19 — Buzzer skill is the whole ballgame

Every player with 20+ contested buzzes across the first nine matches; blue marks the three strongest regulars (P1, P2, P3)

20%40%60%80%100150200250300median contested reaction (ms)share of buzzer races wonP1: 93 min, 48.2% back-half winsP1P3: 123 min, 47.7% back-half winsP3P2: 128 min, 50.8% back-half winsP2P5: 101 min, 78.6% back-half winsP5P6: 138 min, 57.1% back-half winsP6P7: 221 min, 48.6% back-half winsP7P8: 162 min, 47.4% back-half winsP8: 137 min, 48.3% back-half wins: 131 min, 51.2% back-half wins: 179 min, 42.4% back-half wins: 187 min, 37.1% back-half wins: 247 min, 42.9% back-half wins: 234 min, 43.8% back-half wins: 198 min, 45.2% back-half wins: 160 min, 32.4% back-half wins: 233 min, 37% back-half winsP10: 278 min, 28.1% back-half winsP10: 139 min, 23.8% back-half wins

The trend is monotone: faster median reaction, more races won. P1, P2, and P3 entered six of the first eight decided matches and won five of them (P2 alone is 3-for-4). P5 won 78.6% of races in the one match without a comparable rival. Note the x-axis carries the host-clock caveat above — the ranks are solid, the absolute milliseconds are host-relative.

Figure 20 — Who wins the match, by within-field speed rank

All ten decided matches, by the winner's median-reaction rank in their own field

0246810The field's fastest playerThe field's fastest player — Matches won: 44The second-fastestThe second-fastest — Matches won: 33Everyone slowerEveryone slower — Matches won: 33

Wins from outside the top two exist — but both of the early 'upsets' came in fields that were later found to be uniformly strong or uniformly slow, and the two new matches went to the second-fastest (P11 over P2, heads-up at the wire) and a fourth-fastest whose rank is muddied by late entrants with few buzzes. In every clearly mixed field, within-field speed rank has ruled.

QPAL makes the concentration vivid: P2 and P1 took 125 of 165 buzzer races between them — 86% with P4 included — while the other five entrants split 23 races across 48 minutes. P9 was eliminated at clue 10, revived, and was out again at 66 having won two races all night. And the match's final latecomer (fed into ×2 overtime with a fixed stake, eliminated in six clues; revived at half stake into ×4, eliminated in one) shows the endgame compounding the problem: the deeper the match runs, the more it belongs to whoever is already on top.

Figure 21 — Races won in the first two matches on 0.95.x

Contested-clue wins per player; 99 races in match 12, 54 in match 13

0102030P11P11 — Match 12 (GMRF, Host A): 3434P11 — Match 13 (SLDV, Host C): 00P12P12 — Match 12 (GMRF, Host A): 2525P12 — Match 13 (SLDV, Host C): 55P2P2 — Match 12 (GMRF, Host A): 00P2 — Match 13 (SLDV, Host C): 2323P9P9 — Match 12 (GMRF, Host A): 55P9 — Match 13 (SLDV, Host C): 1414P13P13 — Match 12 (GMRF, Host A): 1919P13 — Match 13 (SLDV, Host C): 55P14P14 — Match 12 (GMRF, Host A): 1010P14 — Match 13 (SLDV, Host C): 55Match 12 (GMRF, Host A)Match 13 (SLDV, Host C)

Both wins went to the fast lane again — P12 (140ms settled median, 93% first-buzz accuracy) and P2 (121ms) — making it 12 of 13 decided matches for a shark-tier buzzer. But the floor has risen visibly: every player won at least five races in both matches, and P13, the slowest buzzer in the field at 269ms, took 19 races in match 12 — third most — on the back of two lives and the comeback edge. P11 led match 12 on races with 34 and did not win it: most races is not the match, twice now. Each match had five players; P2 and P11 each hosted the match they sat out.

Everything above is the room that existed when the comeback shipped: one group, resident sharks, the fastest buzzer taking the match twelve times in thirteen. Then the game reached rooms that had never played it. The streak did not survive contact.

Figure 22 — The fastest buzzer won once in five matches

Winner's settled buzz median against the fastest settled median in the same room (first six buzzes per player excluded); a single dot means the winner WAS the fastest

0ms200ms400ms600msM14 · Tournament · Host DM14 · Tournament · Host D — Fastest settled buzzer in the room: 325msM14 · Tournament · Host D — The winner: 325ms325ms325msM15 · Arcade · Host DM15 · Arcade · Host D — Fastest settled buzzer in the room: 346msM15 · Arcade · Host D — The winner: 597ms597ms346msM16 · Arcade · host unrecordedM16 · Arcade · host unrecorded — Fastest settled buzzer in the room: 300msM16 · Arcade · host unrecorded — The winner: 365ms365ms300msM17 · Arcade · Host CM17 · Arcade · Host C — Fastest settled buzzer in the room: 122msM17 · Arcade · Host C — The winner: 182ms182ms122msM18 · Arcade · host unrecordedM18 · Arcade · host unrecorded — Fastest settled buzzer in the room: 215msM18 · Arcade · host unrecorded — The winner: 215ms215ms215msFastest settled buzzer in the roomThe winner

In this batch the fastest settled buzzer won exactly once — M14, the Tournament match, where the comeback is off and the buzzer is supposed to be the whole game. Everywhere the comeback was on, the match went somewhere else. M17's winner was seventh of eight by settled speed. M15's was fifth of six. That is not a fluke in one room; it's two rooms, two hosts, five matches. It's exactly the outcome the whole balance program was built to produce, and it showed up in exactly the rooms that didn't have a 130ms-class resident. Lace 'em up.

Figure 23 — So what won instead? Accuracy.

First-buzz accuracy against settled buzz median, all players with enough contested clues in the batch; match winners highlighted

25%50%75%100300500700900settled buzz median (fast → slow)first-buzz accuracyP29: 126, 67%P29P31: 157, 50%P31P12 won M17: 182, 81%P12 won M17P27: 182, 71%P27P32: 190, 78%P32P28 won M18: 213, 81%P28 won M18P30: 214, 67%P30P11: 225, 77%P11P17 won M14: 312, 77%P17 won M14P21: 346, 75%P21P23 won M16: 365, 61%P23 won M16P19: 442, 56%P19P25: 487, 55%P25P18 won M15: 540, 70%P18 won M15P20: 637, 55%P20P24: 901, 29%P24

Four of the five winners sit on the accuracy frontier of their speed class — P12 and P28 at 81%, P17 at 77% on the most races of the batch (44), P18 at 70% from the slow half of the field. In a room where nobody owns the buzzer, the pot economy pays for being right: a 600ms press that lands beats a 350ms press that whiffs, twice over, because the miss pays the pot. That's the game we wanted. The one low-accuracy path that still worked was M14's — the Tournament room, where there's no comeback to punish a runaway, and speed is allowed to be king.

Two things carried over from that batch and neither is a footnote. Overtime is still where these matches live — 37–83% of clues in the four Arcade rooms — and every one of them ended in a full wipe, the winner banking above the fallen roof. And the new rooms buzz on a different clock entirely: settled medians of 300–900ms where the established room runs 120–270.

Which is the whole point. That slower room is the population the comeback, the kickout research and the pace work were built to serve. Participation held there too: 21 of 22 player-matches won at least one race.

One room, two modes, two different games

Matches 19 and 20 are the cleanest experiment the live record has produced. Same seven players. Same host. Same night. Played once as Arcade and once as Tournament, and the two matches do not resemble each other.

Figure 24 — Flip the mode switch and the whole room reshuffles

Contested-clue wins per player, Arcade match beside Tournament match; eight players appear across the two (P36 played only M19, P13 only M20)

0102030P35P35 — M19 · Arcade (comeback on): 3030P35 — M20 · Tournament (comeback off): 66P32P32 — M19 · Arcade (comeback on): 1919P32 — M20 · Tournament (comeback off): 11P34P34 — M19 · Arcade (comeback on): 1313P34 — M20 · Tournament (comeback off): 1111P37P37 — M19 · Arcade (comeback on): 44P37 — M20 · Tournament (comeback off): 1010P30P30 — M19 · Arcade (comeback on): 44P30 — M20 · Tournament (comeback off): 2929P31P31 — M19 · Arcade (comeback on): 44P31 — M20 · Tournament (comeback off): 00P36P36 — M19 · Arcade (comeback on): 00P36 — M20 · Tournament (comeback off): 00P13P13 — M19 · Arcade (comeback on): 00P13 — M20 · Tournament (comeback off): 11M19 · Arcade (comeback on)M20 · Tournament (comeback off)

P35 took 30 races and the Arcade match, then 6 in Tournament. P32 took 19 races and second place in Arcade, then one race and the first elimination (clue 14) in Tournament. P30 went the other way entirely: 4 races in Arcade (a late draw, in at clue 60) to 29 of 64 in Tournament, at a 46ms settled median and 91% accuracy, and the match. Comeback off, the fastest press owns the room. Comeback on, the same room spreads its races across four or five people. That's the entire Arcade/Tournament distinction, shown by one room in one evening. Same people. Same host. Same night.

That is the Arcade/Tournament distinction, demonstrated by one room in one evening rather than argued from a simulator. Widen it to every match on record and the pattern is the same shape.

Figure 25 — Twenty matches in: the mode picks the winner, not the room

How often the match went to the fastest settled buzzer in the field (first six buzzes per player excluded), grouped by era and mode

0%25%50%75%100%Matches 1–13 · established room, Arcade eraMatches 1–13 · established room, Arcade era — Share of matches won by the fastest settled buzzer: 92%92%Matches 15–24 · new rooms, ArcadeMatches 15–24 · new rooms, Arcade — Share of matches won by the fastest settled buzzer: 0%0%Matches 14, 20 & 22 · TournamentMatches 14, 20 & 22 · Tournament — Share of matches won by the fastest settled buzzer: 100%100%

The first thirteen matches — one established room with resident sharks — went to the fastest tier twelve times. Five Arcade matches in new rooms since then went to the fastest buzzer zero times (winners placed 5th of 6, 2nd of 4, 7th of 8, 2nd of 3, and 3rd of 7 by speed). Both Tournament matches went to the fastest, and not by a little — 23 of 48 races, then 29 of 64. The comeback is doing precisely what it measured in six thousand simulated matches: in Arcade, accuracy and staying power decide; in Tournament, the buzzer does. Hosts now have a real choice when they pick the mode, and the data say what each one buys you.

Which means the mode is not a flavor setting. It is the single biggest lever a host touches, and it decides what kind of skill wins the night. Pick Tournament and the fastest press owns the room. Pick Arcade and the room shares it out.

What was measured, and what shipped

The design goal, verbatim: "I want everyone to have SOME chance to win in this."

Not an equal chance. The sharks should stay favored. But a real one.

So no rule changed until it had been measured. Every candidate lever went through the shipping engine at 1,500 matches per configuration, with a race model calibrated to the live data and a field of two sharks, one mid-tier player, and three casuals.

Figure 26 — What actually moves a casual player's chance to win

Simulated match-win probability for the three casual archetypes combined, 1,500 matches per configuration

0%5%10%15%No help (baseline)No help (baseline) — Casual players' match-win probability: 0.1%0.1%Kickout on 2 (70% after 10 misses)Kickout on 2 (70% after 10 misses) — Casual players' match-win probability: 0.1%0.1%Comeback for everyone (70%, 40 races)Comeback for everyone (70%, 40 races) — Casual players' match-win probability: 1.9%1.9%Gated comeback (70%, 40 races)Gated comeback (70%, 40 races) — Casual players' match-win probability: 7.3%7.3%Gated comeback + stablesGated comeback + stables — Casual players' match-win probability: 14.2%14.2%

The intuitive fix — a kickout on 2 that upgrades your buzzer after ten straight losses — measured at exactly zero: it fires rarely, steals one race, and resets. Even making the casuals the fastest buzzers on the floor topped out under 5%, because the match is won by the economy, not the buzzer. What worked was the comeback with a gate: eliminated with fewer than three races won means an instant return at half stake with a 70% buzz boost for forty races. Ungated, the same mechanic is a shark subsidy (their free lives feed the second shark); gated, it is a targeted underdog rule that is nearly sandbag-proof, since staying under three race wins means not scoring. Top-shark win rate drops from 81% to 55–58%; cost is about seven extra clues.

Shipped — first night's returns

The gated comeback shipped in 0.85 with the measured configuration (gate: three race wins, half stake, 70% boost, forty races) and was live for both August 16 matches. Its first night delivered the intended second act: in VWQW, P11 was eliminated at clue 36, came back at half stake, and climbed to a $10,700 peak by clue 119 — reaching the final three of a match that went to the full ×8 overtime. The books stay honest, though: all ten decided matches to date have been won by shark-tier buzzers, and one comeback return during deep overtime came back with the wrong stake (a bug now with the developers).

And then time to eat a correction. The attribution above is overturned: the logs give P11 eight race wins at elimination, so he would not have passed the wins<3 gate at all. Both revival and the comeback were switched on in VWQW, and that return was revival's queue re-entry. The comeback's real first-night clients were the nought- and one-win players — exactly who it was built for. The story stands; only the credit moves. Whether the rule moves the win column — the probability that a bottom-half buzzer takes a match, ~zero with a P1-class player in the room — is a question the next dozen match logs get to answer.

Where the trigger belongs

Eighty-nine eliminations across the eleven recorded matches, parsed to wins-since-entry, tenure, overtime flag and life number. Three questions that had been argued from intuition, now answered by the record.

Written down so nobody re-litigates them without new data.

Figure 27 — Race wins at the moment of first elimination

All 63 first-life eliminations across the eleven recorded matches

05101520250 wins0 wins — First eliminations: 14141 win1 win — First eliminations: 882 wins2 wins — First eliminations: 223 wins3 wins — First eliminations: 334 wins4 wins — First eliminations: 665 wins5 wins — First eliminations: 336 wins6 wins — First eliminations: 227 wins7 wins — First eliminations: 118+ wins8+ wins — First eliminations: 2424

Bimodal, with the valley at 2–3 wins: 24 players went out with zero or one race win — the flattened — then a gap, then the long tail of players who genuinely played. A gate at 'fewer than 3' cuts almost exactly at the valley, catching 24 of 63 first eliminations. Tenure turns out to be the wrong axis entirely: the zero-win group mostly lasted 9–17 clues in the ring (they weren't eliminated fast, they were invisible), so a tenure gate misses them. And 24 of these 63 eliminations happened during overtime, which is why the overtime question below matters as much as the threshold.

Figure 28 — What each gate does to who wins

Simulated match-win probability by trigger design; 1,500 matches each, boost fixed at 70% for 40 races

0%5%10%15%No comebackNo comeback — Casual player wins: 0.3%0.3%No comeback — Mid player wins: 1.4%1.4%Wins < 1Wins < 1 — Casual player wins: 4.6%4.6%Wins < 1 — Mid player wins: 3.5%3.5%Wins < 2Wins < 2 — Casual player wins: 9%9%Wins < 2 — Mid player wins: 6.2%6.2%Wins < 3 (shipped)Wins < 3 (shipped) — Casual player wins: 11%11%Wins < 3 (shipped) — Mid player wins: 10%10%Wins < 5Wins < 5 — Casual player wins: 11.7%11.7%Wins < 5 — Mid player wins: 15.2%15.2%No gate (everyone)No gate (everyone) — Casual player wins: 2.9%2.9%No gate (everyone) — Mid player wins: 12.5%12.5%Tenure < 12 cluesTenure < 12 clues — Casual player wins: 1.9%1.9%Tenure < 12 clues — Mid player wins: 2.7%2.7%Casual players (combined)Mid player

Top shark's win probability under the same gates

0%25%50%75%No comebackNo comeback — Top shark wins: 79.5%79.5%Wins < 1Wins < 1 — Top shark wins: 74.3%74.3%Wins < 2Wins < 2 — Top shark wins: 68.5%68.5%Wins < 3 (shipped)Wins < 3 (shipped) — Top shark wins: 61.9%61.9%Wins < 5Wins < 5 — Top shark wins: 50.6%50.6%No gate (everyone)No gate (everyone) — Top shark wins: 48.7%48.7%Tenure < 12 cluesTenure < 12 clues — Top shark wins: 76.4%76.4%

The threshold is a real dial with a clear shape. Casual benefit saturates at wins<3 (11.0% → 11.7% going to 5); the marginal fires from a looser gate flow to the mid player instead (10.0% → 15.2%). Removing the gate entirely collapses the casual benefit to 2.9% — the sharks' free lives wash the subsidy out — and a tenure gate does almost nothing, confirming the empirical picture. The top shark falls from 79.5% (no comeback) to 61.9% at the shipped gate and 50.6% at wins<5.

Figure 29 — Who the free lives actually go to

Comeback fires per match by tier, for three gate designs

0123Wins < 3 (shipped)Wins < 3 (shipped) — Casual fires: 22Wins < 3 (shipped) — Mid fires: 0.30.3Wins < 3 (shipped) — Elite fires: 0.10.1Wins < 5Wins < 5 — Casual fires: 2.62.6Wins < 5 — Mid fires: 0.50.5Wins < 5 — Elite fires: 0.20.2No gate (everyone)No gate (everyone) — Casual fires: 33No gate (everyone) — Mid fires: 11No gate (everyone) — Elite fires: 1.81.8CasualMidElite

The gate's whole job in one chart: at wins<3 the elites essentially never qualify (0.1 fires/match); ungated they take 1.8 free lives per match — more than anyone, because they're in the ring longest. The shipped gate is a targeted subsidy; no gate is a shark benefit with a casual garnish.

Figure 30 — Three overtime behaviors, gate fixed at wins < 3

Simulated match-win probability when the comeback fires flat, not at all, or with a multiplier-scaled stake during overtime

0%5%10%15%Fire at flat half stake (shipped)Fire at flat half stake (shipped) — Casual player wins: 11%11%Fire at flat half stake (shipped) — Mid player wins: 10%10%Don't fire during overtimeDon't fire during overtime — Casual player wins: 6.3%6.3%Don't fire during overtime — Mid player wins: 4.7%4.7%Stake scales with the multiplierStake scales with the multiplier — Casual player wins: 14.1%14.1%Stake scales with the multiplier — Mid player wins: 13.7%13.7%Casual players (combined)Mid player

Refusing to fire in overtime guts the rule (11.0% → 6.3% casual) — too many eliminations live there. Firing at a flat half stake is what shipped, and it works, but a flat $1,500 into ×4 values is one pot payment from re-elimination — the conveyor-belt pattern the live logs show (returns during deep overtime lasted 1–6 clues). Scaling the return stake with the current multiplier is strictly better: casual 14.1%, mid 13.7%, top shark 56.1%, at the same fire rate. It also formalizes the fix for the wrong-stake bug already filed.

Wins-since-entry is the axis; tenure is not. The obvious alternative — gate on how long somebody lasted — fails, and fails for an interesting reason: the players who never got going mostly lasted 9 to 17 clues in the ring. They die slowly while invisible, so tenure does not separate them from anybody. A tenure<12 gate scores the casuals 1.9% against 11.0% for wins<3, and a hybrid (wins<3 or tenure<12) fires for the identical 24 of 63 live first eliminations as wins<3 by itself. Tenure adds nothing.

Three is an empirical valley, not a round number. First-elimination win counts are bimodal: fourteen players at zero, eight at one, a gap at two and three, then a long tail. Three cuts at the gap. Casual benefit saturates there too — 11.0% at wins<3 against 11.7% at wins<5 — while the looser gate mostly feeds the middle of the field (10.0% to 15.2%) and hands the strongest player 50.6%. comebackGate is a setting for a host who wants the broader shape.

The whole sweep, re-run against the shipping engine with tools/trigger-study.mjs at 1,500 matches a row, boost fixed at 70% for 40 races:

GateTop sharkMidCasualsFires per match (casual)
no comeback77.8%1.5%0.5%—
wins < 172.1%3.3%5.5%0.8
wins < 267.8%6.2%9.1%1.5
wins < 3 — shipped62.9%9.9%10.8%2.0
wins < 551.1%14.8%11.3%2.6
no gate at all48.2%11.8%3.0%3.0
tenure < 12 clues76.4%2.3%1.3%0.1
wins<3 or tenure<1262.8%9.9%10.8%2.0

The hybrid row is the tell: identical to the wins gate to a tenth of a point, because tenure adds nobody the wins gate had not already caught. And the no-gate row is the shape of the whole argument — the shark falls furthest there, but the casuals collapse to 3.0%, because the free lives go to the people already winning. Lowering the shark is not the same as raising the field.

Ungated is a subsidy for the strong, and once-per-player has to stay. With no gate the elites take 1.8 free lives each per match and casual benefit collapses to 2.9%. The gate is the mechanic, not a refinement of it. And the live logs show what a re-firing trigger would produce: repeated overtime deaths at nought wins, one to six clues per extra life.

The falling ceiling eats the scaled stake — open

Found by David reading the 0.88.0 comeback path, and larger than the comeback: the ceiling clamp silently undoes stake scaling at high multipliers, and it has been doing so to entrants since scaleEntryStake shipped.

Every arrival is clamped to the current ceiling — admit(), revival, and now the comeback. That clamp was written when stakes were flat, and its comment still states the intent: a ceiling below the entry stake “would clip newcomers on arrival, which would quietly undo the thing the stake is for.” The stake then learned to ride the multiplier and the ceiling did not. Worse, the two move in opposite directions — the multiplier climbs through overtime while the roof is falling — so they are guaranteed to cross if overtime runs long enough.

At the defaults (start 3,000, ceiling 10,500, overtimeCeilingDrop 120 a clue, doubling every 6 up to ×8), taking the main-match decay as zero:

MultiplierOT clues inCeilingEntry stakeComeback stake
×269,7806,0003,000
×4129,06012,000 → 9,0606,000
×8188,34024,000 → 8,34012,000 → 8,340

So a fresh entrant stops getting the full multiple from about ×4, and a comeback from about ×8. The 3,000 floor is not what binds — the decayed ceiling itself is, and it would take roughly 62 overtime clues to reach the floor.

The argument above for not fixing it was wrong, and the sweep says so. It claimed that scaling the ceiling floor by the multiplier “kills the drain that makes the endgame resolve”. It does not. Kept here because the correction is the useful part.

Measured 2026-08-17, once node existed on the machine doing the shipping. The floor was scaled by the overtime multiplier raised to a power — 0 is the shipping behavior, 1 the obvious repair — and run through test/harness.js at 6,000 matches a row, ten times the default:

Floor exponent30p length30p p90back-halfskill (3 strongest)
0 — shipping31m35m57%32%
0.532m36m57%32%
1 — the repair35m41m56%36%

The drain survives: the repair costs about four minutes at the median and six at the ninetieth percentile, against the 53 minutes that removing ceiling decay altogether produces in the same harness. That was the stated objection and it does not hold. Half-scaling is decorative — it is inside the noise on every column, the same shape of non-result as the 50% comeback boost.

The real objection is the one that matters more. The harness scores skill as knowledge, not as buzzer speed, so the question was put to tools/comeback-study.mjs, which models real reaction times — a 95ms elite down to a 270ms casual — at 1,500 matches a row, in the shipped configuration:

Floor exponentTop sharkSecond eliteMidCasuals
0 — shipping56.9%16.5%12.3%14.3%
0.557.5%16.3%12.5%13.7%
1 — the repair63.1%15.0%11.8%10.0%

The repair takes roughly a third of the casuals' share of wins — 14.3% to 10.0% — and hands the strongest buzzer 6.2 points.

Those figures are the second set. The first, published here as 10.8% falling to 7.5%, were produced by comeback-study.mjs before anyone noticed the tool was returning a flat half stake — the pre-0.88 rule — while labeling the row SHIPPED. The engine's comeback had learned to ride the overtime multiplier and the study had not, so every row it printed understated the mechanic. The conclusion is unchanged and the direction and relative size both held on re-measurement, which is the only reason it is not a retraction. The tool is fixed; level-study and trigger-study now agree with it at 14.5%. Caught by the modeling chat, 2026-08-17.

So the ceiling clamp is an accidental leveller. It clips whoever has accumulated most, and deep in overtime that is almost always the elite; raising the floor stops the clipping and lets them keep the pile. The defect is real — entrants genuinely stop getting their full multiple from about ×4 — but this repair for it runs directly against the one axis this project cares most about, and by a wide margin. That is a much stronger reason to leave it than the one first given here, which was simply not true.

Where a real fix would have to go. Not through the roof. It would have to exempt an arrival from the clamp without raising the ceiling for accumulated scores — which means the cap at the top of resolveClue, not get ceiling.

That was then measured, one shape of it works, and it shipped in 0.90.0: nobody is clipped to a roof they never touched. An arrival lands capped by the ceiling as it stood when overtime opened, and is exempt from the per-clue cap until its score first falls to the current roof. Four candidates were measured, all reproduced here against the corrected study:

Arrival handlingTop sharkMidCasuals30p back-half
shipping clamp56.9%12.3%14.3%57%
grace of 1 clue55.6%13.1%14.6% (noise)—
grace of 4 clues55.7%13.3%14.5% (noise)—
floor rises to the arrival55.4%13.1%14.9% (noise)57%
grace until first dip, flagging every arrival54.3%13.4%16.2%60%
same, capped at the overtime-open roof54.5%13.4%16.0%60%
SHIPPED — as above, arrival-only flag55.6%12.9%15.2%57%

The last row is the one that shipped, and it is not the one that was proposed. The two middle rows flag every arrival and clear the flag on the first dip to the roof. That also exempts a player who landed comfortably under the roof and then climbed above it by winning pots — an open-ended ceiling exemption on accumulated score, which is precisely the property this whole design exists to avoid. It is a materially larger rule than its own one-line description: same number of arrivals landing above the roof, 299 against 298 in an identical harness, but 7,777 skipped clips against 1,882.

Shipping the strict reading — flag only an arrival whose landing is above the roof — takes 15.2% rather than 16.2%. Half the headline benefit came from the loose part. It is also a better trade: about a point of casual win share for about a point of back-half at twenty players and nothing at thirty, against two points bought with three or four. The point left on the table is deliberate, and widening the flag to recover it would be re-introducing the thing the ceiling-floor repair was rejected for.

Fixed-length graces are no-ops, and the reason is worth keeping. An arrival sitting above the roof loses score to every pot it does not win, so roughly half of them are back under the roof within one clue by payment rather than by clipping. Protecting that state for one clue, or four, protects something that mostly liquidates itself. The value is in never being step-function-clipped while still fighting — which is why the exemption has to run until the player first touches the roof, and why bounding the landing at the roof as it stood when overtime opened costs nothing: stakes above that reference burn down through it almost at once.

The change is real rather than cosmetic, which was the trap this had to clear. Over 400 matches and 3,317 arrivals, the shipping clamp puts zero arrivals above the roof; each candidate puts 115 above it, and 52 of those are still above the roof after the next resolveClue. The drain is untouched — length 31 to 32 minutes, p90 flat at 35.

The cost is on the other axis, and the strict version is cheap. The loose variants take back-half from 57% to 60% at thirty players and 62% to 66% at twenty. The shipped strict version costs a point at twenty and nothing measurable at thirty. Either way the direction is mechanically inevitable — protecting arrivals is protecting late arrivals — and this is the second time the two fairness axes have pulled against each other here, after the overtime return stake. Shipped on by default; the setting is arrivalGrace.

The boost is a threshold, not a dial

The 70% figure is not a setting with a comfortable middle.

It shipped at 50% for one release, on the reasonable-sounding grounds that 70% was more help than the moment called for. Measured, 50% does very nearly nothing. A discount only counts if it puts a slow player under a fast one, and against the study field's 95ms elite the three casuals need 54.8%, 60.4% and 64.8%:

medianat 50% offat 70% offneeded to beat the elite
CasualA210ms105ms63ms54.8%
CasualB240ms120ms72ms60.4%
CasualC270ms135ms81ms64.8%

At 50% not one of them reaches him and the rule is decorative; at 70% all three clear him. Below about 55% it does nothing and the useful range starts near 65%, so retune it against that table rather than by feel — feel is what produced 50%. The lesson generalizes past this setting: a parameter that has to cross another player's number to matter has no moderate position.

Superseded in 0.101.0 — the edge is a band, not a percentage. The table above is right about where the help has to land and wrong about how to get it there: a percentage takes 70% off whatever was pressed, so the help grows without limit as the press gets slower. The analysis chat replayed 249 real contested races under it and found a boosted player taking 41% of them without having pressed fastest, once by 1.56 seconds — the thing a room sees and calls rigged. comebackBand replaces comebackBoost: a returning player's press is ranked at the floor of the 60 ms band it lands in (155 ms counts as 120; a press under one band is untouched), so the edge is worth one band at most. Same replay: 12% taken without pressing fastest, never by more than 54 ms — a photo finish rather than a theft. Two returning players in the same band are split by who actually pressed first. Those replay figures are the analysis chat's and no tool in this repository has reproduced them yet, so they are quoted here as theirs; the measurements above and in comeback-study were taken under the percentage and stand as its record, not as figures for the band.

Two measurements of the same rule appear in this document and they are not in conflict. Figure 26 gives a single casual's match-win probability — 1.9% ungated against 7.3% gated. comeback-study reports the three casuals' combined share of wins, which is 2.9% ungated against 11.0% gated, with the elite falling from 74.4% at a 50% boost to 61.9% at 70%. One is a player's own chance, the other is the group's share of the table; quote whichever the question asks for, and say which.

One more correction on the record: the figures first published with this mechanic — "bottom four 1.5% to 41.9%, elite 93.8% to 38.6%" — match no row the study produces, and they have been withdrawn. A measurement nobody can reproduce is worse than no measurement, because it survives review.

The leveling budget

With the comeback shipped, the obvious next question is what else narrows the gap between an ordinary player and a shark.

Every remaining candidate went through tools/level-study.mjs at 1,500 matches a row. The answer is mostly a warning.

Each lever on its own, nothing else changed:

LeverTop sharkMidCasuals
baseline, nothing77.8%1.5%0.5%
photo finish, 25ms window75.8%2.4%0.1%
photo finish, 50ms71.8%2.3%0.3%
photo finish, 100ms63.9%4.6%0.5%
winner cooldown, sit one race47.7%10.1%2.5%
trailing player picks the category75.7%3.1%0.7%
the shipped comeback55.9%12.9%14.5%

The photo finish has a geometry problem. Sending every buzz within W milliseconds of the fastest to a random draw only randomizes players quick enough to reach the window. It flattens shark against shark and leaves the field untouched: casuals sit at or below their baseline at every width, while the top shark drops fourteen points. It is a fairness lever that redistributes strictly among the fast.

Letting the trailing player pick the category does almost nothing, for the same underlying reason the comeback works and this does not: a trailing player's problem is winning the race, not choosing which clue to lose.

The winner cooldown is the one real find, and it is not what it looks like. Sitting out the race after taking one mechanically caps anybody's share near half, and it cuts the top shark to 47.7% — the largest single effect of any non-comeback lever here. But look where the winnings go: mid 10.1%, casuals 2.5%. It does not lift the field, it promotes the second shark. Lowering the leader is not the same as raising anybody.

Levers interfere; they do not stack

This is the finding worth keeping.

Bolt a second race-structure lever onto the comeback and the casuals come out worse off, not better:

ConfigurationTop sharkMidCasuals
comeback alone55.9%12.9%14.5%
comeback + 50ms photo finish59.7%8.3%11.2%
comeback + winner cooldown47.1%9.4%12.1%
comeback + both46.3%8.5%11.3%

The interference is mechanical rather than statistical. The photo-finish window dilutes the returner's discount — a boosted press now shares a coin flip with whoever it just beat, so the thing that was bought with the boost is given away again. The cooldown benches the returner after every win, throttling the exact engine the comeback runs on. Race allocation is one budget, and the comeback already spends it; adding a second spender makes the first one poorer.

And notice which way the best-looking rows point. The two configurations that cut the top shark hardest — the cooldown rows, at 46–47% — are among the worst for the casuals. Stated as "beat the shark" that reads as progress. Stated as "give everyone a chance" it is a regression.

The metric you pick decides the design you get.

Where this lands. Do not stack race-structure levers on the comeback. What headroom remains is in channels orthogonal to the buzzer — the stables even-pot split, which is the largest economic dent in elite dominance measured here and does compose with the comeback; the gate exposed as a host dial; and format answers such as drafted teams. The winner cooldown earns one line in the notebook for one context: Tournament runs without the comeback, so the cooldown would be the only leveler on the floor, and it gets the shark to 47.7% with times still meaning exactly what they look like. Whether that suits Tournament's pure-speed identity is a design call, not a measurement.

These sections arrived as a written study rather than as raw numbers, and the figures here are the ones reproduced by re-running the shipped scripts against the current engine, not the ones delivered — they differ by a few tenths in places, and by two points for "comeback + both", which the delivery put at 9.2%.

A correction to a correction, recorded because the first one was written with more confidence than it had earned. This document said the v4 delivery was wrong to point at "a delta handed to adjustScore" on the comeback path, on the grounds that no such call exists there. In src/engine.js that is true — the comeback is a direct assignment. But the delivery was reading tools/comeback-study.mjs, where the comeback is applied through adjustScore, and where that call does clamp to the ceiling exactly as described. They had found a real divergence between the study and the engine; the reply checked only the engine and declared it imaginary. The claim it was attached to — that scaling the overtime stake "formalizes the fix" — still does not hold, because the engine's clamp survives the scaling. But the pointer was good, and following it would have caught the mislabelled figures above months of measurements earlier.

The price of aiming, and who targeting is for

Targeting shipped with a fixed price: aim at somebody, miss, and if they take the clue you pay the whole pot.

That price decided who the mechanic was for. It was not the people who needed it.

Measured over a twelve-player field at 6,000 matches a row, six ordinary players in a stable all aiming at the richest shark took 5.4% of the wins — against 17.6% for the same group holding its fire. Coordinating against the runaway leader, the one answer a room has to somebody running away with the match, was the worst thing they could do to themselves, because every aimed miss paid the entire pot and the shark wins most races. The strategy was sound. The price was not.

Figure 31 — The backfire dial — what an aimed miss pays, and who wins the match

Everyone always targets the current leader (the dominant strategy once aiming is cheap); backfire payment swept from the shipping whole-pot rule down to nothing; mutual targets pay the focused pot once, never backfire on top; stables off, 6,000 matches per row

0%15%30%45%60%Whole pot — the shipping rule (×1)Whole pot — the shipping rule (×1) — Normies (6 players): 13%13%Whole pot — the shipping rule (×1) — Sharks (2 players): 52.6%52.6%Half the pot (×0.5)Half the pot (×0.5) — Normies (6 players): 19.7%19.7%Half the pot (×0.5) — Sharks (2 players): 42%42%A quarter of the pot (×0.25)A quarter of the pot (×0.25) — Normies (6 players): 28.2%28.2%A quarter of the pot (×0.25) — Sharks (2 players): 35.1%35.1%No backfire at allNo backfire at all — Normies (6 players): 42.8%42.8%No backfire at all — Sharks (2 players): 24.4%24.4%Normies (6 players)Sharks (2 players)

Weakening the backfire rule is the largest casual-side lever measured in this project — the normies walk from 13.0% up to 42.8% while the sharks fall from 52.6% to 24.4%, and the dial is usably smooth in between. The voltron's coordinated-targeting strategy tells the same story (5.5% → 17.3% → 34.4% → 30.4% across the same dial). The cost is in kind, not degree: without a real backfire price, holding a target is a free option, permanent leader-focus becomes the dominant strategy, and the score lead itself turns into a liability — a game where the best move may be staying second. These runs model naive always-on targeting, not players adapting to that incentive. Decided (Aug 17): Arcade and chaos games default backfire off; Tournament defaults to ×0.5; the host holds the dial in every mode; and backfire never stacks with being the winner's own target — a mutual pair pays the focused pot, once.

So the price is a dial. targetBackfire is the share of the pot an aimed miss pays, and the defaults follow the mode:

Aimed miss paysStag, everyone aims at the leaderStable aims at the richest shark
the whole pot — the original rulesharks 52.7% · ordinary 13.1%sharks 66.9% · ordinary 5.4%
half — Tournamentsharks 42.1% · ordinary 19.3%sharks 48.2% · ordinary 17.2%
nothing — Arcade and Chaossharks 28.0% · ordinary 39.4%sharks 38.0% · ordinary 31.0%

Arcade and Chaos set it to nothing, so ganging up is affordable and the room gets its answer back. Tournament keeps half, because there is no comeback there to soften a bad match and a free-to-coordinate leader hunt would simply decide it. The host can move the dial anywhere in any mode, and the whole pot is still on it.

Aiming back is neither a discount nor a surcharge. If the winner's own target also aimed at the winner, that player owes the focused pot once — never the focused pot plus a backfire on top. And turning backfire off does not turn off focused fire: a winner who aimed still lands the whole pot on one head.

The accepted cost is that cheap backfire makes permanent leader-focus the dominant strategy and a score lead a liability, which shortens matches under heavy targeting — 117 clues to 105 in the stag rows. The figures model naive always-on aiming rather than players adapting, which is the honest caveat and the reason Tournament does not follow Arcade here.

One row of the delivered study did not reproduce at first, and the cause turned out to be worth more than the row. The backfire-off stag case came back at 28.0% / 39.4% here against 24.4% / 42.8% as delivered — and the two runs were on different engines. arrivalGrace shipped on by default midway between them, and no study pinned it, so each run inherited whatever its engine defaulted to. Both sets of numbers reproduce to the tenth once the setting is fixed. Every other row matched all along because under a real backfire price the grace barely moves this study; the whole divergence lives in the single configuration that combines a free aimed miss with universal leader-targeting.

That corner has a finding in it. With backfire off and everybody hunting the leader, the arrival grace hands about 3.4 points back to the strongest players and stretches matches from 97 clues to 105. The likely mechanism is that a graced return lands on its overtime-scaled stake above the decayed roof, which makes it instantly the richest player on the board — and in a kill-the-leader equilibrium the protection is a bullseye, so the focused pot lands on exactly the player it just protected. That reading is plausible rather than proven; the attribution to the setting is exact. Nothing to act on at the shipped defaults, but worth remembering if a preset ever pairs cheap targeting with heavy overtime re-entry.

The draw slot is the biggest lever measured

Everything else in this document moves the casual row by a point or two. Where a strong player happens to be drawn moves it by tens. Twelve-player field, 6,000 matches a row, stables off:

Where the two strongest players are drawnTheir combined share of wins
random — the reference55.0%
opening together, slots 1 and 234.3%
early, slots 4 and 537.3%
middle, slots 6 and 741.2%
last, slots 11 and 1289.8%

The late-slot advantage is generic, not a shark effect. Put two merely average players in the last two slots and they take 69.7% between them. The draw is doing the work: an early entrant spends the whole match being ground down by a full ring, while a late one walks into a thinned field holding a fresh stake. Same asymmetry the ceiling and the entry interval were tuned against, seen from the other end — and much larger than either.

Nobody pulls this lever. The shuffle pulls it for you.

It is also a flow-rate phenomenon. With the two strongest players opening together, feeding the rest in every 5 clues takes them to 20.9%; every 15 clues gives them 36.6%. Fast entry keeps the ring full, which is what grinds an early leader down. The same dial pointed the other way — strong players drawn last — hands them 86.7% at interval 5 and 94.3% at 15.

Figure 32 — What the stable does to the match, by configuration

Combined match-win probability of the six normies against the two sharks; sharks in their own two-stable except where noted

0%20%40%60%Normie voltron only (sharks stag)Normie voltron only (sharks stag) — Normies (6 players): 24.6%24.6%Normie voltron only (sharks stag) — Sharks (2 players): 46.5%46.5%Voltron vs shark stable, share=surplusVoltron vs shark stable, share=surplus — Normies (6 players): 18.4%18.4%Voltron vs shark stable, share=surplus — Sharks (2 players): 59.7%59.7%Voltron vs shark stable, share=evenVoltron vs shark stable, share=even — Normies (6 players): 17.3%17.3%Voltron vs shark stable, share=even — Sharks (2 players): 60.4%60.4%Voltron vs shark stable, share=winnerVoltron vs shark stable, share=winner — Normies (6 players): 16.4%16.4%Voltron vs shark stable, share=winner — Sharks (2 players): 58.7%58.7%Voltron + mutual targeting warVoltron + mutual targeting war — Normies (6 players): 15.5%15.5%Voltron + mutual targeting war — Sharks (2 players): 38.9%38.9%Baseline — everyone stagBaseline — everyone stag — Normies (6 players): 9%9%Baseline — everyone stag — Sharks (2 players): 55%55%Voltron targets the richest sharkVoltron targets the richest shark — Normies (6 players): 5.5%5.5%Voltron targets the richest shark — Sharks (2 players): 66.9%66.9%Shark stable only (no voltron)Shark stable only (no voltron) — Normies (6 players): 5.5%5.5%Shark stable only (no voltron) — Sharks (2 players): 70.3%70.3%Normies (6 players)Sharks (2 players)

The voltron is real: pot immunity plus shared winnings take the six normies from 9% (everyone stag) to 17.3% against a shark stable, and 24.6% when the sharks stay stag — while a shark stable with no answer to it is the worst measured configuration (normies 5.5%, sharks 70.3%). Share mode barely matters (16.4–18.4%). The trap is targeting: pointing six buzzers at the richest shark hands the match back to the sharks (5.5%), because every normie aiming at a shark who then wins the race pays the whole pot — and the sharks win most races. The backfire rule is the shark's bodyguard.

Figure 33 — How often the voltron clears the ring as a unit

Matches in which every live player left standing is a voltron member — the stable's own victory condition, after which it dissolves and members settle it between themselves

0%5%10%15%20%25%Voltron only (sharks stag)Voltron only (sharks stag) — Voltron clears the ring: 22.6%22.6%Vs shark stable, share=surplusVs shark stable, share=surplus — Voltron clears the ring: 18.1%18.1%Vs shark stable, share=evenVs shark stable, share=even — Voltron clears the ring: 16.2%16.2%Vs shark stable, share=winnerVs shark stable, share=winner — Voltron clears the ring: 14.6%14.6%Vs shark stable + mutual targetingVs shark stable + mutual targeting — Voltron clears the ring: 12.6%12.6%Targeting the richest sharkTargeting the richest shark — Voltron clears the ring: 4.5%4.5%

Answering the question as asked — "do they defeat them" as a team — the voltron sweeps the ring in about one match in six against a shark stable (16.2%), one in four-and-a-half when the sharks go stag. A caveat that is really a finding: the staged release plus the join cap (a stable may hold at most half the live ring) means the six-normie voltron never fully assembles in a typical match — it peaks at about 3.5 members, because by the time the ring is big enough to admit six, somebody is already gone.

Figure 34 — Shark win share by draw position — earlier is weaker

Two sharks placed at fixed draw slots, everyone else shuffled; entry interval 10 clues (the auto value for this room); stables off

0%25%50%75%random draw 55%Sharks open together (slots 1–2)Sharks open together (slots 1–2) — Sharks' combined win probability: 34.8%34.8%Sharks early (slots 4–5)Sharks early (slots 4–5) — Sharks' combined win probability: 37.4%37.4%Sharks middle (slots 6–7)Sharks middle (slots 6–7) — Sharks' combined win probability: 41.7%41.7%Sharks split (slots 4 and 9)Sharks split (slots 4 and 9) — Sharks' combined win probability: 59.5%59.5%Sharks split (slots 1 and 12)Sharks split (slots 1 and 12) — Sharks' combined win probability: 80.1%80.1%Sharks last (slots 11–12)Sharks last (slots 11–12) — Sharks' combined win probability: 89.6%89.6%

The draw is the biggest lever measured in any of these studies. Sharks opening together — the configuration that feels most dangerous at the table — is the minimum, 34.8%: an early shark spends the whole match pinned to the ceiling, paying pots, exposed to every elimination risk the game has. Sharks in the last two slots win 89.6% — they arrive with overtime-scaled stakes into a thinned, tired field. Splitting them (1 and 12) is nearly as bad as both late, and the split's early shark wins only 15.2% while the late one takes 65%: same player, same buzzer, opposite end of the draw.

Figure 35 — Flow rate — the wishbone is a slow-flow phenomenon

Sharks' combined win probability by entry interval, sharks first versus sharks last

0%25%50%75%100%51015entry interval, clues between arrivals (fast flow → slow flow)Sharks last (slots 11–12) — 5: 86.7%Sharks last (slots 11–12) — 10: 89.6%Sharks last (slots 11–12) — 15: 93.9%Sharks lastSharks first (slots 1–2) — 5: 20.9%Sharks first (slots 1–2) — 10: 34.8%Sharks first (slots 1–2) — 15: 37.5%Sharks first

The observed pathology — two early sharks picking off a slow trickle of arrivals one at a time — is real, and it is the slow-flow variant of the early line: at interval 15 the early sharks climb back to 37.5%, and each arrival faces the pair nearly alone (2.5 players per match are eliminated having taken ≤1 clue). Feeding players in fast is what breaks the wishbone: at interval 5, early sharks fall to 20.9%, the best anti-shark configuration measured. Note the two lines never cross — flow rate tunes the early-shark case but nothing rescues a late shark.

Figure 36 — The control — the last two slots win for whoever holds them

Combined win probability of whichever pair is placed in slots 11–12, everyone else shuffled; interval 10

0%25%50%75%Sharks hold slots 11–12Sharks hold slots 11–12 — That pair's combined win probability: 89.6%89.6%Average players hold slots 11–12Average players hold slots 11–12 — That pair's combined win probability: 69.3%69.3%Normies hold slots 11–12Normies hold slots 11–12 — That pair's combined win probability: 22.5%22.5%

The late-draw advantage is generic, not shark-specific: even two normies holding the last slots more than double their tier's share (22.5% against 9.0% under a random draw), and two average players there win 69.3% of matches. Skill multiplies the slot, it does not create it. Which is the whole recommendation: give the late slots to the weakest players you have, and put your sharks in first.

A stable of ordinary players is real, and it never fully assembles. Six of them banding together takes the group from 9.1% to 25.4% against stag sharks, which is the largest economic effect measured here. But the join cap means it peaks at about three and a half members in practice, so what the chart shows is the ceiling of the idea rather than its typical case.

None of this is currently a lever the host can pull: the draw is random by design, and making it otherwise would be seeding a battle royale. It is recorded because it sets the scale — any balance change worth arguing about has to be measured against a background that already swings this far on nothing but the shuffle.

Handbook section for merge · Part IV · rev. Aug 31, 2026

The bone pile: ideas we tried, measured, and buried

This handbook keeps its wrong answers on the page, marked as wrong. This section is that rule applied to design. Every leveling idea below sounded good, most of them sounded good to me, several got proposed more than once, and every one of them got measured and lost. Here they are, with the number that killed each one.

Why keep a graveyard? Because re-litigating a dead idea costs a study, and because the shape of a failure is the useful part. If a new idea reduces to one of these shapes, it owes us a reason it will measure differently. "It feels like it should work" is how every one of these started.

Figure 37 — the bone pile, by the only number that matters

Casual players' combined match-win probability under each rejected lever, six-player study field (two elites, one mid, three casuals), 1,500–6,000 matches per row; reference line is the shipped gated comeback

0%5%10%15%shipped comeback 15.2%Race levers stacked on the comebackRace levers stacked on the comeback — Casual players' match-win probability: 11.3%11.3%Ceiling-floor scaling repairCeiling-floor scaling repair — Casual players' match-win probability: 10%10%Warm-up start boost (boost everyone)Warm-up start boost (boost everyone) — Casual players' match-win probability: 4.6%4.6%Permanent 80% buzz boostPermanent 80% buzz boost — Casual players' match-win probability: 4.5%4.5%Ungated comebackUngated comeback — Casual players' match-win probability: 3%3%Winner cooldown aloneWinner cooldown alone — Casual players' match-win probability: 2.5%2.5%Tenure-based comeback gateTenure-based comeback gate — Casual players' match-win probability: 1.3%1.3%Trailing player picks categoryTrailing player picks category — Casual players' match-win probability: 0.7%0.7%Photo-finish window (100ms)Photo-finish window (100ms) — Casual players' match-win probability: 0.5%0.5%Kickout on 2 (boost after 10 lost races)Kickout on 2 (boost after 10 lost races) — Casual players' match-win probability: 0.3%0.3%

Every bar is a measured configuration, at current-engine levels where we re-ran them (the race-lever, warm-up and permanent-boost rows are the prior generation's numbers; the story didn't move). The comeback's 15.2% is the bar to clear. Nothing here gets two-thirds of the way there, and the best of the bunch — race levers stacked on top of the comeback — is a case of adding help and getting less.

Buried because the trigger measures the wrong thing

Three ideas died the same death: they fired on a signal that doesn't separate the player who never got going from the player who's just slower.

Kickout on 2 — a 70% buzzer boost after ten lost races — moved the casual row from 0.5% to 0.3%. Losing races isn't a distress signal. Everybody below the fastest buzzer loses races constantly, so the boost shows up late, to everyone, after the pot has already emptied their pockets. Tenure gates (time in the ring instead of clues taken) are the same mistake on the other axis — 1.3% — because time served says nothing about whether you ever got to play. The warm-up start boost had two things wrong with it at once: warm-up buzzes are biased (everyone reads slow for their first half-dozen presses — 302 and 242ms for players who settle at 85 and 60), so a warm-up classifier boosts the whole room, which is worth 4.6%; and warm-ups are free, so the signal is free to fake. The comeback's trigger survives for exactly the opposite reason: getting eliminated with fewer than three clues taken is expensive. Sandbagging it means actually not scoring. You can't fake your way into it.

Buried because lowering the shark is not raising the field

This is the most common way a good-sounding idea fails, and it's why this handbook judges every candidate by the casual row and never the shark row. The winner cooldown is the biggest solo anti-shark lever we've ever measured — top shark 79.5% → 49.2% — and it leaves the casuals at 2.5%. Every win it confiscates goes to the second-fastest buzzer. You didn't level the field; you crowned a different shark. The ungated comeback (3.0%) and a permanent 80% buzz boost for everyone (4.5%) fail identically: help the strong can also use is a subsidy for the strong. The photo-finish window (0.5%) has a geometry problem on top of that — it only randomizes among players fast enough to reach the window, which is precisely the players who need no help. The trailing-player category pick (0.7%) is real, and too small to matter.

Buried because you can't spend the race budget twice

Add a photo-finish window or a winner cooldown on top of the comeback and the casuals do worse than with the comeback alone: 14.5% falls to 11.2–11.3%. Race allocation is one budget. The comeback already spends it on the players the gate picks out, and a second race-structure lever takes races away from them as often as it hands them over. The same interference killed the warm-up start boost as a supplement (oracle-perfect boosting stacked on the comeback: 12.7%, below the comeback by itself). One race lever at a time. Pick the one that works.

Buried because the repair was worse than the bug

The overtime ceiling clamp really does eat the arrival-stake scaling. That defect is real and it's documented. Scaling the ceiling floor with the multiplier — the obvious fix — hands the strongest buzzer 5.6 points and takes the casuals from 14.3% to 10.0%, because the clamp it removes is an accidental leveler that mostly binds on whoever has piled up the most. (An earlier claim that this repair "kills the drain" was wrong and is corrected above; it costs about four minutes, not the match.) Fixed-length arrival graces (1 or 4 clues) and floor-rises-to-arrival are measured no-ops — a protected arrival above the roof mostly pays itself back under it within a clue anyway. The one variant that did anything was landing-above-the-roof protection, which shipped in 0.90.0 as arrivalGrace at the strict reading (casuals 15.2%, back-half +1). The loose every-arrival reading scored higher (16.2%) because it quietly smuggled back the open-ended exemption we'd just rejected. The missing point is deliberate. We'll take the smaller number that's the right rule.

Buried, then dug back up: the whole-pot backfire

From the twelve-player stable studies: six casual-tier players in a stable pointing every buzzer at the richest shark was the worst thing they could do under the old whole-pot backfire rule — 5.5%, against 17.3% for the same stable holding its fire. An aimed miss paid the entire pot, and the shark wins most races. The strategy wasn't wrong. The price was. That finding didn't stay a rejection; it became the case for the backfire dial (Arcade and Chaos default off, Tournament ×0.5, mutual targets pay the focused pot once), which has its own section. It stays on this list as the cleanest example we have of a rule quietly deciding who a mechanic is for.

Not a rejection — a warning: some dials are thresholds

One near-miss belongs here. The comeback boost shipped at 0.5 for one release on the perfectly reasonable-sounding argument that 70% was more help than the moment needed. Measured, 0.5 is not half of 0.7. It's about a tenth: the casual row fell from 15.2% to 4.9%. The discount only counts if it puts a slow press under a fast one, and at 50% none of the study's casuals crossed the elite. Some settings are thresholds wearing a dial's clothing, and the only way to find out which is to measure both sides of the knee. Retune against the table. Never by feel.

Sources: kickout-study, trigger-study, level-study (1–2), the arrival-exemption run, and the voltron/backfire studies — six-player rows on the calibrated study field, twelve-player rows noted as such; 1,500–6,000 matches per configuration; every figure reproduced in this project's own runs. House rules: sections not file swaps; corrections recorded, not deleted; handles only.

Two belts in one night, and the buzzer's secret

Matches 21 and 22 ran on the same evening, in front of the same host, with the largest field yet recorded — ten players. One person won both. P6 took the Arcade match with 13 of 82 races and then took the Tournament match with 32 of 73, and those two wins have almost nothing in common.

That is the mode system doing exactly what it was built to do.

Figure 38 — Win Arcade like an Arcade player, win Tournament like a Tournament player

Contested-clue wins per player; eleven players appear across the two matches (P27, P36 and P1 played only M21, P41 only M22)

0102030P6P6 — M21 · Arcade (comeback on): 1313P6 — M22 · Tournament (comeback off): 3232P14P14 — M21 · Arcade (comeback on): 1111P14 — M22 · Tournament (comeback off): 1818P30P30 — M21 · Arcade (comeback on): 1212P30 — M22 · Tournament (comeback off): 11P41P41 — M21 · Arcade (comeback on): 00P41 — M22 · Tournament (comeback off): 1212P38P38 — M21 · Arcade (comeback on): 99P38 — M22 · Tournament (comeback off): 00P27P27 — M21 · Arcade (comeback on): 77P27 — M22 · Tournament (comeback off): 00P13P13 — M21 · Arcade (comeback on): 66P13 — M22 · Tournament (comeback off): 33P39P39 — M21 · Arcade (comeback on): 55P39 — M22 · Tournament (comeback off): 00P40P40 — M21 · Arcade (comeback on): 55P40 — M22 · Tournament (comeback off): 00P36P36 — M21 · Arcade (comeback on): 44P36 — M22 · Tournament (comeback off): 00P1P1 — M21 · Arcade (comeback on): 22P1 — M22 · Tournament (comeback off): 00M21 · Arcade (comeback on)M22 · Tournament (comeback off)

In the ten-player Arcade match every single player won at least two races and the winner needed just 13 of 82 — P6 won on staying power, banking above the roof three times in the drain. In Tournament the same player took 32 of 73 races at a 127ms settled median, the fastest in the room, and the fastest settled buzzer has now won all three Tournament matches on record. Arcade winners since the new rooms arrived: zero for six by settled speed at that point, and zero for eight by match 24. The mode still picks how you have to win. P6 is just the first player to win both ways in one evening.

In a ten-player Arcade match every single player won at least two races, and the winner needed thirteen of eighty-two. P6 won that one on staying power, banking above the roof three times in the drain. In Tournament the same player took thirty-two of seventy-three at a 127ms settled median, the fastest in the room. Every Tournament match on record has gone to the fastest settled buzzer — three for three. Arcade since the new rooms arrived: zero for six as of this batch, and still zero eight matches in.

Half the room was timing the read the whole time

Version 0.96.1 started counting presses that land under 150ms — faster than a human can react to a light. The count settled an argument. That sub-100ms reading back in match 18 was not an outlier, and it was never one unusual player.

Figure 39 — Half the room is timing the read, and they always were

Share of live buzzes landing under 150ms — faster than human reaction to the window — per recorded match; orange marks the two newest communities

0%20%40%60%M8M12M16M20M22recorded matchlive buzzes under 150msM8: M8, 55.4%M8M9: M9, 49.3%M9M10: M10, 43.9%M10M11: M11, 31.5%M11M12: M12, 43.4%M12M13: M13, 40.8%M13M14: M14, 1.2%M14M15: M15, 5.8%M15M16: M16, 2.8%M16M17: M17, 42.9%M17M18: M18, 26.8%M18M19: M19, 40.8%M19M20: M20, 49%M20M21: M21, 48%M21M22: M22, 48.7%M22

Engine 0.96.1 started counting presses under 150ms, and the count settles an open question. M18's sub-100ms outlier wasn't an outlier. In every match played by the experienced rooms, 27–55% of live presses land under 150ms — that's cadence-timing, pressing on the rhythm of the read instead of reacting to the window. It goes back to match 8. Only the two newest communities (M14–M16, in orange) buzz almost entirely by reaction. The practical upshot: a "settled median" mixes two different skills now, and one M21 player who reads 72ms overall reads 667ms on the presses where the timing missed. Fast hands are real. But the fastest numbers are rhythm.

In every match played by the experienced rooms, 27–55% of live presses come in under 150ms. That is not reaction. That is cadence — pressing on the rhythm of the read rather than waiting for the window — and it goes back to match 8. Only the two newest communities buzz almost entirely on reaction.

Which means a settled median now measures two different skills at once, and the document should say so plainly. One M21 player reads 72ms overall and 667ms on the presses where the timing missed. The M20 figure of 46ms is 965ms on its misses. Fast hands are real and some of these people have them. But the fastest numbers are rhythm.

Time to eat a correction: the overtime tail

The last batch recommended budgeting about twenty clues for the overtime drain. That recommendation came from three matches, and the full record shows those three were the three shortest tails ever recorded — 17, 19 and 26.

Figure 40 — The overtime tail: 17–34 clues, median 29. Time to eat a correction.

Clues from the entry queue emptying to the final wipe, for all eleven revival-off matches on record

0102030M9 (4p)M9 (4p) — Overtime tail, clues from queue-empty to the wipe: 1818M11 (11p)M11 (11p) — Overtime tail, clues from queue-empty to the wipe: 3030M14 (6p)M14 (6p) — Overtime tail, clues from queue-empty to the wipe: 2929M15 (6p)M15 (6p) — Overtime tail, clues from queue-empty to the wipe: 2626M16 (6p)M16 (6p) — Overtime tail, clues from queue-empty to the wipe: 3030M17 (8p)M17 (8p) — Overtime tail, clues from queue-empty to the wipe: 1919M18 (3p)M18 (3p) — Overtime tail, clues from queue-empty to the wipe: 2929M19 (7p)M19 (7p) — Overtime tail, clues from queue-empty to the wipe: 2626M20 (7p)M20 (7p) — Overtime tail, clues from queue-empty to the wipe: 1717M21 (10p)M21 (10p) — Overtime tail, clues from queue-empty to the wipe: 3434M22 (8p)M22 (8p) — Overtime tail, clues from queue-empty to the wipe: 2929

Last batch I recommended budgeting a ~20-clue overtime tail, measured from three matches. The full record says those three (17, 19 and 26) were the three shortest tails we've ever had. Across all eleven revival-off matches the tail runs 17–34 clues with a median of 29, and it doesn't track field size cleanly — the 4-player match drained in 18, a 6-player match in 30. I'd have shipped a number that was 30% short. The estimator itself had a great night anyway: M21 predicted 82 clues and got 82; M22 predicted 75 and got 73, with both rooms pacing 13.5–15.1 seconds a clue, faster than the 17.5 assumption for the first time on record.

Across all eleven revival-off matches the tail runs 17 to 34 clues, median 29, and it does not track field size cleanly: a four-player match drained in 18 and a six-player one in 30. A number 30% short would have shipped. The estimator now budgets 29, and only when revival is off — revival refills a ring the drain has already emptied, and those four matches ran tails of 28, 31, 47 and 72. That is a different measurement and it keeps the number it was fitted with until somebody measures it properly.

The estimator had a fine night regardless. M21 predicted 82 clues and got 82; M22 predicted 75 and got 73. Both rooms paced 13.5–15.1 seconds a clue — under the 17.5-second assumption for the first time on record. Rooms are getting faster as their hosts get fluent.

Two things stayed stubbornly true. Stables were switched on for the ten-player match — precisely the field size the join cap was designed around — and nobody formed one. Targeting was available in M21 and nobody aimed. Stables, targeting and bounties have now gone twenty-two matches without a single use between them. That is not a balance problem. Nobody can find them.

Sources: match logs for the twenty-two recorded human matches, the 0.96.1 anticipation instrumentation, and the eleven revival-off tails. Figures supplied pre-anonymized; P-labels and host letters as delivered.

Somebody finally jumped

Matches 23 and 24 are the two biggest fields on record — eleven players, then ten, the same night, with the hosts trading places: Host B ran the first and then walked next door to play in the second. Both went to players who had never won before. And after twenty-two matches of nobody laying a finger on an optional mechanic, the top rope got jumped. Four times.

Figure 41 — Big fields spread the races wide — and the winner isn't the fastest

Contested-clue wins per player; fourteen humans appear across the two matches (P11 and P36 played only M23; P8, P45 and Host B only M24)

0510152025M23 · 11 players · Host B's roomM24 · 10 players · Host C's roomP442013P431913P42519P11230P8113P2979P3049P3907P3650P3833P1353P4505Host B (played M24)06

Eleven players is the largest room the game has hosted, and the comeback kept it honest: only two player-matches ended raceless (both in M23, and both of those players won races in M24 an hour later). P11 took the most races in M23 — 23 of 98 — outlasted everyone but the winner, and still lost: races won and staying power are different currencies in Arcade. The fastest settled buzzer in both matches was P30 (54ms and 63ms), who won neither. That's now zero for eight for the fastest buzzer in Arcade since the new rooms arrived, against 12 of 13 in the shark era and 3 of 3 in Tournament. The mode predicts the winner. Twenty-four matches in, it hasn't missed.

Eleven players is the largest room the game has hosted, and the comeback kept it honest: only two player-matches ended without a single race won, and both of those players won races an hour later in the other room. P11 took the most races in M23 — 23 of 98 — outlasted everyone but the winner, and still lost. Races won and staying power are different currencies in Arcade.

The fastest settled buzzer in both matches was P30, at 54ms and 63ms, and P30 won neither. That is zero for eight for the fastest buzzer in Arcade since the new rooms arrived, against 12 of 13 in the shark era and 3 of 3 in Tournament.

Twenty-four matches in, the mode has not once failed to predict the shape of the win.

The tail correction met the two hardest tests available

Three weeks ago the estimator ran about 20% long in experienced rooms — 100 predicted against actuals of 78, 85 and 76. The overtime tail was corrected from roughly 20 clues to 29, and the first two matches on the new model were the two biggest rooms ever recorded.

Figure 42 — The overtime-tail fix shipped, and the estimator hit the two biggest rooms ever

Predicted versus actual clues, recent revival-off matches; the tail correction (≈20 → ≈29 clues) landed between M22 and M23

0255075100Predicted cluesActual cluesM17 (8p) · old tail10078M19 (7p) · old tail10085M20 (7p) · old tail10076M21 (10p)8282M22 (7p)7573M23 (11p) · new tail9498M24 (10p) · new tail105106

Three weeks ago the estimator ran about 20% long in experienced rooms — 100 predicted against 78, 85, 76. The tail correction went in, and the first two matches on the new model were the two hardest tests available: predicted 94 against 98 actual, and 105 against 106. Overtime tails ran 36 and 35 clues, right on the corrected band (17–34, median 29 — these two nudge it to 17–36). Pace medians were 16.4 and 17.2 seconds per clue, the closest to the 17.5 assumption any rooms have run. When the model is right, the predictions get boring. Boring is the goal.

Predicted 94 against 98 actual, and 105 against 106. The tails themselves ran 36 and 35 clues, which nudges the recorded band to 17–36 with the median holding at 29. Pace came in at 16.4 and 17.2 seconds a clue, the closest any rooms have run to the 17.5-second assumption.

When the model is right, the predictions get boring. Boring is the goal.

Every fast number in the room is two numbers

Figure 43 — Every fast number in the room is two numbers

Settled median (all presses after warm-up) against reaction-only median (presses ≥150ms), per player, both matches pooled; ~46% of live presses landed under 150ms

0ms100ms200ms300ms400msSettled median (all presses)Reaction-only median (presses ≥150ms)P3063ms366msP877ms280msP1186ms226msP43 won M24105ms246msP38110ms251msP44 won M23126ms346msHost B (played)158ms230msP29159ms282msP42193ms215msP13236ms406ms

This is the anticipation finding made visible. P8 reads 77ms settled and 280ms on the presses where the timing missed — four of every five of his presses ride the rhythm. P44 won M23 on the same trade: 126ms settled, 346ms reaction-only, and 27 locked-out presses in one match paid as the cost of playing the beat. The gap between a player's two dots is how much of their speed is timing; the right-hand dot is the hand they were born with. Note P42, whose dots nearly touch (193 vs 215) — a pure reaction player, new this week, and it earned him 19 races in M24. Both styles win races. Only one of them shows up honestly in a single median.

This is the anticipation finding made visible. P8 reads 77ms settled and 280ms on the presses where the timing missed — four of every five of his presses ride the rhythm. P44 won M23 on the same trade: 126ms settled, 346ms reaction-only, and 27 locked-out presses in one match, paid as the cost of playing the beat.

The gap between a player's two dots is how much of their speed is timing. The right-hand dot is the hand they were born with. P42's two dots nearly touch, at 193 and 215 — a pure reaction player, new this week, and it earned him 19 races. Both styles win. Only one of them shows up honestly in a single median.

The top rope is live: it pays double, and it drains double

Four jumps, the first four ever recorded. Host B, playing rather than hosting, jumped three times: a 400 clue that paid 800, a 100 clue that paid about 300 through the pot, and one doubled loss. Host B went out first anyway — the highlight-reel price of showing a room how a mechanic works. The other jumper ate a doubled stake on somebody else's race and was eliminated two clues later.

So the early book: the top rope does exactly what this document says it does, in both directions, and both jumpers who lost their jump were gone within two clues. Four uses is not a measurement and nobody should treat it as one. What it settles is narrower and more useful — the mechanic works in live play, and the reason it had never been used was never that it was broken.

Stables and targeting were switched on in both matches. Twenty-four matches in, neither has ever been used, and neither have bounties. One mechanic broke the seal, and it broke it because the host demonstrated it at the table. That is the entire finding, and it is a lesson about teaching rather than about balance.

Sources: match logs for matches 23 and 24, Sep 14 2026, engine 0.96.1; a robot lobby from the same night is excluded as a test. Settled medians exclude each player's first six buzzes. All 33 above-roof rows across both matches are overtime banking by the clue winner, including the two top-rope wins, which banked ceiling-free as designed.

Method, in one paragraph: the rules engine is a plain module with no framework and no network, driven by a Monte Carlo harness that replays hundreds of matches per configuration before every deployment. The harness once caught a scoring regression that would have stretched matches from 57 minutes to 746 — invisible in ordinary testing, fatal on a Saturday night. Clue material comes from a public dataset, and match logs — every clue, every buzz, every correction — supply nearly every number in this document.

J! Royal Rumble · j-royal-rumble.net · The winner is the only result. The rest is an argument for the group chat.