Newest first. npm run ship adds an entry automatically, so this stays current without anybody remembering to update it.
Two public data leaks, both from a guard that defaulted to open. Found in an external review by Colin Davy against 0.94.2; verifying his findings turned up a third and worse one the review missed.
/api/logs was public. It listed all 52 saved matches, and /api/logs/<file> served a complete log to anybody who asked — in-game handles, every buzz time, every answer. The guard existed and returned true when RUMBLE_LOG_KEY was unset, on the reasoning that an unkeyed deployment is a test deployment. The live site ran that way for months. This project anonymizes players to P-labels in the handbook and keeps legal names out of the repo precisely to avoid that, and one route undid all of it.
The admin key had a published default. ADMIN_KEY fell back to the literal string 'daymay', committed to a public repo, so /api/control?key=daymay returned 200 on the live site: list every match, download every log and report, and end a game in progress. It read as protected because an unauthenticated request correctly returned 403 — which is where the review stopped, and why this one is worth remembering as the shape of a near miss.
Both now fail closed with no default, and the server warns loudly at boot when either key is missing rather than failing silently. RUMBLE_ADMIN_KEY and RUMBLE_LOG_KEY are now required in the service environment.
The admin key also satisfies the log guard. The control room downloads individual logs and authenticates with the admin key — a dependency that was invisible while the guard was open, and would otherwise have locked the host out of their own match records.
/api/health is thin in public: status and version only. The full body still goes to 127.0.0.1, because deploy-remote.sh reads matchesInPlay there before restarting and ending people's games. The local check reads socket.remoteAddress, never req.ip, so it cannot be spoofed with X-Forwarded-For if trust proxy is ever switched on. Version stays public because /history prints it on a keyless page anyway.
Also: helmet for the headers the site was sending none of — HSTS, nosniff, frameguard, referrer policy — and x-powered-by disabled. CSP is off on purpose: every page is one self-contained file with inline script, and a default CSP stops the buzzer working. Enabling it properly means extracting and hashing every page's script, which is its own change.
test/security.mjs runs against a deliberately keyless server on its own port and asserts the refusals, including that key=daymay is dead. It is a separate CI step because the main test server runs *with* keys.
arrivalGraceThe one backfire row that would not reproduce across the two chats is solved, and the answer was a fourth instance of the same trap: neither run pinned arrivalGrace. It shipped on by default in 0.90.0, between their measurement and ours, so each run inherited its own engine's default. Both sets of numbers were right — 28.0%/39.4% with the grace on, 24.4%/42.8% with it off, each reproducing to the tenth once the setting is fixed. The analysis chat isolated it and their probe reproduces here exactly.
All nine study tools now pin arrivalGrace and targetBackfire explicitly. Every anchor still holds after pinning: the voltron ×1 row at 66.9/27.7/5.4, the shipped comeback at 15.2% casuals, sharks-last at 89.8%.
A small finding falls out of the corner that caused it. With backfire off and everybody hunting the leader, the arrival grace hands about 3.4 points back to the strongest players and stretches matches from 97 clues to 105 — a graced return lands on its overtime-scaled stake above the decayed roof, making it instantly the richest player, and in a kill-the-leader equilibrium the protection is a bullseye. The mechanism is plausible rather than proven; the attribution to the setting is exact. Nothing to act on at shipped defaults, but it matters if a preset ever pairs cheap targeting with heavy overtime re-entry.
The handbook now records this as resolved rather than open.
The rejected buzz-boost-after-losses mechanic is called kickout on 2, not "pity powerup". David's rename. Same instinct as calling the field casual players rather than low-skill: "pity" names the player as pitiable. The wrestling term also fits the format better — you kick out at two rather than being counted out.
Renamed in the handbook prose, in three chart figures including their tooltips and legends, and in the version history. tools/pity-study.mjs is now tools/kickout-study.mjs, since the handbook cites it as a source and a document saying "kickout on 2" while pointing at pity-study is the kind of drift this project keeps getting bitten by.
test/guides.mjs asserts the new name is present and the old one is absent. The analysis folder still ships its copy as pity-study.mjs, so the old name can arrive again in a delivery — that guard is what catches it.
An aimed miss no longer always pays the whole pot. targetBackfire is the share it pays: Arcade and Chaos default to 0, Tournament to 0.5, host-adjustable in every mode, and 1 reproduces the old rule exactly. David's decision.
The whole-pot price quietly decided who targeting was for. Over a twelve-player field at 6,000 matches a row, six ordinary players in a stable all aiming at the richest shark took 5.4% of the wins — against 17.6% for the same group holding its fire. Ganging up on a runaway leader is the only answer a room has to somebody running away with the match, and the price made it self-harm.
Mutual targeting does not stack: if the winner's own target aimed back, they pay the focused pot once, never that plus a backfire. Turning backfire off does not turn off focused fire. Nine engine tests cover the dial, including f=1 as a byte-for-byte regression anchor against the old rule.
The accepted cost, recorded with the decision: cheap backfire makes permanent leader-focus the dominant strategy and a lead a liability, shortening matches under heavy targeting (117 clues to 105). The figures model naive always-on aiming, not players adapting — which is why Tournament keeps a real price.
Six figures merged, including the draw-slot studies: where a strong player is drawn moves the field by tens of points where everything else in this document moves it by one or two. Two strong players drawn last take 89.8% of wins against 34.3% opening together, on a random reference of 55.0%. The effect is generic — two average players last take 69.7% — and flow-sensitive. Not a lever anybody pulls, since the draw is random by design, but it sets the scale for every other balance argument.
A third tool was found measuring the wrong thing by relying on a default it did not set. voltron-study and backfire-study both assumed the engine's whole-pot backfire; the moment this release made it a dial defaulting to 0, their "backfire ON" rows silently became "off" — in backfire-study two rows printed identical numbers, which is what gave it away. Both now set the value explicitly. comeback-study hit the same trap in 0.90.0. A study must set every setting it claims to vary.
One delivered row did not reproduce: the backfire-off stag case came back 28.0%/39.4% against 24.4%/42.8% as delivered. Every other row matched to a tenth, including the whole-pot anchor, so it is recorded rather than resolved; it does not move the decision.
A new Part IV section: every leveling idea that was measured and turned down, with the number that killed it. Merged from the analysis chat's rejected-ideas page, which David asked for. The point is not history — most of these sound good, several were proposed more than once, and each re-litigation costs a study.
Six entries: triggers that measure the wrong thing (kickout on 2, tenure gates, the warm-up start boost), lowering the shark instead of raising the field, spending the race budget twice, repairs worse than their bug, the whole-pot backfire against coordination, and one near-miss — dials that are actually thresholds.
Every figure in it was re-measured before merging, and most had moved. The page was written against 0.89.0, before comeback-study.mjs was found to be a release behind and before arrivalGrace shipped. Corrected: the reference bar 14.3% -> 15.2%, the ceiling-floor repair 10.8%->7.5% becomes 14.3%->10.0%, the 0.5 boost 11.0%->2.1% becomes 15.2%->4.9%, the winner cooldown's casual row 1.7% -> 2.5%, ungated 2.9% -> 3.0%, tenure 1.9% -> 1.3%, photo finish 0.1% -> 0.5%, trailing pick 0.1% -> 0.7%, kickout on 2 0.3% -> 0.5%->0.3%. The story is unchanged everywhere; the numbers now match what this repo's own tools produce.
Two entries needed more than a number. The arrival exemption is no longer "awaiting a decision" — it shipped in 0.90.0, and at the strict reading, so the section says +0.9 rather than the +1.9 the looser variant measured. And the backfire paragraph cross-referenced a section that does not exist here, since that rule is decided but not built; it now says so.
tools/kickout-study.mjs and tools/level-study2-start-boost.mjs adopted from the analysis folder, because two of the claims could not be reproduced without them. They both do: start boost for everyone 4.6%, oracle start stacked on the comeback 12.7%, below the comeback alone.
Figure 22 is the bone pile chart. Its bars are the original run and carry a note saying so, since re-measuring moved several rows and the reference line; the ordering it shows is unaffected. The handbook gained an h4 level for the six entries.
The "Does being good on television carry over?" section is removed from the handbook, three figures with it. Several of the regulars are identifiable broadcast contestants, and a per-player read on their televised record sitting next to how they do here is not this project's to publish — the handbook already anonymizes to P-labels for exactly that reason.
Permanent rule, not a one-off edit, and it overrides the standing rule that the analysis chat's graphs always go into the handbook. The chart page still exists in the analysis folder; test/guides.mjs now asserts the handbook contains none of its phrases, so a future session merging it back fails the suite instead of shipping it.
The line is individual versus aggregate. Out: anything correlating a player's televised record with their Rumble results. In: the aggregate population data the robot model is calibrated on — 3,339 J!ometry player-games, 1,772 contestants — which is provenance for a model with no individual in it.
Figures renumbered 1–21 in document order. The configuration map stays unnumbered as the shark section's index.
Nine figures merged from ~/Developer/j-royal-rumble-data/charts/, per David's standing rule in that folder's README: the analysis chat's graphs always go into the online handbook. A chart left in that folder is half-delivered.
- Broadcast transfer (3 figures), as a new Part III section: does a strong televised record predict a strong Rumble? Real but loose, holds only where the gap between two players is wide, and accuracy is the half that does not transfer. It doubles as a caution about the robot model, which is calibrated on broadcast attempt and accuracy rates and inherits exactly that. - The configuration map (2 panels) at the top of the shark problem, as the index to that section — every configuration measured across the project on one chart, with almost nothing moving the casual axis. Left deliberately unnumbered: it is the contents page for the section, not another result in it. - The trigger study (4 figures) into "Where the trigger belongs", which until now carried the prose and tables without the pictures.
Figures are renumbered 1–24 in document order, which fixes a duplicate. There were two Figure 11s — the entry-interval chart and the length estimator — and nothing referenced either by number, so the collision had gone unnoticed. The one genuine prose cross-reference was remapped by caption rather than by number.
Also fixed on the way in: the trigger figures carried a bare < in "wins < 3", which is invalid HTML in text content. Escaped.
Corrected: a claim from 0.89.0 that there was nothing to compare until somebody else hosted a match. There is. David has identified which matches somebody else ran, and Figure 17 already quantified it — the three players who played both Aug 16 matches all shifted slower together under the other host at an identical delay setting, +45ms, +90ms and +170ms. That is a calibration offset, not a skill difference, so absolute times compare within a match or across matches with the same host, and not otherwise. The set is small, which is why the host field exists going forward rather than a reason to discount it.
Arrivals were clamped to the current ceiling. That clamp was written for flat stakes; the stake then learned to ride the overtime multiplier and the ceiling did not, and the two move in opposite directions through overtime, so they crossed — a fresh entrant stopped getting the full multiple from about x4, a comeback from about x8. The P26 fix, undone in the phase it was written for.
The repair does not go through the roof. Scaling the ceiling floor by the multiplier was measured and is worse than the defect: lifting the roof hands the elite the clipping they currently suffer, taking casuals 14.3% -> 10.0%. The ceiling clamp is an accidental leveller, because whoever has accumulated most deep in overtime is almost always the elite.
So the carve-out is arrival-only and lives at the cap in resolveClue. An arrival lands capped by the roof as it stood when overtime opened, and is not clipped down to the falling roof until its score first touches it. On by default, arrivalGrace: false restores the old clamp.
Casuals 14.3% -> 15.2%, top shark 56.9% -> 55.6%, drain untouched. It costs about a point of back-half at twenty players and nothing measurable at thirty — the second time the skill and draw axes have pulled against each other here.
The flag covers the arrival, not what the arrival then wins. The variant this was measured from flagged every arrival, which also exempts a player who landed under the roof and climbed above it by winning pots — an open-ended ceiling exemption on accumulated score, which is the property the floor repair was rejected for. It scores 16.2% because it is a bigger rule: 7,777 skipped clips against 1,882 off an identical count of above-roof arrivals, and 3-4 points of back-half instead of 1. The point left on the table is deliberate.
Two shapes measured as no-ops and are recorded so they are not retried: a fixed-length grace of 1 or 4 clues, and letting the cap rise to the highest arrival. An above-roof arrival pays itself back under the roof within about a clue, so a brief grace protects something that liquidates itself.
Also: tools/comeback-study.mjs drove its own comeback and had fallen a release behind, returning a flat half stake while labelling the row SHIPPED. Every figure it produced described the pre-0.88 rule, and some had reached the handbook. Fixed, and it now routes arrivals through the engine hook so it cannot silently diverge again. All three study tools agree.
Deploys no longer trip over their own lockfile. npm install on the box rewrites package-lock.json in place, so the working tree there is dirty by the time the next deploy pulls. That was invisible for months because the lockfile never changed upstream; the moment it did, git pull --ff-only refused and the deploy failed with nothing wrong with the code. deploy-remote.sh now discards the box's copy before pulling — it was never the source of truth for it.
Climb down before the clue is read and there is no cooldown. The five-clue wait prices *riding* a clue at double, not declaring one, so a player who changes their mind in the same gap between clues now pays nothing. setTopRope clears the stamp when cluesRevealed has not moved since it was set — which is exactly "no clue came and went while they were up there". It cannot be used to peek: declarations are only accepted between clues, so there is never a clue on the board to look at. The buzzer needed no change; it already re-enables on topRopeWait reaching zero.
The Discord copy never mentioned the cooldown at all. It does now.
The setup page asks who is hosting, under the room code, in both quick and expert mode. It is saved with the match record as host and shown as a column at /history, and remembered per browser because the same person runs nearly every match. Metadata, not a rule: it travels beside the settings like blend and never reaches the engine.
The point is that the read is a large part of how a match plays and nothing measured so far can see it — every recorded match has had the same host, so the pace figures are one person's delivery as much as the game's. Grouping results by host is not built yet; there is nothing to compare until somebody else hosts.
Recorded, not fixed: the falling ceiling eats the scaled stake. David spotted the comeback landing on the ceiling rather than the configured stake. It does, and the same clamp has been capping *entrants* since scaleEntryStake shipped — the stake rides the overtime multiplier, the ceiling floor does not, and the two move in opposite directions, so they cross. At the defaults a fresh entrant stops getting the full multiple from about ×4 and a comeback from about ×8.
Not repaired here because every candidate moves the ceiling, which is the dominant fairness lever: scaling the floor by the multiplier puts it above the starting ceiling at ×8 and kills the overtime drain, and letting arrivals land above the roof is cosmetic — the next resolveClue clips them straight back. The numbers and the failed candidates are in the handbook; the player-facing rules now say the ceiling applies to what you walk in with rather than promising a multiple the game does not always pay.
From a study of the trigger against all 89 eliminations in the eleven recorded matches.
41 of those 89 happened during overtime, and a flat return stake there bought one to six clues of life — half a stake against quadrupled values is one pot payment from going straight back out. This is the P26 pattern that scaleEntryStake already fixed for entrants and revivals; the comeback was the one re-entry path left flat. It now rides the same switch, and is clamped to the ceiling like admit() — the cap is applied earlier in resolveClue than the comeback runs, so an unclamped stake would have sat above the roof for a clue.
Measured at the shipped gate and boost, same fire rate:
| | casuals | mid | top shark | |---|---|---|---| | fire flat (was) | 11.0% | 10.0% | 61.9% | | don't fire in overtime | 6.3% | — | — | | stake × multiplier | 14.1% | 13.7% | 56.1% |
Strictly better, so there was no trade to weigh. Not firing in overtime guts the rule, since half the eliminations live there.
The handbook's VWQW attribution is corrected. Part IV credited P11's arc — out at clue 36, back, $10,700 peak — to the comeback. The logs give him eight race wins at elimination, so he would not have passed the wins<3 gate; both rules were on and that was revival's queue re-entry. Marked overturned rather than rewritten. The comeback's real first-night clients were the nought- and one-win players, which is exactly who it was built for.
Three findings recorded so they are not re-argued: tenure is the wrong gate axis (the invisible dead mostly last 9–17 clues, so it separates nobody), three sits at an empirical valley in a bimodal distribution, and ungated the rule is a subsidy for the strong. Full working in the handbook.
The Discord rules now need three messages rather than two — the previous posting note predicted the next rule would force it, and this was that rule.
The robots were reading other answers off the board instead of inventing one, which on a loose category is nonsense. The Claude path existed; it was not running.
The bug was that nobody could tell. Every failure went through a bare catch {} to the local fallback with nothing logged, so a missing key, a rejected key, a wrong model name and a timeout were indistinguishable from outside — all four look like robots talking nonsense. It ran that way in live matches.
GET /api/health now reports wrongAnswers: the model, whether a key is configured, how many answers were asked for, written and fallen back, and the reason for the last fallback. mode is what the room is actually hearing — claude, local or mixed — not what was configured. The server logs the reason once per distinct cause rather than once per clue.
Haiku, and the official SDK. claude-haiku-4-5 because this is a throwaway sentence per clue rather than a reasoning task. The hand-rolled fetch is replaced by @anthropic-ai/sdk, which brings typed errors — so the health readout can say *which* failure it was — and a retry on the 429s and 503s that used to fall straight through to the board. The latency ceiling is unchanged at 4s: one retry at 2s each, sized to finish inside the host's read.
test/wrongs.mjs pins the fallback path, which is what CI runs with no key.
Quick setup offers Tournament, Arcade, Chaos. "Standard" is now Arcade, which is what it always was.
The rename caught a real contradiction. All three rule sets left the comeback on, so all three produced a match whose screens read ARCADE MODE — including the one called Tournament. Tournament now drops the comeback along with targeting, which is what makes it a different mode rather than a different preference: no player is ranked with time off their press, so buzz order is buzz speed and the times are published.
Chaos is Arcade plus the advanced mechanics, and says so.
Part IV replaced with the version built from all eleven recorded matches: sixteen figures, the estimator diagnosed rather than just humbled, the host-relative clock finding, and the balance-lever simulation with the shipped gated comeback's first night in it.
Merged, not swapped. The incoming file was built from an older Parts I–III and would have dropped the gemstone stables section, the 10,500 small-field ceiling, One foot on the floor, Arcade/Tournament and the comeback correction — seven of the needles guides.mjs asserts, which is that guard doing exactly what it exists for. Parts I–III are the current ones; Part IV is the new one.
Two measurements of the comeback now sit in one document, so it says which is which. Figure 16 gives a single casual's match-win probability (1.9% ungated, 7.3% gated); comeback-study reports the three casuals' combined share (2.9% against 11.0%). One is a player's own chance, the other the group's share of the table. The threshold finding is restored alongside them.
docs/analysis/README.md carries the analysis without the identities. The delivered note named eight identifiable broadcast contestants alongside a per-player critique of how well each did. This repository is public and git history is permanent, so the reasoning is kept and the real-name mapping is not; the P-labels map to in-game handles, which were already committed in the CSVs.
A match is now one shape or the other, named on the host console, the watch screen and every player's buzzer for the whole match.
Arcade is the comeback switched on. Because a player on the way back is ranked at a fraction of the time they pressed, buzz order stops being buzz speed — so the room sees 1st, 2nd, 3rd rather than times. A player still sees their own reaction time on their own buzzer, to the tenth of a millisecond. Nobody sees anybody else's.
Tournament is the comeback off. buzzEdge is always 1, the fastest press wins every race, and the times stay public because they mean exactly what they look like.
The times are absent from the payload, not merely unrendered — watchView and raceView omit ms entirely in Arcade, the same discipline that keeps the answer off the watch screen. test/arcade.mjs asserts both directions and that a player's own number still reaches them.
This replaces the 0.85.1 fix, which sent the ranked time beside the real one so the host could see why a slower press held the clock. Showing the order beats explaining the arithmetic behind it, and the edge field is gone.
A bug fixed with it: rerank ranked on raw ms while rankRace ordered on ms * buzzEdge, so the place a player was told could disagree with who actually held the clock. It only ever showed up once the comeback existed, because every edge was 1 before that — and it would have made the new place display wrong from its first match.
The Discord rules are now at capacity: 1,970 / 1,707 against a 2,000 cap, and no other two-message split fits. The next rule added there needs a third message.
The console locking up on an entry, third and final piece. 0.85.2 stopped the banner eating the host's keystrokes; it went on eating their mouse clicks.
The banner is anchored at bottom:92px against a dock whose height is min-height:92px — a minimum, not a fixed size — and the buzz chips wrap. So as soon as several people buzz the chip row runs to two lines, the dock grows past 92px, and the banner lands on top of Correct / Wrong / Nobody got it. At z-index:60 with its own click-to-dismiss, it swallowed the click and closed itself: the host clicked Correct, the banner vanished, nothing happened.
That is why entries with robots were the worst of it. Every robot buzzes, so the chip row wraps every time and the overlap is guaranteed — which is exactly the "harder when it's a bot" report, and not a robot bug at all.
pointer-events:none makes the overlap harmless at any dock height: clicks reach the control underneath. It still closes on any key and on its own timer. The overtime splash keeps click-to-dismiss deliberately — it covers the screen and is obviously in the way, where this one is a thin strip that looks harmless.
pagerefs now fails if the banner takes input by any route.
Reported from the full buzzer: a player could not read his own board because the title cards would not show a second row.
The board was six separate grids, one per column, so row one was sized per column and a long category made its own card taller than the rest. The fix for that had been a hard max-height with the title clamped to three lines — which is what cut a title off. Beside a 420px buzzer on a narrow window each column is about 93px, where the text wraps every 13 characters or so, so anything past roughly 28 characters lost its tail with nothing on screen to say it had one.
It is one grid now: display:contents on the column lifts the cards and cells into it, so row one is shared across all six and sized to the tallest card. The cap is gone, the clamp goes from three lines to five, and the cells stay aligned because they are literally in the same grid rows.
This is the hang. The entry banner stays up for 4.2 seconds after somebody walks in, and the console's keyboard handler opened with if (document.querySelector('.entry')) return;. So for four seconds after every entry the host's keys did nothing: pick the next clue with the mouse, press space to arm, and the game sits there. Press again and it works — which is both the "it hung when it was the clue when someone was coming in" report and the reason a host ends up sending two adjudications for one clue, which is what threw the destructure error in 0.85.1.
The banner still dismisses on any key; it just no longer swallows it. Eating the host's input to close a decoration was the wrong trade — a banner in the way is a nuisance, a dead keyboard is a stopped game.
And the comeback edge outlived its 40 races. racesRun was incremented on entry.missed, where the field is entry.missedIds, so it read as "clues somebody won" and a clue everybody buzzed and nobody converted did not count. The edge is measured in races, so it lasted longer than it should have — with nothing on the console saying who held it, the effect from the outside was a player who seemed to keep an unexplained advantage on the buzzer. Reported as exactly that.
The expiry test only played clues with a winner, which is why it passed for several releases. It now plays contested clues nobody converted as well, and checks that a clue nobody buzzed still does *not* burn the edge.
Adjudicating a clue twice took the game down. The host pressed Y twice and got "cannot destructure property 'slot' of 'match.clue' as it is null", with the console apparently stuck. Y fires on keydown and the guard in front of it reads the console's own copy of the state, which still shows the old clue until the server's push lands — so a second press arrives after the clue is settled and match.clue is null. resolve was the only handler of its kind with no null check, so it destructured null and threw, and hostOnly handed the TypeError to the host verbatim.
It is refused in words now ("that clue is already settled — pick the next one"), the console will not send the second one, and test/doubleresolve.mjs plays the sequence and checks the match is still playable afterwards.
And the console now shows why a slower time is winning the race. A player on the way back from a near-elimination is ranked at a fraction of the time they pressed, so first place can legitimately show a bigger number — reported from the same match as "P11 actually was fastest but P12 was highlighted as the first person in". The chip shows both numbers, which the comment above rankRace had claimed for a while without it being true.
The setup page opens on two dropdowns instead of forty controls: a rule set (Tournament, Standard, Chaos) and a game speed (Blitz, Standard, Extended), the roster, and a button that adds robots. Expert is the old full card, one click away and remembered per browser.
Tournament is every standard rule except targeting — nobody can be ganged up on, so it comes down to the buzzer. Chaos is everything, advanced mechanics included. Blitz feeds a player in every 5 clues; Standard and Extended pace to 20 and 40 minutes and let the entries scale to whoever turns up.
The presets write into the real controls rather than living beside them, so switching to Expert shows exactly what a preset did, and the dropdown reads Custom the moment a hand-edit stops matching it. Expert renders in both modes and is hidden with CSS: collect() reads every control by id when it saves, and a card that is genuinely absent takes the Save button down with it.
Somebody knocked out before they ever got going comes straight back on half a stake, with 70% off their buzz for the next 40 races. Once each, and only for players with fewer than three clues to their name.
The gate is the design. Ungated, the same mechanic is a subsidy for the strong — they take their free life too and end up further ahead. In the study's six-player field (one 95ms elite, one at 130, a mid, and three casuals at 210/240/270), gated at 40 races:
| | elite | second | mid | the three casuals | |---|---|---|---|---| | ungated, 70% | 48.7% | 35.9% | 12.5% | 2.9% | | gated, 70% | 61.9% | 17.1% | 10.0% | 11.0% |
Ungated pulls the elite down further and does nothing for the people it was built for. Gated, the casuals take ten times the share.
The boost is a threshold, not a dial — found the hard way. It shipped at 50% and was corrected to 70% the same evening, because 50% measures as very nearly nothing: casuals 2.1% against 11.0%, with the elite taking back 12.5 points. The reason is that the discount only counts if it puts a slow player under a fast one, and against a 95ms elite the three casuals need 54.8%, 60.4% and 64.8% respectively. At 50% none of them get there. Below ~55% the mechanic is decorative; the useful range starts around 65%.
The figures first published with this entry were not reproducible — "bottom four 1.5% → 41.9%, elite 93.8% → 38.6%" matches no row the study produces, and has been replaced above with its actual output. Retune against npm run comeback-study and the threshold table in engine.js, not by feel.
It costs draw fairness, which is worth knowing. The back half of the draw goes from 50% to 59% at sixteen players, because a late entrant is likelier to still be under the gate when they hit trouble. Fairness by skill and fairness by draw are pulling against each other here for the first time.
It also quietly preempts bounties and revival for anyone under the gate: they are never eliminated, so nothing pays out and no second life is spent.
Part IV replaced with an analysis of all nine recorded human matches rather than the two it was built on. Four new figures: the estimator against live play, pace by match, race-win rate against reaction time, and the competitive picture.
The estimator comes out humbled — seven of nine within a handful of clues, with two structural misses it cannot currently model: a four-player match that ran 2.3x its prediction because tiny fields trade points without eliminating, and one that ran 180 clues against 75 predicted because of latecomers, revival and a long overtime. It models none of those three.
It also puts a number on something David had only suspected: the winner was the fastest buzzer in the field in half the decided matches, and race-win rate is monotone against reaction time across every player with enough buzzes to judge. Fairness by draw is measured and fine; fairness by skill has never been measured at all. That is now written down as an open question with a definition of done.
Measured, and written up as Figure 12 in the handbook. The entry interval turns out to be nearly free at small fields — every value from 3 to 20 lands within a few points of an even draw — and dangerous at scale: twelve players on a 20-clue gap hand the back half 66% of the wins. What it really controls is length. The cap drops from 15 to 10, which leaves fairness alone and takes a third off a small match.
Eight measured just as well and was tried first, but it made a six-player fifteen-minute game unreachable on auto, so the estimator warned every time. That also turned up a long-standing flaw: the "set the target to N" button suggested a number that clamped again and re-raised the same warning.
Two sounds on one entrance. The horn played under the player's own music. If somebody brought a theme, that is the entrance; the horn only fills in for people who did not choose one.
Two players called Dave were one player. A duplicate name was accepted and the room had nothing else to tell them apart. A second Dave is now turned away with an explanation. Worse, the join screen never came back after a refusal — return sfx.preload(); sat in front of the renderJoin() meant to redraw it, so anybody rejected was stuck looking at nothing.
Back to home and Start a new game on the finished-match screen. The console was a dead end and the host had to know the URLs.
Entries come faster in a small field. The cap was 15 clues and four-to-six player games always hit it, so every small match had identical pacing. Ten now.
Reports can be marked resolved or deleted from the control room. Settled ones sort below the rest rather than scrolling the live ones away.
A live bug report said "entrance music didn't play", filed from the host console — which never played it. The mechanism was working; it was pointed at the wrong screen. Music played only on a watch screen with sound enabled, and a watch screen is optional; nobody had one open that night.
It now comes out of every player's buzzer and the host console, which always exist. The player code is one shared module rather than a copy per page, and the music stops when the buzzers arm.
A new entrant's stake now rides the overtime multiplier, and so does a revival. A fixed 3,000 walked into a ring where single clues paid 2,000: in a real match P26 entered at clue 150 into x2, lasted six clues without winning a race, was revived at 1,500 into x4 where the top row paid 2,000, and was gone after one. Measured over 3,000 matches, entrants at x2 or above who died within three clues fell from 58% to 9%.
Also: board control lights the whole score row rather than a 7px dot; the watch screen now draws the category hint at all, which is why a live room could not play a category that depended on it; and the report button moved off Arm buzzers on the host console.
A real 53-clue six-player match had the winner pinned at the 6,000 ceiling for 20 clues, with 11,930 points swallowed: half the match, they answered correctly and gained nothing. Small fields now get 10,500. The ladder is deliberately no longer monotonic — a small field needs the most headroom, not the least. Ceiling decay was tried again and rejected again: a falling cap is one the leader meets sooner, so it made pinning worse and brought back the late-draw bias.
Jumping the lights while waiting in the queue counted as a live attempt. One player finished a real match credited with 28 attempts across a tenure of one clue. The buzz path checked for this; the early-buzz path did not.
YouTube clips play through a covered player rather than a hidden one, because browsers refuse autoplay to an iframe with no size. Everything stops after five seconds, the test button included.
An illustrated guide that assumes nothing, and the Discord explainers rendered as a page. The guide warns that the space bar only works when the buzzer window is in front, which caught people out in testing.
Stables are named from a list — Diamond, Ruby, Emerald, Sapphire, Onyx, Topaz — each with a colour and a line-art badge that tints its members' rows on every scoreboard. The pot is now split evenly across the stable.
When a stable member wins, the pot stays full size and the teammates' share is loaded onto whoever is outside. The first version let the pot shrink, which protected the stable and did nothing to anybody else.
Teams, with betrayal costing half your stack. Off by default.
A straight shuffle regularly dealt three or four robots in a row, and a stretch of match where nobody real walks in is the part a room notices.
Every match on the server, with a way to end one, and the saved logs. Matches with no activity for ten minutes end themselves: a forgotten test match blocked a deploy for an hour.
Twelve original 8-bit themes in three moods, synthesised rather than sourced so the licensing is unambiguous.
/ used to serve the host setup page, so anybody who typed the domain landed on the controls for running a match.
The ceiling turned out to be a stronger fairness lever than the entry interval, and it had been confounding every earlier measurement by jumping at 25 players.
A public, read-only board at /watch/CODE, with the answers never sent to it.
Robots that buzz on a real clock, with reaction times and accuracy drawn from 3,339 real player-games rather than a fitted curve.
Stakes double every six clues with nobody eliminated, up to eight times face value. Took three attempts; the failures are recorded in the handbook.
Buzzers, a board, scoring where every other player pays the winner, and elimination below zero.