The Engine
The half of this paper that makes falsifiable claims. It ranks hitters on how they actually strike a baseball, files the resulting calls with a metric, a threshold, a window and a timestamp — and then grades itself in public, misses included. Everything below is either a claim, the scoreboard for the claims, or a door to check them.
15 calls filed · 7 studies pre-registered · 2 published as nulls · holdout n=721
Next resolution: August 12, 2026 · 10 days
Also on the engine: The RadarThe BaselineThe One ChartMethodologyThe receiptsThe LegendOpen DataThe Ethos
What we tested, and what it came back
Every study here was pre-registered: the protocol and the scoring script frozen in git before the seasons that judge it were examined, run once, published either way. Two of the seven came back no. Four came back yes and still did not enter the engine— a pass is not a ship, and that distinction is the interesting half.
CAME BACK YESin the engine
The luck gap and EV95, out of sample
The engine rests on this one. Both signals held on seasons the model had never seen, and the boards are built on the result.
CAME BACK NOnot shipped
It worked in both training years and died on its one real test. It does not enter the engine.
CAME BACK YESnot shipped
Our expectation model grades a batted ball on how it was struck and cannot see legs, so fast men beat it every year. The finding held out of sample to the third decimal — and it still did not enter the engine, for the reason directly below.
CAME BACK NOnot shipped
Having found the bias, we built the board that corrects for it. It picked worse fade candidates than the plain luck gap, in both test years, so we did not ship it — and the sprint speed on the Fade board is printed, never ranked on. The previous receipt carries a correction we appended the same day, when this run refuted a sentence in it.
CAME BACK YESnot shipped
A pitcher's ERA barely repeats from one year to the next; his rate of called strikes and whiffs (CSW%) does, and out of sample it carries real information about next year's ERA beyond the ERA he already has. But it did not out-predict ERA the way EV95 out-predicts results for hitters — it is a complement, not a replacement. So CSW% earns a declared number on the pitcher pages and nothing it can rank on.
CAME BACK YESnot shipped
The process signal on a cleaner outcome
The follow-up: we re-ran the whiff-rate test against FIP, the defense-independent version of ERA, in case ERA's fielding luck was hiding the signal. It was not. CSW% adds even more to next year's FIP than to next year's ERA — and still does not beat it. A complement on both, a replacement on neither.
CAME BACK YESnot shipped
The hitter luck gap, pointed at arms: a pitcher's FIP against his ERA. FIP carries real forward information about next year's ERA beyond the ERA itself (the gate passed) — but out of sample it is co-equal to ERA, not the superior gauge the training years showed. It authorizes forward-graded arm calls, which we are filing by hand first; it does not authorize a ranked board.
Every door
The claims, the scoreboard, and the ways to prove us wrong — in that order.
The Radar →
The nightly boards — who is about to get better, who is about to give it back, and who nobody is talking about yet. They rank MISPRICING, not talent.
The Baseline →
Every qualified MLB hitter by EV95, in one list. The unselected population the boards are a selection out of — so you can see what we passed over, not just what we picked.
The Ledger →
Every call ever filed, on the terms it was filed on, with the countdown to its verdict. Append-only: misses stay up, and nothing is edited after filing.
The One Chart →
The whole thesis in a single scatter: contact quality against results on contact, and the gap between them. If you read one thing here, this.
Methodology →
How the engine is built, what it was validated on, and where it stops. Written to be argued with rather than admired.
The receipts →
Every pre-registered study, protocol frozen in git before the data that judges it was examined — published either way. The ones that came back no are the reason to believe the ones that came back yes.
The Legend →
Every statistic this site prints, defined in plain language. Nothing is printed here that is not defined here — a test suite fails if it is.
Open Data →
The boards, the record and the movers as JSON. Free, CC BY 4.0, no key — because being checkable is the product, and a claim you cannot download is a claim you have to take our word for.
The Ethos →
The eight rules the engine is held to, including the one that says nothing is edited after filing. Each one shows up as behaviour in the code, not just as copy.
Fever Baseball · RE24 corr 0.975 with FanGraphs