Research preview · methodology

Methodology

The platform's rule: separate questions get separate numbers, and every number carries its unit, status, model version, and evidence profile. Nothing below is live yet — this page states the contracts the models must earn their way into.

Four families, four questions

FamilyConstructUnit
CORE (rate)Retrospective, context-adjusted rate contribution in a declared role. Published as three mechanically separate products: Actual (realized run value), Expected (process the event model deserved), and Forecast (point-in-time prediction). These never share an unlabeled leaderboard.Runs above average per role-standard workload (600 PA; 750 BF starters; 250 BF relievers), plus a role-cohort index centered at 100
CROWN (cumulative)Value above a role-specific, empirically defined replacement baseline, as a fully visible ledger: batting + baserunning + fielding + catching + position + league + replacement + declared residual. Team totals reconcile publicly; residuals are shown, never redistributed invisibly.Runs, converted to wins by a versioned league-season calibration
PULSE (form)Recent evidence against the full-season baseline, with exact window, opponent/park mix, and role changes disclosed. A rolling number never asserts a mechanical cause.Windowed rate deviation with uncertainty
SIGNAL (reliability)Six separate evidence axes — sample, identifiability, granularity, noise, availability, lineage — displayed, never silently multiplied into the estimate.Per-axis profile; the overall band is capped by the weakest load-bearing axis

Foundations being built first

Base-out run expectancy by season and era (24 states); event run values (RV = runs scored + RE(after) − RE(before)); RE24; win probability and leverage as separate context-added lenses; park factors fit only on seasons before the season being adjusted; era boundaries as versioned contracts (including the 2026 ABS break). Every model run is replayable from an immutable source capture and a hashed fitted state.

What we refuse to do

No blending of actual and expected value into one score. No "better than WAR" claim without a task-matched, held-out test against declared comparator vintages. No cross-era framing series spanning the 2025 model update and 2026 ABS change. No missing data silently becoming zero. No proprietary metric as a training label. No catcher game-calling value without an identifiable causal design.

Status ladder

concept → scaffold → research → provisional → validated · (blocked / retired)

Promotion requires predeclared mechanical, construct, predictive, and reconciliation gates. Correlation with an existing metric proves similarity, never superiority.