Mithrandir Metrics
3 identitiesaudit-backed
Section VI / Methods / Assumptions, validation, change history

Every number has a paper trail.

Risk archetypes, validation examples, formulas, and version notes. Methodology cards summarize what is live, what validated, and where coverage is intentionally limited.

Lanes
3
Markets
6
HR ECE
.0149
Method audit 2026-04 Calibration harness available Lane health 18 rows Scheduler 1 failed / 9 stale
Published methodology

Definitions, formulas, validation gates, calibration results, caveats, and null findings are public-facing by design.

Protected implementation

Raw data licenses, credentials, scheduler internals, model artifacts, and exact feature-pipeline code paths stay out of public copy.

Reader promise

If a surface is experimental, missing coverage, descriptive-only, or parked after validation failure, the page should say that directly.

Three Lanes Beta / current market coverage
Post-beta coverage

What each market can honestly show today.

Validation receipts linked

Validated non-HR challengers remain unavailable until their real production inference outputs are wired and logged. Full production status still requires forward evaluation and direct posted-line calibration.

Market Palantir Fangorn Valinor Triple Confirmed Consensus Receipt path
Home Runs Production Production Production Full Full HR production validation receipts
Total Bases Production CatBoost validated, not wired LightGBM + beta validated, not wired Waiting on live output Insufficient lanes outputs/validation/three_lanes_beta/total_bases/
Hits Production Random Forest validated, not wired In active development Waiting on live output Insufficient lanes outputs/validation/three_lanes_beta/hits/
Pitcher Outs Production Random Forest validated, not wired In active development Waiting on live output Insufficient lanes outputs/validation/three_lanes_beta/pitcher_outs/
Strikeouts Production In active development In active development Waiting on multi-lane output Insufficient lanes outputs/validation/three_lanes_beta/strikeouts/
Futures Production Intentionally limited Intentionally limited Waiting on multi-lane output Insufficient lanes docs/codex_context/three_lanes_beta_preregistration_amendment_001_2026_05_24.md
Mithrandir+ / flagship derived metrics
Stuff+ v1.1 / promoted

Run-value target, full backfill, original feature set.

Y-Y r² .504
Training

2020-2023 Statcast pitch-event backfill via pybaseball. Validation 2024; held-out test 2025.

Target

Statcast delta_run_exp. Lower run value is better for pitchers; 100 is league average on the published plus scale.

Feature set

Velocity, release point, extension, movement, location, zone height, count, pitch type, pitcher hand, and fastball-relative deltas.

STABILIZATION 100 pitches RMSE VS FIP -0.168 ABLATION features tested
Ablation finding

Spin axis, batter handedness, and platoon context were tested. They added noise for v1.1 and were dropped from production despite a proxy-target variant producing shinier raw Y-Y numbers.

Arsenal Decomposition

Phase 2 initially attempted a Kirby Index. Audit showed public Kirby is a command metric built from release-angle consistency, while Mithrandir had built pitch-type Stuff+ composition. We renamed it Arsenal Decomposition and surface it only as descriptive support under Stuff+, not as a predictive metric.

Two-Strike Adjustment null

Two-Strike Adjustment Index v1.0 followed the Chamberlain chase-delta framework but missed the locked gate: contact-rate r² .147 vs .300 target, and raw two-strike K% remained the better next-year K% predictor (.445 r² vs .080). This stays parked as an honest null, with bat-tracking integration deferred for any v1.1 attempt.

Pitch-Type Vulnerability partial promote

Per-type validation cleared FB, CH, and SI but exposed sample gaps for CB/FC and a slider miss. Family aggregation by hitter-recognition class validated Fastball, Sinker, and Changeup; Breaking horizontal remains caution-only at r .239; Breaking vertical and Cutters/Splitters stay insufficient-sample with explicit tags.

Conditional OAA internal reframe

Conditional OAA was tested as a standalone predictive metric and missed the locked Y-Y gate: r² .174 vs .450 target. The small-sample test was excellent, though: RMSE 2.073 vs raw OAA 6.377. We therefore apply shrinkage internally below 200 defensive outs and label those rows "regression applied" instead of promoting a separate metric page.

Stuff+ v1.1 = 100 + z(predicted pitch run prevention) * 10 production target = delta_run_exp / production features = original pitch-shape + location set Arsenal Decomposition = pitch-type-relative Stuff+ by pitcher and pitch family Pitch-Type Vulnerability = regressed same-family wOBA-against, inverted to resistance percentile

Open Stuff+ methodology + leaderboard

Open Pitch-Type Vulnerability methodology + leaderboard

SEAGER+ v1.0 / promoted

Plate discipline index, reframed after validation.

BB% r .648
Original gate

SEAGER+ was initially pre-registered against next-year ISO at r ≥ .300. The first formula produced r .023 and failed.

Audit correction

A methodology audit found a denominator error: public SEAGER is rate-based ST - HPT, not a per-pitch run-value average. Correcting it moved ISO r to .274.

Reframed signal

Alternative-target diagnostics showed the corrected metric is a plate discipline monster: next-year BB% r .648 and chase-rate r -.692.

PRIMARY TARGET BB% / chase ISO OBSERVATION r .274 ROADMAP called-strike weighting
SEAGER+ = 100 + z(Selection Tendency - Hittable Pitches Taken) * 10 Selection Tendency = good takes / non-hittable opportunities; HPT = hittable takes / total takes
Transparency note

The ISO threshold miss is documented, not hidden. v1.1 roadmap: called-strike probability weighting plus bat-tracking quality on correct swings if we want to chase the power-prediction framing again.

Open SEAGER+ methodology + leaderboard

Lane health / calibration drift moat
Epistemic chrome

Models are allowed to be wrong; they are not allowed to hide it.

2026-07-27
Scheduler freshness

Last cycle 2026-07-28T03:50:09.543655+00:00 / 1 failed / 9 stale

Failed: train_stuff_plus

Market Lane Logged Brier ECE 14d trend Calibration Flag
Home Runs Palantir 11499 0.122 0.011 flat recalibrated healthy
Home Runs Fangorn 11499 0.122 0.010 flat recalibrated healthy
Home Runs Valinor 5904 0.116 0.006 flat recalibrated healthy
Strikeouts Palantir 964 0.247 0.031 degrading recalibrated drift
Strikeouts Fangorn 53 0.257 0.119 insufficient sample skipped insufficient sample insufficient sample
Strikeouts Valinor 460 0.247 0.038 degrading recalibrated drift
Total Bases Palantir 5326 0.239 0.073 degrading recalibrated drift
Total Bases Fangorn 0 -- -- tracking skipped insufficient sample tracking in progress
Total Bases Valinor 0 -- -- tracking training time tracking in progress
Hits Palantir 344 0.177 0.049 degrading skipped insufficient validation split drift
Hits Fangorn 0 -- -- tracking training time tracking in progress
Hits Valinor 0 -- -- tracking training time tracking in progress
Pitcher Outs Palantir 512 0.259 0.056 degrading recalibrated drift
Pitcher Outs Fangorn 0 -- -- tracking training time tracking in progress
Pitcher Outs Valinor 0 -- -- tracking training time tracking in progress
Futures Palantir 0 -- -- tracking skipped insufficient sample tracking in progress
Futures Fangorn 0 -- -- tracking training time tracking in progress
Futures Valinor 0 -- -- tracking training time tracking in progress

Home Runs / Palantir

Healthy

Tracked
0-10% bucket: predicted 6.5%, observed 5.2%, n=1089 10-20% bucket: predicted 15.6%, observed 14.8%, n=8856 20-30% bucket: predicted 21.6%, observed 18.7%, n=1554 predicted observed
11499props 0.122Brier 0.011ECE

Home Runs / Fangorn

Healthy

Tracked
0-10% bucket: predicted 5.7%, observed 6.4%, n=564 10-20% bucket: predicted 15.4%, observed 14.5%, n=10419 20-30% bucket: predicted 24.4%, observed 21.4%, n=504 30-40% bucket: predicted 31.1%, observed 0.0%, n=12 predicted observed
11499props 0.122Brier 0.010ECE

Home Runs / Valinor

Healthy

Tracked
0-10% bucket: predicted 8.9%, observed 9.5%, n=315 10-20% bucket: predicted 14.1%, observed 13.6%, n=5547 20-30% bucket: predicted 21.0%, observed 28.6%, n=42 predicted observed
5904props 0.116Brier 0.006ECE
Rolling projections / Bayesian shrinkage
Daily update layer

The prior stays visible; the season earns weight one game at a time.

live M1 + rolling
Preseason

The frozen projection from the original season simulator. It remains on-page as the baseline.

Current pace

The raw extrapolation of observed 2026 performance. Useful, but noisy.

Rolling

The Bayesian blend: preseason prior plus observed performance, weighted by sample size.

TEAM WINS 60-game regression constant HITTER RATES 220 PA PITCHER RATES 240 BF
observed_weight = opportunities / (opportunities + regression_constant) rolling_projection = prior * (1 - observed_weight) + current_pace * observed_weight team game lines = rolling team talent + opponent context + starter rolling ERA + bullpen FIP + lineup status
Limitation

Small samples mean high uncertainty; early-season rolling projections intentionally mean-revert toward preseason until the observed sample becomes persuasive.

Per-model write-ups / status
Validation

HR Valinor is the worked example.

Non-stacked LightGBM + beta calibration promoted on ECE strength; Brier difference was inside bootstrap noise.

Glossary / formulas
ECE = sum_b (n_b / N) * abs(avg_pred_b - avg_outcome_b) Brier = mean((p_i - y_i)^2) RA_FV = FutureValue * P(reaching ceiling)
Versioned changelog