MLB Pitch Quality (Stuff+) Evaluation System
Featured
A physical-characteristics pitch-quality model — grading pure stuff independent of location, count, and batter handedness. 100 = MLB average.
An independent reimplementation of Stuff+ that grades a single pitch on its physical
characteristics alone — velocity, spin rate and axis, induced vertical and horizontal break,
extension, release height and side, arm angle, and vertical/horizontal approach angle — plus
ten “primary-pitch differentials” measuring how each pitch diverges from that
pitcher's most-used offering that season. Location, count, batter handedness, and
seam-shifted wake are deliberately excluded: the first three aren't stuff, and the last
can't be computed on TrackMan, which keeps a future college port a recalibration rather
than a redesign. Trained on roughly 1.76 million MLB swings (Statcast, 2021–2026), each of the
seven pitch types gets its own paired sub-model — a three-class classifier for swing outcome
(whiff / foul / in-play) and an xwOBAcon regressor for contact quality — composed into a single
run-value grade scaled so that 100 is major-league average. Pitchers, never pitches, are split
80/20, so no arm appears on both sides of the wall.
Nearly every structural choice replaced a simpler one that measured worse, and the dead ends
are documented as carefully as the wins. Seven independent models replaced a single shared one
after a per-type generalization gap appeared that tracked sample size almost exactly — the
49K-row splitter overfit by 0.117 log-loss while the 571K-row four-seamer overfit by 0.041, a
spread one shared complexity budget can't express. A contact-quality weight of 2.0 replaced
the unexamined default of 1.0 after a sweep against three independent criteria. Along the way a
two-stage classifier built to isolate the noisy “foul” class scored no better and was
scrapped; two of three candidate differentials degraded the model and were dropped; and a subtle
standardization bug — normalizing per pitch type instead of pooling — was caught only after it
had quietly made 912 of the 930 pitches ever graded above 135 cutters.
Because there is no ground truth for “stuff,” the model is validated on four separate
claims rather than one, every number computed from leakage-safe out-of-fold grades (5-fold
cross-validation grouped by pitcher). It agrees with FanGraphs' Stuff+ at ρ = 0.697 across
3,096 pitcher-seasons, holds a pitcher's grade across adjacent seasons at r = 0.813, and
forecasts next-season ERA at r = −0.317 (innings-weighted) against a naive baseline's
+0.205 — roughly 1.6× better than the stat predicting itself. The less flattering result is
stated too: FanGraphs' own Stuff+ still forecasts ERA slightly better (−0.358 vs −0.317),
though this model leads on whiff rate (+0.514 vs +0.355) and strikeout rate (+0.417 vs +0.390).
- Python
- LightGBM
- Optuna
- pybaseball
- Statcast
- Run-value modeling
- xwOBA
- Grouped cross-validation
GitHub →
Stuff+ leaderboards — 2025, vs FanGraphs →
Free-Agent Contract Valuation Model
Complete · Personal
A front-office-style surplus-value model for an online baseball simulation league — valuing players and pricing the market as two independent problems.
A model that recommends a single contract — years and average annual value (AAV) — for any free agent in
Frostfire, a long-running online OOTP simulation
league with 21 seasons of data. The design mirrors how real front offices think: estimate what a player
is worth and what the market will pay as two independent models, then recommend a
signing only when value clears price by enough margin to absorb projection risk. I built it through
directed, end-to-end AI-assisted development — supplying the domain knowledge, modeling decisions, and
validation judgment while iterating the implementation through conversation with an AI coding assistant.
Value is assembled bottom-up: box-score stats are park-neutralized, decomposed into components (power,
contact, baserunning, defense, and catcher framing), converted to runs above replacement using linear
weights fit from the league's own run environment, aged forward with delta-method curves fit per
position and component, then run through a 40,000-iteration Monte Carlo that separates true-talent
uncertainty from season-to-season luck before pricing on a convex dollars-per-run curve. Price is a
deliberately simple ridge regression on real signings, and a length optimizer picks the contract that
maximizes risk-adjusted surplus — or recommends not signing at all. I excluded an available
player-ratings access token by choice: using information other managers couldn't see would be
cheating, not analytics.
On 280 held-out signings the market model reached an R² of 0.568 but cleared my pre-committed accuracy
bar — 85% of deals within ±15% of actual AAV — just 18.6% of the time. Rather than lower the target, I
diagnosed the gap: three independent rounds of experiments (gradient boosting, robust regression, added
features, and more training data) all failed to beat the simple baseline on held-out data. With no
scouting or player-ratings data exposed by the league and only ~250–280 unbiased signings to learn from,
the ceiling was structural — a limit of the available data, not the architecture. Identifying that
distinction, rather than chasing a number, was the real result.
- Python
- Surplus-value modeling
- Monte Carlo simulation
- Ridge regression
- Sabermetrics
- StatsPlus API
- AI-assisted (Claude)
GitHub →
Computational Modeling of T₂ Relaxation via Single-Sided NMR
Manuscript in prep
Simulating molecular T₂ relaxation for binary mixtures from first principles.
A research project in the Meldrum Spin Lab predicting T₂ relaxation times of substances and
molecular mixtures from physics-based first principles. Simulates molecular behavior in the
presence of a single-sided NMR magnet using molecular dynamics, with target manuscript
submission in October 2026.
Built an automated simulation pipeline combining OpenMM, MDAnalysis, SLURM, MPI, and tmux on
W&M's HPC cluster, with a validation suite comparing simulated vs. experimental T₂
values across a panel of test mixtures.
- OpenMM
- MDAnalysis
- SLURM
- MPI
- Bash
- Python
- SciPy
Repository private — available on request
Automated Google Business Profile Pipeline
Shipped · Client work
Replacing daily manual social posting with a fully automated content pipeline for a Northern Virginia realtor.
Designed and deployed an end-to-end content automation system for
Andrew Capuano,
a realtor and certified appraiser serving the Gainesville and Bristow area. The pipeline replaces
roughly 20 minutes of daily manual work with a fully hands-off system that publishes professional,
localized Google Business posts every morning.
On a daily schedule, the workflow draws a randomized hook–topic pair from a curated content library,
selects an image at random from a pool of 20 pre-sized assets, generates the post copy via a ChatGPT
step keyed to the day's inputs, and publishes the formatted post directly to the Google Business
Profile API through Zapier. The output is consistent, on-brand, and indistinguishable from manual posting.
- Zapier
- ChatGPT API
- Google Business Profile
- Workflow Automation
- Content Systems
Workflow template — available on request