In this post
- The institutional failure mode
- The answer, in one minute
- A profit is a lead, not a verdict
- Updating made the strategy busier, not better
- The market already owned the probability
- What a backtest cannot prove
- The next test
- When this shows up on your desk
The institutional failure mode
A result can be profitable, reproducible in code, and still fail a capital or promotion gate.
That is the failure mode this research is about. The domain happens to be MLB betting models. The transferable problem is broader: a backtest looks good, the team wants to fund or ship it, and the evidence still cannot separate a repeatable edge from a favorable run of outcomes, incomplete timing, or a price that was only observed, not accepted.
This is not a betting tip sheet. It is not a claim that these models should be funded. It is a case study in what a positive backtest is allowed to authorize.
The answer, in one minute
We tested whether MLB models could earn more than the sportsbook price implied. In a corrected replay of 2022–2025 games, three strategies finished ahead. The two strongest moneyline candidates were:
- LightGBM control: +$1,645, or +7.18%, across 229 settled wagers.
- XGBoost: +$2,032, or +8.03%, across 253 settled wagers.
That is encouraging. It is not enough to put money behind either model.
When we measured how much the result could move if we replayed comparable sets of game dates, both ranges still included a loss. LightGBM ran from about -$1,271 to +$4,573; XGBoost ran from about -$1,251 to +$5,065. The observed profit may reflect a repeatable advantage, or a favorable run of games. This sample cannot tell us which with enough confidence.
So the decision is deliberately narrow: lock these two rules, test each once on new games, and require its full after-cost uncertainty range to stay above break-even. Until then, they are hypotheses, not products and not betting advice.

A profit is a lead, not a verdict
“Profitable in this sample” and “reliably profitable” are different claims.
These candidates were trained before each season and then left alone during it. We call that setup annual-frozen. The model, market, betting rule, and estimated costs were locked before the test. That is the right shape for a capital decision: freeze the rule, then measure it.
They earned another look because reported profits were positive. They did not earn capital because the likely profit range crossed zero. A range that crosses zero does not mean the models are bad. It means the data still cannot reliably separate an edge from ordinary luck.
| Model | Profit | Return | Settled wagers | Profitable seasons |
|---|---|---|---|---|
| LightGBM control | +$1,645 | +7.18% | 229 | 3 of 3 |
| XGBoost | +$2,032 | +8.03% | 253 | 2 of 3 |
There is another reason to stay cautious: the models shared 146 wagers. Two positive rows do not automatically mean two independent discoveries. Much of their success may come from the same underlying bets.
The useful conclusion is not “we found a winner.” It is: we found two ideas that deserve one fair, precommitted test.
Updating made the strategy busier, not better
We also tested a tempting idea: retrain or switch models during the season. More recent data sounds like it should help. In this research, it usually did not.
The changing approach beat the fixed approach in only 8 of 22 matched comparisons. Its typical change in return was slightly worse when models were trained before the season (-0.09 percentage points) and materially worse when they were updated during it (-2.29 points).
The bets show why. Updating replaced most of the choices. Dynamic versions shared almost no wagers with their matched fixed counterparts (median overlap about 1%). Fixed versions still changed substantially, but overlapped far more (about 33%).
The newly introduced wagers did not pay for that disruption. Across the model groups in this analysis, updated-only selections reduced realized profit. In total, those overlapping research sleeves lost about $287,000 on wagers that appeared only after within-season updating.
That is not a claim that one live account lost $287,000. These were overlapping research portfolios, not one combined bankroll. The practical lesson is simpler: more frequent changes made the strategy busier without making the evidence stronger.

The market already owned the probability
Before treating the two candidates as interesting, we asked a harder question: did the baseball data know something the sportsbook price did not already know?
A model can predict games reasonably well yet add nothing to a price built from a deep market. So we did not only ask “can a model predict baseball?” We asked whether it could improve the market forecast on later, unseen seasons.
The stronger design locked the sportsbook’s margin-free probability as the starting forecast. In log-odds form, that market starting point was fixed at a coefficient of one: the model could not quietly downweight or rewrite the price. It could only add a correction from pregame baseball data: team form, bullpen use, schedule, park, starter history, availability flags, and the rest of a 156-input catalog, not odds, lines, or other market signals.
Across the four expanding 2022–2025 test seasons, none of six model families passed. Five had reliably worse log loss than the market. Elastic Net was inconclusive, which is not evidence of an improvement.
Supporting checks pointed the same way:
- Standalone learned classifiers scored worse than the sportsbook’s margin-free probability (Brier 0.247067 for the market benchmark), and every executable classifier strategy lost in forward replay (-2.99% to -12.01%).
- Six regression models still all lost in every tested season and every combined market (returns about -4.99% to -6.47%), even though they were not simply copying one another’s bets.
- Giving every family all 156 available inputs did not produce a reliable broad correction. Five were detectably worse than the market probability; Elastic Net remained inconclusive.
The 156 inputs were the full governed catalog available to this experiment, not the full universe of baseball information. Advanced FanGraphs data, including WAR, and comparable player-level projections were not available here. Those sources may help. The limited claim is that this catalog did not add a reliable market residual. Any new source still needs a timestamp and proof it was available before the decision; otherwise a historical model can look smarter than it could have been in real time.
The market benchmark is not itself a betting strategy. Its job is honesty: a proposed strategy must add useful information beyond the price already on the board.

What a backtest cannot prove
The hard work was not only building models. It was keeping every number’s economic meaning intact from the data source to the conclusion.
These are separate things:
- a model forecast;
- a sportsbook price;
- a bet the model selected;
- a simulated wager in historical data;
- a settled game result; and
- a wager a sportsbook actually accepted.
They are connected, but they are not interchangeable.
The replay uses historical FanDuel observations to calculate wager economics. That supports a simulation. It does not prove FanDuel would have accepted a stake at the displayed price. A production-grade record needs the sequence: when the source saw the price, when the bookmaker updated it, what offer was observed, whether it was accepted, and how it settled.
Timing matters for inputs too. In the available archive, starter information could be shown to exist at or before the quoted decision time for only seven games. For 9,322 games, it first appeared later. That is a gap in the evidence, not proof that the models used future information. It does mean we cannot make a strong claim about decision-time starter data without collecting it as it happens.

Finally, many rows do not automatically mean many independent chances to be right. Rolling windows can reuse the same dates, markets, and decisions. We used 2,000 resamples of game dates to estimate uncertainty, but resampling does not create new games.
The next test
The next experiment should be boring enough to be credible:
- Freeze the LightGBM-control and XGBoost moneyline rules before the test.
- State the price definition, estimated costs, uncertainty method, and protection against trying too many variants in advance.
- Separate wagers selected by both models from wagers selected by only one.
- Test each model once on new games.
- Continue only if the entire after-cost uncertainty range remains above break-even.
- If both fail, stop recycling the same inputs. Collect new, time-stamped information instead: probable starters, lineups, scratches, availability, weather, roof status, and multi-book prices recorded at decision time.
No weekly selector rescue. No fresh six-model sweep. No remix of the same 156 inputs because the latest chart disappoints.
When this shows up on your desk
If a strategy is profitable in sample, but even after costs the likely outcomes still include a loss, optimization is probably not the next step.
If research, paper, and live paths no longer agree, and nobody can prove whether the gap is timing, costs, model version, or fill, the same is true.
If a report shows a price or result you cannot tie to when the price was seen, whether the bet could be filled, and which model owned the result, the same is true.
First make the result trustworthy. Then decide whether it earns money or a launch. Not another model sweep.
Reproducible code and a proven edge are not the same thing. A pipeline can run cleanly while still using a stale quote, a changed model, an unclear decision time, or a price that was observed but never executable.
Those are process and software gaps: who owns the number, when the price was real, which model version ran, and what pass/fail rule authorizes the next step. The Quant Assurance Framework (QAF) is built to close gaps like that so a capital decision rests on evidence you can defend.
Conclusion
Two MLB models made money in the corrected 2022–2025 replay. That is worth testing once more. It is not enough to trust yet.
The profit was real in this sample. It still did not earn capital.
The disciplined move is to freeze the rules, collect new evidence, and let one fair test decide whether there is anything real to pursue.
Explore the full experiment
Open the complete slide deck in a new tab.
Enjoyed this post?
Subscribe for more research and trading insights through the BlackArbs newsletter.