Beating a Baseline Is Not Beating the Market
A model that beats guess-the-average and a model that beats the closing line are two different claims. Here is the test, and our own honest numbers.
Two Claims That Sound Identical
Here are two sentences a betting model can produce. One: this model predicts game outcomes better than a simple rule of thumb. Two: this model prices games better than the sportsbook does. They sound like the same boast and they are not remotely the same claim. The first is easy, and almost every competent model clears it. The second is the only one that can make anyone money, and very few models have ever demonstrated it. Most model marketing quotes the first number and lets you hear the second.
What a Baseline Actually Is
A baseline is the deliberately dumb rule you measure a model against. In sports the standard ones are trivial: always pick the home team, always pick the favorite, or assume every team scores the league average. They are useful precisely because they are not worthless. Home teams really do win more often, so always-home already posts a respectable-looking score. In our own stored college football sample the home team wins 59.7 percent of the time, and in our stored baseball sample it is 52.2 percent. A model that cannot beat those numbers has learned nothing at all.
Why Beating One Proves Almost Nothing
The reason a baseline is a low bar is that the market cleared it a long time ago. By game time, the closing line has absorbed every roster note, every injury report, and every dollar from every bettor who thought they knew something. Team quality is already inside that number before you arrive. So a model that beats always-home has proved it knows the good teams are good, which the price also knew. Beating a baseline tells you the model is not broken. It says nothing whatsoever about whether it can beat a price.
The Market Is a Much Harder Test
Testing a model against the market needs something the naive test never asks for: a stored price for the same game, on the same number, captured before the result existed. Then you ask whether the model would have taken the side the closing number ended up favoring, over a sample large enough to mean something. You also have to clear the sportsbook margin, because the two sides of a line add up to more than 100 percent in implied probability. A model that is right slightly more often than the close, and still pays that margin, is a losing model. That is the real bar, and it sits a long way above always-home.
Our Own Numbers, Told Straight
We run score projections for several sports, so here is what we can and cannot say about them. Against a naive baseline they look good: our college football module gets the sign of the margin right 71.3 percent of the time, against that 59.7 percent always-home rate. That is a real improvement, and it is also exactly the kind of number we just told you not to be impressed by. The same table has our baseball module at 53.6 percent against a 52.2 percent home rate, which is indistinguishable from noise. Neither figure has been tested against a price, so neither one is evidence that the model beats a sportsbook.
The One Market Test We Have Run
Once, on one sport. In WNBA, where we had 43 games carrying both a projection and a comparable stored closing price, the projection went 12-16 against the close. That is a losing result, and 43 games is far too small to conclude anything in either direction, including that it failed. For the other sports the count of comparable games is currently zero, and the reason is boring rather than damning: we only began storing prices on July 14, 2026, and the stored football and basketball seasons had already finished by then. The first real evidence needs a priced game to finish, which means NFL preseason on August 6 and college football on August 29.
What Runs in Shadow, and What No Longer Does
For a long time the answer was everything. The projection was computed daily, stored, and displayed beside the pick so a reader could see when our own model disagreed with our own call, and nothing selected a bet from it. One cell has since opened: baseball moneylines, where the projection now sources real published picks. Every other sport and every other market is still shadow only, computed and shown and never staked. Cells open one at a time, only after that exact sport and market has cleared the test above, and the public record is where you check whether we were right to open it. Showing a number and betting a number stay separate decisions.
The Question to Ask Any Model
Next time a service quotes an accuracy figure, ask the one question that separates the two claims: accurate against what? If the comparison is a coin flip, the home team, or last season average, the number is a sanity check that the code runs. If the comparison is the closing price on the same game, captured before the result, you are looking at something worth paying attention to. Then ask how many games are in that sample, because a few dozen is a story and not a finding. Most services will have no answer to the first question. We have one, and right now it is a losing result on 43 games.
Related reading
- How Are AI Sports Picks Actually Graded?
- Can AI Predict Sports Betting? How AI Betting Analysis Actually Works
- Best AI Sports Betting Apps: The Five Checks That Separate Them