Building Your Own Projection Models
Every post in this series has told you to build your own number before looking at the market. Seventy-five posts later, it is time to be specific about how.
The sport-by-sport roadmap ended with Post 75. Five sports, sixty-five posts, and one instruction repeated in every one of them: build your own projection first, then look at the price.
That instruction is easy to state and harder to execute, and most bettors never actually do it. They read a rating, glance at a line, and construct a story about why the line is wrong. That is not projection. It is anchoring with extra steps.
This post is about the alternative. Not a proprietary black box, not machine learning, not something requiring a statistics degree. A structured, repeatable method for turning inputs you can gather into a number you can defend, built in a spreadsheet, maintained across a season, and honest about what it can and cannot do.
01Why You Need Your Own Number
The case is simpler than it sounds. Without an independent projection, you have no way to identify a mispriced game, because you have nothing to compare the price against.
What most bettors do instead is evaluate the line directly. They look at Kansas -8.5 and ask whether that feels right. The problem is that the number itself carries information, and once you have seen it, your judgment is contaminated by it. Psychologists call this anchoring, and we covered it in Post 10. It is not a weakness you can willpower your way out of. It is how human estimation works.
The only reliable defense is procedural. Produce your number before you see theirs. Write it down. Then compare.
A second benefit matters as much as the first. A written projection is falsifiable. If you say a team should be favored by six and the market says nine and the game lands at eleven, you learned something specific. If you merely felt the line was too high, you learned nothing, because your position was never precise enough to be wrong.
An opinion you did not write down before seeing the price is not a projection. It is a reaction, and reactions cannot be graded.
— Bang the Over02What a Projection Model Actually Is
Strip away the mystique and a projection model is three things:
- A set of inputs. Numbers describing each team or competitor.
- A rule for combining them. Arithmetic that turns inputs into an expected outcome.
- A conversion. Turning that expected outcome into a spread, a total, or a price you can compare to the market.
That is the whole thing. A professional syndicate's model and a spreadsheet you build this weekend differ in sophistication, data quality, and speed. They do not differ in structure.
What matters is that the model be yours, in the specific sense that you understand every step and can explain why each input is there. A model you downloaded and do not understand is worse than no model, because it produces confident numbers you cannot interrogate.
03The Three Model Families
Assign every team a single number on a points-above-average scale. The projected margin is the rating difference plus a venue adjustment. Simple, transparent, fast to maintain, and the structure books themselves use for the bulk of their pricing. This is where every bettor should begin and where many should stop.
Instead of one rating, project the components separately: possessions and efficiency in basketball, pass and rush efficiency in football, run environment in baseball. More inputs, more places to be right, and more places to be wrong. Worth building once your power rating is stable and you understand where it fails.
Run the game thousands of times using distributions rather than point estimates. The output is not a single margin but a full probability distribution, which is what you actually need for totals, moneylines, and alternate lines. Powerful, and genuinely unnecessary until the simpler families are working.
The honest sequencing advice: build a power rating, use it for a full season, and only add complexity where you can point to a specific, repeated failure that the added complexity fixes. Complexity added for its own sake makes a model harder to debug without making it more accurate.
04Building a Power Rating From Scratch
Here is a working method you can implement in a spreadsheet this week.
Step one. Establish a starting rating. Take last season's final ratings, regress each toward the league average by roughly a third, and adjust for known roster change. Regression to the mean is not pessimism, it is arithmetic: extreme results contain more luck than moderate ones.
Step two. After each game, calculate the performance rating. Take the actual margin, adjust for venue, and compare to what your ratings predicted. If you projected Team A by 6 at home and they won by 14, they outperformed by 8.
Step three. Update both ratings. Move the winner up and the loser down by a fraction of that surprise. A common starting point is a quarter of the difference, split between the two teams. Team A gains, Team B loses, and the total rating in the league stays roughly constant.
Step four. Cap the input. Do not let a 40-point blowout move ratings four times as far as a 10-point win. Cap the margin you feed the update at something like 20 points, for reasons we cover in Section 06.
That is a functioning Elo-style rating system, it takes an afternoon to build, and it will produce numbers within a few points of the public systems within about a dozen games.
The update fraction is the single most important dial in the model. Set it too high and your ratings chase last week's result. Set it too low and they never notice a team has genuinely changed. Start at a quarter, then test a third and a fifth across a full season of past results and see which produced the best predictions. That test is your first real piece of modeling work.
05Inputs: What Belongs and What Does Not
The temptation with a first model is to include everything. Resist it. Every input you add is another thing that can be wrong, another thing to maintain, and another opportunity to fit noise.
Inputs that earn their place:
- Margin or efficiency, opponent-adjusted. The core signal in every sport.
- Venue. Home advantage, ideally venue-specific rather than league-average.
- Availability. Who is actually playing. The largest single manual adjustment in most sports.
- Pace or possessions, where the sport has meaningful variation in it.
- Rest and schedule, as a small structured adjustment rather than a feeling.
Inputs that usually do not:
- Win-loss record. Already contained in margin, and worse.
- Recent form as a separate term. If your update rule is calibrated, recency is already handled.
- Head-to-head history. Small samples across different rosters. Almost always noise.
- Streaks of any kind. The classic case of a pattern that exists in the data and predicts nothing.
- Anything you cannot update reliably every week. A brilliant input you stop maintaining in January is worse than a mediocre one you keep.
06Why You Cap Margin
A 45-point win and a 12-point win are not proportionally different pieces of evidence, and treating them that way is one of the most common modeling errors.
Large margins are contaminated. Garbage time is played by bench units. Late fouling inflates basketball margins. A defensive touchdown turns a competitive football game into a blowout on the scoreboard without changing what the game told you about either team. The information content of margin flattens out well before the final score does.
Every serious public rating system caps or damps margin for this reason, and yours should too. A practical approach is to apply the full margin up to a threshold and then compress everything beyond it. The exact threshold matters less than having one.
07Home Advantage as a Model Parameter
Most models apply a single home-field constant per sport. That constant is the average of a distribution, and almost no venue sits at the average.
The upgrade path is straightforward and it is one of the few places an individual can genuinely outperform a large model. Start with the league average. Then, over a season, track how each venue performs relative to that average and build a venue-specific table. We recommended exactly this for college basketball in Post 64, and the logic generalizes.
Be disciplined about sample size. One season of home games is a small sample, and a venue that looked worth five points may regress to three. Update your table gradually rather than replacing it wholesale each year.
08Converting a Rating Into Every Market
A power rating produces an expected margin. Every market you bet is a transformation of that.
To a spread: the expected margin is the spread. Direct.
To a moneyline: convert the margin to a win probability using the sport's margin standard deviation, then convert the probability to a price. For college basketball we used a standard deviation near 11 points in Post 66. Every sport has its own value, and finding it is a useful exercise: take a season of results, compare actual margins to closing spreads, and take the standard deviation of the differences.
To a total: your model must project the two scores separately rather than only their difference. In possession-based sports that means projecting pace and efficiency, as we did in Post 65.
To alternate lines and props: this is where simulation earns its keep, because you need the shape of the distribution rather than its center.
A model that projects only margin cannot price totals, and a model that projects only totals cannot price sides. If you find yourself deriving one from the other with a rule of thumb, you have left the model and started guessing. Build the components you need for the markets you actually bet.
09Backtesting Without Fooling Yourself
Backtesting means running your model against historical games to see whether it would have made money. It is essential and it is the easiest place in this entire post to deceive yourself.
The rules that keep it honest:
- Test against closing lines, not opening lines. The closing line is the market's best estimate. Beating openers is easy and means little.
- Use only information available at the time. If your model uses season-long stats to predict a November game, you have used the future to predict the past. This is the most common and most fatal backtesting error.
- Hold out data. Build on one set of seasons, test on a different one you never looked at while building.
- Measure closing line value, not just record. Per Post 9, a model that consistently beats the close is real. A model with a good record and no CLV got lucky.
- Include the vig. A 52 percent model is a losing model at standard pricing.
- Be suspicious of good results. If your first backtest shows a 58 percent win rate, you have a bug. Find it.
10Overfitting: The Central Danger
Overfitting is building a model that explains the past beautifully and predicts the future badly. It is the failure mode of nearly every homemade model, and it happens through a completely reasonable-feeling process.
You build a model. It misses some games. You add a variable that would have caught them. It misses others. You add another. Each addition improves the backtest. Eventually you have a model that would have been magnificent last season and is worthless this one, because you fitted the noise rather than the signal.
The defenses:
- Fewer inputs. Every variable should have a mechanism you can articulate before you test it, not a correlation you found afterward.
- Held-out data. The only real test is performance on games the model has never seen.
- Suspicion of complexity. If a simpler version performs nearly as well, use the simpler version.
- Forward testing. Run the model on live games without betting for several weeks. Slow, and the only fully honest test.
11Where the Model Beats You, and Where You Beat It
This is the most practical framing in the post, because a model is not a replacement for judgment. It is a division of labor.
| Task | Who wins | Why |
|---|---|---|
| Consistency across many games | Model | Never tired, never bored, never emotional |
| Opponent adjustment | Model | Simultaneous estimation is not something a person does in their head |
| Resisting recency | Model | Weights the season correctly if calibrated |
| Availability and roster change | You | Requires reading news and judging what it means |
| Style and scheme interaction | You | Requires watching, and does not generalize across a league |
| Venue specifics | You | Small samples where local knowledge beats aggregate data |
| Knowing when the model is blind | You | The model has no concept of its own error bars |
The correct workflow follows directly. The model produces a baseline. You apply adjustments in the areas where you have information it does not. The result is your number.
12Tools: Spreadsheet or Code
Spreadsheets are the right starting point for almost everyone. Visible logic, easy debugging, no setup, and entirely adequate for a power rating across a few hundred teams. Most bettors never need more, and we cover tracking spreadsheets specifically in Post 85.
Code becomes worthwhile when you need to pull data automatically, run simulations, or handle more games than a spreadsheet can manage comfortably. Python with a data library is the standard path, and the learning curve is real but not steep.
The wrong reason to move to code is that it feels more professional. The right reasons are automation of data collection and simulation. If neither applies to you yet, a spreadsheet is not a compromise.
13Maintaining a Model Through a Season
A model is not a project you finish. It is a routine you run.
- Update ratings after every slate. Falling behind compounds and is miserable to reconstruct.
- Log every projection. Your number, the closing number, and the result. This is the raw material for every improvement you will make.
- Review monthly against the closing line. Not against results. CLV is the fast signal.
- Do not adjust mid-season based on a bad month. Note the failure, keep running, and revisit in the offseason with a full sample.
- Rebuild in the offseason. That is when you change parameters, add inputs, and re-backtest, with a complete season of new data.
14The Section Roadmap: Posts 76 Through 100
This final section covers the craft rather than any single sport. Twenty-five posts, ending in a complete synthesis of all one hundred.
- 76Building Your Own Projection Models (you are here)
- 77Line Shopping at Scale
- 78Multi-Sportsbook Strategy
- 79Arbitrage Betting Fundamentals
- 80Middling Opportunities
- 81Hedge Construction
- 82Promotions and Boosted Markets
- 83Tax Considerations for American Bettors
- 84Bookmaker Limits and Account Management
- 85Tracking Tools and Spreadsheets
- 86Same-Game Parlays: The Math
- 87Exotic Bets and Specialty Markets
- 88Information Networks and Beat Reporters
- 89Reading Sharp Action
- 90The Psychology of Discipline
- 91Managing Variance and Drawdowns
- 92Going Full-Time: Realistic Assessment
- 93Building a Professional Network
- 94Year-End Review and Optimization
- 95Cross-Sport Portfolio Construction
- 96NHL Betting Crash Course
- 97Soccer Betting for American Bettors
- 98Tennis and Golf Betting Fundamentals
- 99Combat Sports Betting Framework
- 100The Complete Bang the Over Synthesis (Series Capstone)
15Common Modeling Mistakes
- Looking at the line first. Everything downstream is contaminated. This is the whole reason models exist.
- Using future information in a backtest. Season-long stats to predict an early-season game is the classic version.
- Adding variables to fix misses. That is overfitting, and it feels exactly like improvement.
- Not capping margin. Blowouts contaminate ratings.
- Testing against openers. Beating the closing line is the only meaningful bar.
- Ignoring vig in evaluation. Break-even is roughly 52.4 percent at standard pricing, not 50.
- Trusting the model over obvious information. If it likes a team missing two starters, the model is wrong, not brave.
- Abandoning it after a bad month. A hundred games is noise. Judge on CLV and a full season.
16The Bigger Picture
There is a persistent fantasy in sports betting that somewhere there exists a model good enough to print money, and that the work is finding it. That is not how this functions.
A well-built personal model will land close to the market on most games, which is the correct outcome and the one people find deflating. The market is the aggregate of everyone's models plus everyone's information. Matching it is the baseline achievement, not a failure.
The value is in the small number of games where your model and the market disagree meaningfully, and where you can name the reason. That reason is almost never that your arithmetic is better. It is that you know something the aggregate does not: a roster situation, a venue quirk, a style interaction, a league you actually watch.
The model's job is to make those moments visible. Without it, a game where you have a real edge and a game where you have a feeling look identical.
◆ Final ThoughtsWrite the Number Down First
If this post produces one behavioral change, make it this: before you look at a line, write your number in a place you cannot edit afterward. A notes app, a spreadsheet cell with a timestamp, anywhere that makes revision visible.
That single habit converts betting from a series of opinions into a series of testable claims. Most bettors never make that conversion, which is why most bettors cannot tell you whether they are actually good at this. Within a season you will be able to.
In Post 77 we take the number you just built and go looking for the best available price on it. Line shopping is the least glamorous habit in this series and it produces more measurable return per hour invested than anything else you will do, which makes the fact that most bettors skip it genuinely remarkable.
- Build your number before you see the price. Anchoring is not a weakness you can overcome by trying harder, only by procedure.
- A model is inputs, a combining rule, and a conversion. Nothing more mysterious than that.
- Start with a power rating. Regress last year's ratings, update by a fraction of each game's surprise, and cap the margin you feed in.
- Fewer inputs is better. Every variable needs a mechanism you can state before you test it.
- Backtest against closing lines, use only information available at the time, hold out data, and include the vig.
- Overfitting feels exactly like improvement. If adding a variable fixes last season's misses, be suspicious rather than pleased.
- Divide labor. The model handles consistency and opponent adjustment. You handle availability, style, venue, and knowing when it is blind.
- Matching the market on most games is success, not failure. The edge lives in the few where you disagree and can say why.
The arithmetic of half points and cents across a full season, how many books you actually need, reading a price screen quickly, devigging across books to find the true number, why market-making and retail books disagree systematically, and the execution protocol that turns all of it into measurable return.
Continue the 100-part Bang the Over series for sport-specific strategy, advanced edges, and pro-level American sports handicapping.
Continue the Series