Every night during the season, our Alabama Rankings page re-rates every AHSAA varsity girls flag football team in the state: nearly 150 teams across 6A, 5A, 1A–4A, and the private-school division, all on one scale. This article explains, in plain English, where those numbers come from, what they can tell you, and what they can't. No formulas required, though we'll tell you exactly what's under the hood.
One Scale for the Whole State
The most common question we get is some version of: how can you compare a 6A team to a 2A team they'll never play? The answer is that the schedule is far more connected than it looks. Teams play across classifications constantly, in tournaments, area crossovers, and non-region games. Every one of those games is a thread tying two teams together, and with hundreds of games a season, the threads form a web that touches everyone. Rate all the games at once and every team lands on the same ladder, even pairs that never meet directly.
That's why our rankings are statewide rather than division-by-division. The divisions are how the playoffs are organized; they aren't separate universes of ability, and the data proves it every season.
The Power Rating, in Plain English
The main number on the rankings page is the Power Rating. It answers one question: how many points better (or worse) than an average Alabama team is this team? A rating of +12 means we'd expect that team to beat a perfectly average team by about 12 on a neutral field. A rating of −5 means they'd be about a 5-point underdog to the same average team.
The useful property is that the gap between two ratings is a predicted margin. Team A at +12 versus Team B at +5? The model expects A by 7 on a neutral field, plus a little under 2 points if A is at home. That's what home field has historically been worth in this data.
Where do the numbers come from? Not from a formula applied team by team. The model looks at every rated game in the state simultaneously and finds the single set of ratings that best explains all of the actual margins at once. That "all at once" is what handles strength of schedule automatically: beating a +10 team by 6 moves your rating very differently than beating a −10 team by 6, because the model knows exactly who you played. There's no separate SOS adjustment bolted on afterward; opponent strength is baked into the fit itself.
Blowouts are capped
Flag football produces extreme scores. One 2025 team outscored its opponents 725–14 on the season. If we fed raw margins into the model, a handful of 46–0 games would drown out everything else. So margins are capped at 28 points: winning by 40 counts the same as winning by 28. Style points beyond four touchdowns tell us nothing about how a team handles a close game, and the cap means running up the score does nothing for a team's rating.
Last season counts, then fades
Early in a season there aren't enough games to rate anyone reliably, so each team starts from its final rating from the previous season, pulled 30% of the way back toward average (teams change more between seasons than within them; graduation is real). As this season's games accumulate, that starting point matters less and less, until the rating is effectively all current-season. Nothing switches over on some arbitrary date; the influence fades out on its own.
Because reasonable people can disagree about using last season at all, the rankings page also shows a Cold Start version of every rating, where all teams start equal and only this season's results count, plus a Seeding Impact column showing the difference. A big positive gap means a team's ranking is still leaning on last year's résumé; a negative gap means they're outplaying it.
The Elo Rating
The second number is an Elo rating, the same family of system used in chess and in FiveThirtyEight's old sports models. Where the Power Rating re-solves the whole season from scratch every night, Elo works game by game: two teams put rating points on the table, the winner takes some, and the size of the transfer depends on how surprising the result was and how big the margin. Beat a team you were supposed to beat and little changes; upset a top-ten team by three scores and you take a serious bite of their rating.
An average team sits at 1500. Elo reacts faster to recent form than the Power Rating does, which makes it a better momentum gauge and a slightly worse season-long summary. When the two systems disagree about a team, that disagreement is usually telling you something, most often that the team is trending sharply in one direction. Like the Power Rating, Elo has a cold-start version that ignores last season.
What the Win Probabilities Mean
The matchup predictor on the rankings page turns rating gaps into win probabilities. The logic: the predicted margin is the model's best guess, but flag football games are noisy. Historically, actual results scatter around the prediction by about 9–10 points in either direction. Feed that scatter into the standard bell-curve math and a predicted margin becomes a probability.
| Predicted Margin | Win Probability | In Words |
|---|---|---|
| 3 points | ≈ 62% | Barely more than a coin flip |
| 7 points | ≈ 77% | Clear favorite, loses 1 in 4 |
| 14 points | ≈ 93% | Heavy favorite, upsets happen |
Notice how slowly certainty arrives. A touchdown favorite still loses about a quarter of the time. The randomness is real: a short game with a handful of possessions leaves plenty of room for the worse team to win it. Any ranking system that tells you a 7-point favorite is a lock is lying to you.
Questions We Get Every Season
"Team A beat Team B. Why is B ranked higher?"
Because the rating reflects a whole body of work, not one result. If B has played a brutal schedule close and A has feasted on weak opponents, one head-to-head game (which the model does see and does count) may not outweigh everything else. Head-to-head feels decisive to fans; statistically, it's one noisy game like any other.
"We won. Why did our rating go down?"
Because you won by less than the model expected, or the nightly re-fit learned something new about your past opponents. If you're rated 20 points better than an opponent and win by 6, that's evidence the gap is smaller than 20. The rating moved toward what the result actually showed.
"What does the provisional tag mean?"
Fewer than four rated games. The math still produces a number, but with that little information it's closer to an educated guess than a measurement, and the tag says so. It disappears on its own as games are played.
"Which games count?"
Ratings use completed varsity games between AHSAA-registered flag programs, pulled nightly from the same schedule-and-scores system the schools themselves report into. Records shown on the page count every completed game, including exhibitions and out-of-state opponents, because that's the team's real record. Only rated games move the numbers, though.
What the Model Doesn't Know
The model sees scores. It does not see injuries, transfers, weather, a starting quarterback's sprained ankle, or the fact that your best athlete was at a track meet last Tuesday. When you know something the scoreboard hasn't absorbed yet, you genuinely know something the model doesn't. That gap is why a coach using a model beats either one working alone.
The Short Version
The Power Rating is "points better than an average Alabama team," fit to every game in the state at once, with strength of schedule built in, blowouts capped at 28, and last season fading out as this season fills in. Elo is the faster-twitch second opinion. The win probabilities are honest about how random a short game is. See the current numbers on the Alabama Rankings page, or try any matchup in the predictor and check the model's work all season.