The first few weeks of a new college football season are packed with fun for the whole family. This year, we’ve already had a secret clock, a castaway dramatically returning to town, and an Australian-style onside drop-kick of dubious legality. There have been international excursions to Dublin and London — I managed to attend the Union Jack Classic, and can confirm it was surprisingly great fun — plus a planned-but then-mysteriously-cancelled trip to Rio. Perhaps nothing has been more shocking than the United States Senate actually passing a bill (good to know where their legislative priorities lie) in an attempt to impose some kind of order on the chaos.

However, one thing remains the same every season: the September surprises. At this point in the season, there are typically around 20 undefeated teams, including some that everyone could have guessed, and some that no one would have ever dreamed of. In 2026, the count post-September 26th is precisely 22: Indiana, Miami, and Georgia, of course, but also the likes of Florida, UCLA, and (shock! horror!) UMass. Even FBS newcomers North Dakota State are still undefeated, though that may not really qualify as a surprise; they’ve hardly ever lost before, so why start now?

Of course, most of those 22 won’t survive the season (or maybe even next week) and so begins the flying circus of schedule strength, victory margins, and transitive win chains. We pore over every detail, looking for signs of who will still be a real playoff contender come December. Did they challenge themselves with non-conference power matchups, or settle for FCS “money games”? Is there even a Power 4 anymore, or really just a Power 2? Sure, Texas is undefeated, but they got lucky beating Ohio State by one and Tennessee by three. You can’t just ride one-possession wins to a national title! And if Oregon beat USC, but Oklahoma State beat Oregon, and Tulsa beat Oklahoma State, then what exactly does that make Tulsa and USC? (Absolutely nothing?)

The implicit claim underlying all of this is that team performances in September carry some signal about who’s going to win in November and December — and in particular, that teams which stumbled their way to a perfect September will have a harder time than teams which cruised through it. So in this article, we’ll sort out the eventual contenders from the pretenders over the last five years, then compare how they actually played across the opening weeks of the season. The question is simple: can we read the tea leaves in September and spot the frauds before they lose? Or is the “eye test” just grading wins on style points that don’t actually tell us anything?

The data

As a starting point, we need some measure of expected preseason team strength. This does two jobs for us: it gives us a proxy for opponent quality early in the season, and it gives us a benchmark for judging whether each team ultimately beat or fell short of expectations. We’ll look at the last five years, neatly avoiding the COVID-shortened 2020 season. Using historical data from Covers, originally sourced from BetMGM, we can get preseason over/under win predictions for every FBS team, along with their final totals. Then, using SportsDataverse data, we can pair those season-level expectations with game-by-game schedules and results.

For this analysis, we’ll focus on FBS teams that made it through September undefeated. Across the last five seasons, that’s 107 teams — an average of 21 to 22 per year, almost exactly in line with this year’s 22. Before we can compare the quality of the contenders’ wins to the quality of the pretenders’ wins, we need to actually define who falls into each of those groups. The plot below shows each of the 107 teams’ preseason expected wins compared with their final actual wins, looking only at the regular season (no conference championships or postseason).

From this, we’ll specify two deliberately contrasting groups:

  • Contenders: Teams that finished with ten or more wins and beat their preseason expectation by more than two. Ten wins is the traditional bar for a good season, but the second condition matters too: we don’t just want Alabama or Ohio State doing exactly what everyone expected. These are teams that started September hot, often surprising people, and then kept it rolling all the way into playoff contention. The three labelled green points are the biggest outperformers of the last five years: 2022 TCU, which went undefeated en route to the national title game (we won’t discuss what happened there), plus 2024 Indiana and BYU, whom basically no one saw coming.

  • Pretenders: Teams that finished with nine or fewer wins and failed to beat their preseason expectation by at least one. Some of these teams ultimately did meet expectations, but remember, that was after starting 4–0 or 5–0. If a perfect September didn’t translate into even a modest improvement over what was expected in August, then something went wrong along the way. The three labelled red points are the biggest frauds: 2023 USC, whose only win after September was by one point in Berkeley; 2024 Liberty, which fell to Kennesaw State and Jacksonville State; and 2025 Maryland, the true poster child for this entire project, which started 4–0 and then went 0–8.

The result is nicely balanced: 31 Contenders and 30 Pretenders, both with an average preseason expectation right around 7.5 wins, so we’re not simply falling victim to causal confounding by comparing obvious heavyweights with mediocre teams that got lucky for a month. The remaining 46 teams, shown in gray, sit somewhere in the middle: they neither over- nor under-performed enough to tell us much, so we’ll leave them out and focus on the two groups where hindsight gives us a clear verdict.

The statistical approach

Now that we have our two sets of teams, what exactly are we looking for? We want to know whether the contenders’ September wins actually looked different from the pretenders’ wins. The obvious place to start is margin of victory: did the pretenders simply win by less? Of course, I can already hear your objection — some teams spend September playing Samford, while others get Ohio State. So, to account for differences in opponent quality, we’ll look at how margin of victory changes with respect to expected opponent strength, rather than looking at margin alone. We’ll also drop FCS opponents entirely, since those games truly tell us nothing.

Specifically, we’ll first find the best-fit relationship between margin of victory and expected opponent wins for all 210 games in our sample. We’ll then fit it separately for the 104 Pretender games and 106 Contender games. If either group really played differently, its best-fit line should clearly diverge from the overall one, either in terms of the intercept or the slope. For example, if the Pretenders consistently won by less, the red line should sit below the black; if they held up against weaker opponents but struggled as the competition improved, the red line should fall away towards the right.

To quantify this, there’s an F-test for whether the two group-specific lines are meaningfully different from the common fit; the details are remarkably dull even by the usual standards of hypothesis testing, so I’ll tuck them in the appendix.

The results

In the plot below, each point represents a single September game. For example, a green point at (6.5, 30) would mean that a Contender faced an opponent with a preseason expected win total of 6.5 and won by 30 points. The black line shows the best-fit relationship across all 210 games in our sample, with the gray band showing its uncertainty interval; the dashed red and green lines show the same best-fit relationship separately for the 104 Pretender games and the 106 Contender games.

And all three lines look essentially the same! The red and green lines both fall within the error band of the combined fit, and the F-test gives a p-value of 0.5156, providing no evidence that the relationships are meaningfully different. In other words: whether you ultimately end up being a Contender or a Pretender, your margins of victory in September tend to look roughly the same — you defeat weaker teams by an average of 20-30 points, and stronger teams by an average of 10-20 points.

However, margin of victory across all of September is only one angle, and it’s the broadest one. What if the teams that eventually fall off have systematically weaker defenses, because defense wins championships? Or maybe everyone gets a pass for the chaos of the opening weeks, but the warning signs start to emerge once September settles down. No problem! We can repeat the same analysis for points scored and points allowed, and then again using only Weeks 3 to 5, excluding the games from Weeks 0 to 2 (yes, college football is zero-indexed; they must be Python programmers).

However, the trend lines remain essentially indistinguishable. The plots below show each comparison along with the corresponding F-test p-values, none of which even threaten the social-science “borderline significance” threshold of 0.1.

What does it all mean?

At the risk of giving coaches more press-conference ammunition, it may actually be true that a win is just a win, especially early in the season, no matter how ugly it looks. Of couse, a loss is also a loss, and it tells you something meaningful about the team. But once we restrict ourselves to teams that escaped September undefeated, there seems to be surprisingly little basis for separating the “good wins” from the “bad wins”.

I tried slicing the data a few additional ways beyond the plots above, as I kept finding reasons not to trust the result. Maybe an SEC team projected for seven wins isn’t really comparable to a Sun Belt team projected for seven wins, so let’s look at just major-conference teams against major-conference opponents. Maybe the effect only shows up among the elite teams clustered at the top of the chart, so let’s look at those separately too. But the answer stubbornly stayed the same: no meaningful difference in margin of victory, offense, or defense trends, including between the early and late parts of September. There are only so many dimensions along which you can search the data before accepting that it simply may not be hiding any secrets.

At the end of the day, over the last five years, the hot-start teams that eventually challenged for the playoff looked a lot like the ones that later fell away — both generally handled weaker opponents, both had the occasional close shave, and both sometimes blew good teams out. We spend September searching for warning signs because we know some undefeated teams are going to collapse. But if there’s a canary hiding in the early-season coal mine, it’s remarkably quiet.

Appendix

Code and data are available on Github. I'm always happy to discuss collaboration ideas; my contact information can be found under the CV tab on my website.

The precise statistical method we apply is a standard F-test for nested linear models. For each plot, we first fit a single best-fit line to all of the games, then compare it with a more flexible model that allows the Contenders and Pretenders to have their own intercepts and slopes. The null hypothesis is that the two groups follow the same underlying line; the alternative is that either their average level (intercept), their relationship with opponent strength (slope), or both are different.

The single best-fit line requires only two parameters, while the separate lines require four, so some improvement is inevitable from the two additional degrees of freedom. The F-statistic asks whether the improvement in explained variation, relative to the remaining noise in the data, is large enough to justify the added complexity. A small p-value would suggest that the groups really do follow different trends; the consistently large p-values we observe indicate that allowing separate slopes and intercepts does not materially improve the fit.