System notes · TensorFlow.js · ESPN box scores
How the ranking model actually works
This is not a black-box “AI ranking.” It is a small multilayer perceptron, trained from scratch in TensorFlow.js on the server, whose job is to map a 25-dimensional box-score vector onto a standardized point differential, then average those predictions by team. The real work happens in computeSeasonRankings, inside a Node.js route that can run for up to 300 seconds.
- 01Roster the league32 NFL clubs, or FBS via ESPN group 80
- 02Enumerate weeksRegular season only · seasontype=2
- 03Pull box scoresSummary API · 10 concurrent workers
- 04Featurize25-D vector per team-game
- 05Z-scorePopulation μ, σ across the season
- 06Fit MLP25→128→64→32→1 · Adam · MSE
- 07RankNFL mean-pool · FBS SRS + Top 25
- 08Shrink to priorBlend last season while this one is unfinished
01 · Ingest
Where the bytes come from
Every number in the model originates from ESPN’s undocumented but public JSON APIs. There is no proprietary play-by-play dump, no Next Gen Stats, no PFF grades. We speak HTTP to three families of endpoints, with a 3-attempt retry and linear backoff (400 ms, 800 ms), and we ask Next to cache each response for 24 hours.
| Purpose | NFL | FBS (CFB) |
|---|---|---|
| Team universe | site.api …/nfl/teams?limit=32 | core …/groups/80/teams (FBS) |
| Week calendar + scoreboard | …/nfl/scoreboard?seasontype=2 | …/college-football/scoreboard?groups=80 |
| Per-game box score | …/nfl/summary?event=ID | …/college-football/summary?event=ID |
| Reference ranking | ESPN FPI (fpi, fpirank, W-L-T) | AP Top 25 (live site poll, else latest week / preseason) |
NFL teams are the 32 clubs on the league teams endpoint, sorted by display name. College football is not “every NCAA team”: the roster is ESPN group 80 (FBS), loaded from the core API season group-teams collection. Scoreboards are pinned to the same group. The public /teams?groups=80 site list ignores the filter, so we do not use it. Division III clubs never enter the ranking universe. If an FBS club plays an opponent that is not in group 80, we still ingest the FBS side of that box score; the missing opponent is later treated as a cupcake in the SRS pass.
Season slicing is strict. We only keep regular season games (seasontype = 2). Preseason and playoffs never enter the tensor. A game is admitted only if ESPN marks it completed (status.type.completed or STATUS_FINAL). Week numbers are read from the scoreboard calendar; if that parse fails we fall back to 18 NFL weeks or 16 CFB weeks.
Event ids are collected across every week, dumped into a Set so a game cannot be trained on twice, then fetched with a worker pool of 10 concurrent summary requests. That pool is why the first CFB run is slow: you are hydrating on the order of a thousand game summaries, not running a giant GPU job.
02 · Parse
One game becomes two training rows
A summary payload has a two-team box score plus a header with final scores. We emit two GameStats objects: team A vs B, and team B vs A, with opponent features mirrored. That duplication is load-bearing. It makes the marginal distribution of score and opponentScore identical across the season matrix (every point scored is also a point allowed, from the other row). That identity is what lets the target collapse to a scaled margin, as we'll see in §04.
ESPN stats are annoyingly heterogeneous. Some fields are numeric values. Efficiency stats arrive as display strings like 7-14 or 7/14. We split on [-/] and keep the numerator (makes). Completion percentage is reconstructed from completions and attempts rather than trusted as a pre-baked percentage. Penalties arrive as a single count-yards blob. Possession is either MM:SS or a raw second count. Missing stats become 0, not null — the network never sees NaNs.
Special teams are the flakiest layer. Kick-return yards, punt-return yards, field goals made, and punts are read from the team box score first. If those keys are empty, we walk boxscore.players, find the matching statistical group (kickReturns, puntReturns, kicking, punting), locate the column by key name, and sum across athletes. Slash stats like fieldGoalsMade/fieldGoalAttempts contribute only the left-hand number.
03 · Feature space
The 25-dimensional team-game vector
After parsing, time of possession is converted to seconds and the record is projected into a fixed-order float32 vector of length 25. There is no embedding layer, no opponent id in the tensor, no home/away bit, no week index, no rest days, no QB identity. The MLP can only “see” what happened on the stat sheet that afternoon. Opponent identity is stored alongside the row and used only after predict, in the FBS Simple Rating System pass.
Crucially, the two scoring columns are not in that vector. They are parsed and normalized alongside everything else, but they exist only to build the label. Feeding them in would hand the network the answer, and the MLP would collapse into an identity function on two slots.
Offense (own)
| Tensor slot | ESPN source | Notes |
|---|---|---|
| offensiveYards | totalYards | Net offensive yards from scrimmage. |
| passingYards | netPassingYards | Net passing yards (sacks already deducted by ESPN). |
| rushingYards | rushingYards | Rushing yards. |
| completionPercentage | completionAttempts | Derived: completed / attempts × 100, parsed from C-A or C/A. |
| firstDowns | firstDowns | First downs earned. |
| thirdDownConversions | thirdDownEff | Makes only, parsed from made-attempt strings. |
| fourthDownConversions | fourthDownEff | Makes only. |
| redZoneConversions | redZoneAttempts | Scores inside the 20, makes only. |
Defense (opponent box score)
| Tensor slot | ESPN source | Notes |
|---|---|---|
| opponentOffensiveYards | opp.totalYards | Yards allowed. Mirror of the other sideline. |
| opponentPassingYards | opp.netPassingYards | Net passing yards allowed. |
| opponentRushingYards | opp.rushingYards | Rushing yards allowed. |
| opponentCompletionPercentage | opp.completionAttempts | Opponent C% allowed through the air. |
| opponentFirstDowns | opp.firstDowns | First downs allowed. |
| opponentThirdDownConversions | opp.thirdDownEff | Opponent third-down makes. |
| opponentFourthDownConversions | opp.fourthDownEff | Opponent fourth-down makes. |
| opponentRedZoneConversions | opp.redZoneAttempts | Opponent red-zone scores. |
Special teams / kicking
| Tensor slot | ESPN source | Notes |
|---|---|---|
| kickReturnYards | kickReturnYards | Team total, else Σ player kickReturns.kickReturnYards. |
| puntReturnYards | puntReturnYards | Team total, else Σ player puntReturns.puntReturnYards. |
| fieldGoalsMade | fieldGoalsMade | Makes parsed from FG made/attempts; player kicking fallback. |
| punts | punts | Punt count; player punting fallback. |
Turnovers, flags, clock
| Tensor slot | ESPN source | Notes |
|---|---|---|
| turnovers | turnovers | Giveaways. ESPN’s team turnover total. |
| takeaways | opp.turnovers | Not a native takeaway field — opponent giveaways. |
| penalties | totalPenaltiesYards | Count half of the P-Y display string. |
| penaltyYards | totalPenaltiesYards | Yardage half of the same string. |
| timeOfPossession | possessionTime | MM:SS → seconds. Numeric ESPN values treated as seconds. |
Parsed but withheld from the tensor
| Column | ESPN source | Notes |
|---|---|---|
| score | header.competitors.score | Final points. |
| opponentScore | opp score | Final points allowed. |
04 · Normalize
Z-scoring the season, then defining y
All team-games in the season are flattened into one matrix of 27 parsed columns — the 25 that become the input tensor plus the two scoring columns. For each column j we compute the population mean and standard deviation (divide by n, not n − 1):
μⱼ = (1/n) Σᵢ xᵢⱼ σⱼ = √( (1/n) Σᵢ (xᵢⱼ − μⱼ)² ) x̃ᵢⱼ = (xᵢⱼ − μⱼ) / σⱼ with σⱼ := 1 if the column is constant
This is a season-relative z-score: a 350-yard passing game in a shootout year is not the same coordinate as 350 yards in a mud year. Normalization is global across the league-season, not per-team, so a team that always runs the ball still gets compared to the league centroid.
The supervision target is not win/loss and not a ranking permutation. It is the difference of the already-normalized scoring columns:
yᵢ = x̃ᵢ,score − x̃ᵢ,opponentScore
Because every game is stored twice (once per sideline), score and opponentScore share the same empirical moments: μs = μo and σs = σo. The expression therefore simplifies to yᵢ = (marginᵢ) / σscore. We are asking the net to predict standardized point differential. That is the entire learning problem.
Those two columns are then dropped, and only the remaining 25 are stacked into X ∈ ℝⁿˣ²⁵. The net has to infer the margin from how the game was played.
05 · Architecture
A four-layer MLP, ~13.7k parameters
The network is a vanilla TensorFlow.js tf.sequential dense stack. No dropout, no batch-norm, no residual connections, no attention. Hidden activations are ReLU. The output unit is linear (identity) because this is unbounded regression, not a classification.
input ℝ²⁵
dense 25 → 128 ReLU 3,328 params
dense 128 → 64 ReLU 8,256 params
dense 64 → 32 ReLU 2,080 params
dense 32 → 1 linear 33 params
≈ 13,697 trainable weights
- Optimizer: Adam, TensorFlow.js defaults (α = 10⁻³, β₁ = 0.9, β₂ = 0.999, ε = 10⁻⁷).
- Loss: mean squared error on y. Quadratic penalty, so a 21-point miss is 9× as expensive as a 7-point miss. Blowouts dominate the gradient.
- Reported metric: MAE, which is easier to read in “standardized points” but is not what Adam is minimizing.
- Schedule: 50 epochs, batch size 32,
validationSplit: 0.2,shuffle: true(TF.js default). The last 20% of the shuffled rows are held out for the val curve; they are not held out of the eventual ranking pass.
Sample size vs capacity is the fun part. An NFL regular season is 272 games × 2 rows ≈ 544 examples. After the val split you have ~435 training rows to estimate 13.7k weights. That is an overparameterized interpolating regime. FBS is healthier (~146 teams × ~12 games, still doubled). We are not training a foundation model. We are fitting a flexible function of one season’s box scores, then immediately evaluating it on those same rows.
Weights are ephemeral. After fit we predict, convert the output tensor to a nested JS array, then dispose the input, target, and model and close a tf.engine() scope. Nothing is checkpointed to disk. Hit the endpoint again (or wait out the 24h cache) and you train a fresh draw of Adam from the default random initialization. Rankings can jitter slightly run to run; the cache is what makes the UI feel deterministic.
06 · Inference
Mean predicted margin, then strength of schedule
After training we push the entire season tensor — train and val rows together — back through the net. For team t with games Gₜ the raw strength is mean predicted margin:
sₜ = (1 / |Gₜ|) Σ_{g ∈ Gₜ} fθ(x̃_g)NFL stops there. FBS runs a Simple Rating System on top of those per-game margins so a 40-point win over a cupcake is not the same as a 40-point win over a contender. Ratings start at sₜ, then iterate ~20 times:
rₜ ← (1 / |Gₜ|) Σ_{g ∈ Gₜ} ( fθ(x̃_g) + r_opponent(g) )
r ← r − mean(r)If the opponent is not on the FBS roster (FCS / unknown), r_opponent is the 10th percentile of current ratings — a cupcake floor. Mean-centering each round identifies the system. Teams with zero completed regular-season games contribute nothing to fit; they are carried by the prior described in §07. FBS is then sliced to the Top 25 for the API, CLI, and UI.
Two interpretive caveats, because they matter:
- In-sample, not out-of-sample. The scoring columns are withheld from the input, so the net genuinely has to predict margin from yards, efficiency, turnovers, and clock. But it is evaluated on the same rows it trained on, including the validation split. This is a fitted summary of a season that has already happened, not a forecast of games it has never seen.
- SOS is graph-based, not poll-based. Opponent quality is the opponent’s own evolving rating, not AP votes or FPI. Home field, injuries, weather, and recency are still invisible. A week-1 demolition and a week-18 demolition are still exchangeable except insofar as the opponent’s rating differs.
So when our list disagrees with FPI or the AP poll, that is not always a bug: FPI is a predictive efficiency rating; AP is a human ballot. We are mean-pooling a box-score MLP’s reconstructed margin, then (for FBS) adjusting that margin for who it came against.
07 · Prior
What a season looks like before it has been played
A rating built from one September Saturday is not a rating. In week 1, most of the league has played zero games and would vanish from the list entirely, while the handful of teams that opened in week 0 would take every top slot on the strength of a single result. That failure is loud in FBS, where a 148-team roster can collapse to the sixteen teams that happened to kick off early.
So an unfinished season is not ranked on its own evidence. We detect it directly from the scoreboard: if any regular-season game is still scheduled or in progress, the season is live. Canceled and postponed games sit in ESPN's post state without completing, so they do not keep a finished season looking open forever.
When the season is live we run the whole pipeline a second time on the previous season and use it as a preseason prior. The two rating vectors come from independently initialized nets, so neither scale is meaningful on its own; each is standardized to zero mean and unit variance before they are combined. Every team is then shrunk toward its prior by how much it has actually shown you:
w = n / (n + 4) rating = w · thisSeason + (1 − w) · lastSeason
With n the games played so far, one game buys a team 20% weight on its own season, four games buy 50%, and twelve games buy 75%. A team that has not kicked off sits entirely on last year's rating instead of disappearing. A program with no prior at all — new to the division — starts at the league average, which is generous but never enough to reach a Top 25 built from real ratings.
A completed season skips this path entirely: no second pipeline run, no prior, no extra latency. The response carries priorSeason so you can always tell which of the two you are looking at.
08 · Serving
The HTTP path, the cache, and the fake HUD
The browser calls GET /api/rankings?season=YYYY&league=nfl|cfb. Seasons are clamped to 2000–2026. The handler is a Node runtime (not Edge) with maxDuration = 300, wrapped in unstable_cache under the key power-rankings-v6, revalidate 86,400 seconds. First visitor of the day pays for ESPN hydration plus 50 epochs of Adam. Everyone else that day gets JSON out of cache.
The sci-fi overlay — “UPLINKING ESPN SCOREBOARD,” the ring, the equalizer — is a client-side progress skin keyed off the fetch promise. Its stage list is not hooked to real training callbacks. model.fit runs with verbose: 0. We do not stream epoch loss to the browser.
Side-by-side on the home page, NFL is compared to ESPN FPI (including the FPI scalar and W-L-T from the same payload). FBS is compared to the AP Top 25. For the current season we use ESPN’s live rankings payload (including the preseason poll); otherwise we walk regular-season weeks in reverse, then fall back to preseason. Delta in the UI is espnRank − modelRank: positive means we have you higher than the public list.
09 · Bottom line
What you are looking at
Ingest completed regular-season ESPN box scores. Turn each sideline into a 25-D vector of yards, efficiency makes, return yards, turnovers, flags, and clock — everything except the score. Z-score the season. Train a 25→128→64→32→1 ReLU MLP with Adam/MSE to predict standardized point differential from that vector. Rank NFL by the mean of those predictions; rank FBS with an iterative Simple Rating System on top, then keep the Top 25. If the season is still being played, shrink every team toward last season's rating in proportion to how little it has played. Compare against FPI or the AP poll so you can argue on the internet with numbers that at least came from a well-specified objective.
If you want the code, the interesting files are src/utils/espnApi.ts (ingest + parse), src/lib/powerRankings.ts (z-score, graph, fit, mean-pool), and src/app/api/rankings/route.ts (cache + HTTP).