System notes · TensorFlow.js · ESPN box scores

How the ranking model actually works

This is not a black-box “AI ranking.” It is a small multilayer perceptron, trained from scratch in TensorFlow.js on the server, whose job is to map a 25-dimensional box-score vector onto a standardized point differential, then average those predictions by team. The real work happens in computeSeasonRankings, inside a Node.js route that can run for up to 300 seconds.

  1. 01Roster the league32 NFL clubs, or FBS via ESPN group 80
  2. 02Enumerate weeksRegular season only · seasontype=2
  3. 03Pull box scoresSummary API · 10 concurrent workers
  4. 04Featurize25-D vector per team-game
  5. 05Z-scorePopulation μ, σ across the season
  6. 06Fit MLP25→128→64→32→1 · Adam · MSE
  7. 07RankNFL mean-pool · FBS SRS + Top 25
  8. 08Shrink to priorBlend last season while this one is unfinished

01 · Ingest

Where the bytes come from

Every number in the model originates from ESPN’s undocumented but public JSON APIs. There is no proprietary play-by-play dump, no Next Gen Stats, no PFF grades. We speak HTTP to three families of endpoints, with a 3-attempt retry and linear backoff (400 ms, 800 ms), and we ask Next to cache each response for 24 hours.

PurposeNFLFBS (CFB)
Team universesite.api …/nfl/teams?limit=32core …/groups/80/teams (FBS)
Week calendar + scoreboard…/nfl/scoreboard?seasontype=2…/college-football/scoreboard?groups=80
Per-game box score…/nfl/summary?event=ID…/college-football/summary?event=ID
Reference rankingESPN FPI (fpi, fpirank, W-L-T)AP Top 25 (live site poll, else latest week / preseason)

NFL teams are the 32 clubs on the league teams endpoint, sorted by display name. College football is not “every NCAA team”: the roster is ESPN group 80 (FBS), loaded from the core API season group-teams collection. Scoreboards are pinned to the same group. The public /teams?groups=80 site list ignores the filter, so we do not use it. Division III clubs never enter the ranking universe. If an FBS club plays an opponent that is not in group 80, we still ingest the FBS side of that box score; the missing opponent is later treated as a cupcake in the SRS pass.

Season slicing is strict. We only keep regular season games (seasontype = 2). Preseason and playoffs never enter the tensor. A game is admitted only if ESPN marks it completed (status.type.completed or STATUS_FINAL). Week numbers are read from the scoreboard calendar; if that parse fails we fall back to 18 NFL weeks or 16 CFB weeks.

Event ids are collected across every week, dumped into a Set so a game cannot be trained on twice, then fetched with a worker pool of 10 concurrent summary requests. That pool is why the first CFB run is slow: you are hydrating on the order of a thousand game summaries, not running a giant GPU job.

02 · Parse

One game becomes two training rows

A summary payload has a two-team box score plus a header with final scores. We emit two GameStats objects: team A vs B, and team B vs A, with opponent features mirrored. That duplication is load-bearing. It makes the marginal distribution of score and opponentScore identical across the season matrix (every point scored is also a point allowed, from the other row). That identity is what lets the target collapse to a scaled margin, as we'll see in §04.

ESPN stats are annoyingly heterogeneous. Some fields are numeric values. Efficiency stats arrive as display strings like 7-14 or 7/14. We split on [-/] and keep the numerator (makes). Completion percentage is reconstructed from completions and attempts rather than trusted as a pre-baked percentage. Penalties arrive as a single count-yards blob. Possession is either MM:SS or a raw second count. Missing stats become 0, not null — the network never sees NaNs.

Special teams are the flakiest layer. Kick-return yards, punt-return yards, field goals made, and punts are read from the team box score first. If those keys are empty, we walk boxscore.players, find the matching statistical group (kickReturns, puntReturns, kicking, punting), locate the column by key name, and sum across athletes. Slash stats like fieldGoalsMade/fieldGoalAttempts contribute only the left-hand number.

03 · Feature space

The 25-dimensional team-game vector

After parsing, time of possession is converted to seconds and the record is projected into a fixed-order float32 vector of length 25. There is no embedding layer, no opponent id in the tensor, no home/away bit, no week index, no rest days, no QB identity. The MLP can only “see” what happened on the stat sheet that afternoon. Opponent identity is stored alongside the row and used only after predict, in the FBS Simple Rating System pass.

Crucially, the two scoring columns are not in that vector. They are parsed and normalized alongside everything else, but they exist only to build the label. Feeding them in would hand the network the answer, and the MLP would collapse into an identity function on two slots.

Offense (own)

Tensor slotESPN sourceNotes
offensiveYardstotalYardsNet offensive yards from scrimmage.
passingYardsnetPassingYardsNet passing yards (sacks already deducted by ESPN).
rushingYardsrushingYardsRushing yards.
completionPercentagecompletionAttemptsDerived: completed / attempts × 100, parsed from C-A or C/A.
firstDownsfirstDownsFirst downs earned.
thirdDownConversionsthirdDownEffMakes only, parsed from made-attempt strings.
fourthDownConversionsfourthDownEffMakes only.
redZoneConversionsredZoneAttemptsScores inside the 20, makes only.

Defense (opponent box score)

Tensor slotESPN sourceNotes
opponentOffensiveYardsopp.totalYardsYards allowed. Mirror of the other sideline.
opponentPassingYardsopp.netPassingYardsNet passing yards allowed.
opponentRushingYardsopp.rushingYardsRushing yards allowed.
opponentCompletionPercentageopp.completionAttemptsOpponent C% allowed through the air.
opponentFirstDownsopp.firstDownsFirst downs allowed.
opponentThirdDownConversionsopp.thirdDownEffOpponent third-down makes.
opponentFourthDownConversionsopp.fourthDownEffOpponent fourth-down makes.
opponentRedZoneConversionsopp.redZoneAttemptsOpponent red-zone scores.

Special teams / kicking

Tensor slotESPN sourceNotes
kickReturnYardskickReturnYardsTeam total, else Σ player kickReturns.kickReturnYards.
puntReturnYardspuntReturnYardsTeam total, else Σ player puntReturns.puntReturnYards.
fieldGoalsMadefieldGoalsMadeMakes parsed from FG made/attempts; player kicking fallback.
puntspuntsPunt count; player punting fallback.

Turnovers, flags, clock

Tensor slotESPN sourceNotes
turnoversturnoversGiveaways. ESPN’s team turnover total.
takeawaysopp.turnoversNot a native takeaway field — opponent giveaways.
penaltiestotalPenaltiesYardsCount half of the P-Y display string.
penaltyYardstotalPenaltiesYardsYardage half of the same string.
timeOfPossessionpossessionTimeMM:SS → seconds. Numeric ESPN values treated as seconds.

Parsed but withheld from the tensor

ColumnESPN sourceNotes
scoreheader.competitors.scoreFinal points.
opponentScoreopp scoreFinal points allowed.

04 · Normalize

Z-scoring the season, then defining y

All team-games in the season are flattened into one matrix of 27 parsed columns — the 25 that become the input tensor plus the two scoring columns. For each column j we compute the population mean and standard deviation (divide by n, not n − 1):

μⱼ = (1/n) Σᵢ xᵢⱼ
σⱼ = √( (1/n) Σᵢ (xᵢⱼ − μⱼ)² )
x̃ᵢⱼ = (xᵢⱼ − μⱼ) / σⱼ     with σⱼ := 1 if the column is constant

This is a season-relative z-score: a 350-yard passing game in a shootout year is not the same coordinate as 350 yards in a mud year. Normalization is global across the league-season, not per-team, so a team that always runs the ball still gets compared to the league centroid.

The supervision target is not win/loss and not a ranking permutation. It is the difference of the already-normalized scoring columns:

yᵢ = x̃ᵢ,score − x̃ᵢ,opponentScore

Because every game is stored twice (once per sideline), score and opponentScore share the same empirical moments: μs = μo and σs = σo. The expression therefore simplifies to yᵢ = (marginᵢ) / σscore. We are asking the net to predict standardized point differential. That is the entire learning problem.

Those two columns are then dropped, and only the remaining 25 are stacked into X ∈ ℝⁿˣ²⁵. The net has to infer the margin from how the game was played.

05 · Architecture

A four-layer MLP, ~13.7k parameters

The network is a vanilla TensorFlow.js tf.sequential dense stack. No dropout, no batch-norm, no residual connections, no attention. Hidden activations are ReLU. The output unit is linear (identity) because this is unbounded regression, not a classification.

input ℝ²⁵

dense 25 → 128 ReLU 3,328 params

dense 128 → 64 ReLU 8,256 params

dense 64 → 32 ReLU 2,080 params

dense 32 → 1 linear 33 params

≈ 13,697 trainable weights

  • Optimizer: Adam, TensorFlow.js defaults (α = 10⁻³, β₁ = 0.9, β₂ = 0.999, ε = 10⁻⁷).
  • Loss: mean squared error on y. Quadratic penalty, so a 21-point miss is 9× as expensive as a 7-point miss. Blowouts dominate the gradient.
  • Reported metric: MAE, which is easier to read in “standardized points” but is not what Adam is minimizing.
  • Schedule: 50 epochs, batch size 32, validationSplit: 0.2, shuffle: true (TF.js default). The last 20% of the shuffled rows are held out for the val curve; they are not held out of the eventual ranking pass.

Sample size vs capacity is the fun part. An NFL regular season is 272 games × 2 rows ≈ 544 examples. After the val split you have ~435 training rows to estimate 13.7k weights. That is an overparameterized interpolating regime. FBS is healthier (~146 teams × ~12 games, still doubled). We are not training a foundation model. We are fitting a flexible function of one season’s box scores, then immediately evaluating it on those same rows.

Weights are ephemeral. After fit we predict, convert the output tensor to a nested JS array, then dispose the input, target, and model and close a tf.engine() scope. Nothing is checkpointed to disk. Hit the endpoint again (or wait out the 24h cache) and you train a fresh draw of Adam from the default random initialization. Rankings can jitter slightly run to run; the cache is what makes the UI feel deterministic.

06 · Inference

Mean predicted margin, then strength of schedule

After training we push the entire season tensor — train and val rows together — back through the net. For team t with games Gₜ the raw strength is mean predicted margin:

sₜ = (1 / |Gₜ|)  Σ_{g ∈ Gₜ}  fθ(x̃_g)

NFL stops there. FBS runs a Simple Rating System on top of those per-game margins so a 40-point win over a cupcake is not the same as a 40-point win over a contender. Ratings start at sₜ, then iterate ~20 times:

rₜ ← (1 / |Gₜ|)  Σ_{g ∈ Gₜ}  ( fθ(x̃_g) + r_opponent(g) )
r ← r − mean(r)

If the opponent is not on the FBS roster (FCS / unknown), r_opponent is the 10th percentile of current ratings — a cupcake floor. Mean-centering each round identifies the system. Teams with zero completed regular-season games contribute nothing to fit; they are carried by the prior described in §07. FBS is then sliced to the Top 25 for the API, CLI, and UI.

Two interpretive caveats, because they matter:

  • In-sample, not out-of-sample. The scoring columns are withheld from the input, so the net genuinely has to predict margin from yards, efficiency, turnovers, and clock. But it is evaluated on the same rows it trained on, including the validation split. This is a fitted summary of a season that has already happened, not a forecast of games it has never seen.
  • SOS is graph-based, not poll-based. Opponent quality is the opponent’s own evolving rating, not AP votes or FPI. Home field, injuries, weather, and recency are still invisible. A week-1 demolition and a week-18 demolition are still exchangeable except insofar as the opponent’s rating differs.

So when our list disagrees with FPI or the AP poll, that is not always a bug: FPI is a predictive efficiency rating; AP is a human ballot. We are mean-pooling a box-score MLP’s reconstructed margin, then (for FBS) adjusting that margin for who it came against.

07 · Prior

What a season looks like before it has been played

A rating built from one September Saturday is not a rating. In week 1, most of the league has played zero games and would vanish from the list entirely, while the handful of teams that opened in week 0 would take every top slot on the strength of a single result. That failure is loud in FBS, where a 148-team roster can collapse to the sixteen teams that happened to kick off early.

So an unfinished season is not ranked on its own evidence. We detect it directly from the scoreboard: if any regular-season game is still scheduled or in progress, the season is live. Canceled and postponed games sit in ESPN's post state without completing, so they do not keep a finished season looking open forever.

When the season is live we run the whole pipeline a second time on the previous season and use it as a preseason prior. The two rating vectors come from independently initialized nets, so neither scale is meaningful on its own; each is standardized to zero mean and unit variance before they are combined. Every team is then shrunk toward its prior by how much it has actually shown you:

w = n / (n + 4)
rating = w · thisSeason + (1 − w) · lastSeason

With n the games played so far, one game buys a team 20% weight on its own season, four games buy 50%, and twelve games buy 75%. A team that has not kicked off sits entirely on last year's rating instead of disappearing. A program with no prior at all — new to the division — starts at the league average, which is generous but never enough to reach a Top 25 built from real ratings.

A completed season skips this path entirely: no second pipeline run, no prior, no extra latency. The response carries priorSeason so you can always tell which of the two you are looking at.

08 · Serving

The HTTP path, the cache, and the fake HUD

The browser calls GET /api/rankings?season=YYYY&league=nfl|cfb. Seasons are clamped to 2000–2026. The handler is a Node runtime (not Edge) with maxDuration = 300, wrapped in unstable_cache under the key power-rankings-v6, revalidate 86,400 seconds. First visitor of the day pays for ESPN hydration plus 50 epochs of Adam. Everyone else that day gets JSON out of cache.

The sci-fi overlay — “UPLINKING ESPN SCOREBOARD,” the ring, the equalizer — is a client-side progress skin keyed off the fetch promise. Its stage list is not hooked to real training callbacks. model.fit runs with verbose: 0. We do not stream epoch loss to the browser.

Side-by-side on the home page, NFL is compared to ESPN FPI (including the FPI scalar and W-L-T from the same payload). FBS is compared to the AP Top 25. For the current season we use ESPN’s live rankings payload (including the preseason poll); otherwise we walk regular-season weeks in reverse, then fall back to preseason. Delta in the UI is espnRank − modelRank: positive means we have you higher than the public list.

09 · Bottom line

What you are looking at

Ingest completed regular-season ESPN box scores. Turn each sideline into a 25-D vector of yards, efficiency makes, return yards, turnovers, flags, and clock — everything except the score. Z-score the season. Train a 25→128→64→32→1 ReLU MLP with Adam/MSE to predict standardized point differential from that vector. Rank NFL by the mean of those predictions; rank FBS with an iterative Simple Rating System on top, then keep the Top 25. If the season is still being played, shrink every team toward last season's rating in proportion to how little it has played. Compare against FPI or the AP poll so you can argue on the internet with numbers that at least came from a well-specified objective.

If you want the code, the interesting files are src/utils/espnApi.ts (ingest + parse), src/lib/powerRankings.ts (z-score, graph, fit, mean-pool), and src/app/api/rankings/route.ts (cache + HTTP).