nfelo Power Ratings

Elo-plus-market NFL model, mirrored from the public repository greerreNFL/nfelo @ 0d3f8418 · model v4.3.0 · snapshot 2026-08-06. No reusable software license was verified on 2026-08-07, so Data Dawgs consumes output and does not copy its code. Ratings and projections are the author's. The backtest below is ours.

What this is, and why it's here

nfelo takes 538's Elo framework, adapts it for the NFL, adds a QB adjustment, and then — the part that matters — regresses its own answer toward the betting market. It is the best-documented public model that formally blends a power rating with the Vegas line, and its author publishes the whole thing, weights included.

We cloned it rather than building our own rating, for the reason the research kept landing on: a home-built EPA rating is one more correlated echo of a signal the market has already priced. Borrowing a polished, backtested one costs nothing and starts from a higher floor.

It cleared the gate on the instrument path, not the forecast path. Named sources, reproducible maths, a dated snapshot — all yes. But the forecast path requires receipts against a benchmark set in advance, and against the closing line nfelo does not have them. We checked. The numbers are below, including the ones that don't flatter it.

The receipts — recomputed here

nfelo's projections joined to nflverse final scores. games, , ties dropped. Benchmark: the closing moneyline, devigged — the hardest number in sports to beat, and the one the model is explicitly trying to beat.

The headline claim does not replicate. nfelo's site reports beating the closing line by about +0.14 percentage points of straight-up accuracy. On the identical game set, using the devigged closing moneyline, we get −0.02 — a dead heat, well inside a standard error of ±0.20. Switch the benchmark to the closing spread instead of the moneyline and nfelo comes out +0.20 ahead. Both readings are noise. The honest statement is that nfelo and the closing line pick winners at the same rate, and which one looks better depends on a convention choice nobody published.

The two models only disagree on games out of of the sample. On those, nfelo is right of the time. That is a coin flip, on the only games where the model has anything to say.

Where it does edge ahead is Brier score — probability quality rather than pick quality. vs for the market, a relative improvement. Real, consistent in direction, and very small. If you want a number, take nfelo's. If you want a pick, the close is just as good and free.

Calibration — and the survivor band

Favorite's projected win probability vs how often that favorite actually won. On the diagonal = perfectly calibrated. The shaded band is where survivor picks live.

nfelo closing market (benchmark) —— perfect calibration
For a survivor pool, this chart is the whole answer. In the 75–90% band both are calibrated to within about a point, on a few hundred games each. A survivor pick sourced from nfelo and a survivor pick sourced from the closing line are the same pick at the same price. The edge in a survivor pool is not here — it is in pick popularity and future value, which neither of these numbers knows anything about.

Season by season vs the closing line

Straight-up accuracy, nfelo minus market close, percentage points. Above the line = nfelo won the season.

Against the spread — where the signal actually is

Taking nfelo's side whenever its projected line differs from the close, graded on the closing number. Breakeven at −110 is 52.38%.

Discount this hard. nfelo's parameters were optimized on this same 2009–2025 window, so the large-disagreement buckets are close to in-sample. The pattern — edge concentrated in the games where the model disagrees most — is what you'd expect from a real edge and also what you'd expect from an overfit one. The only test that separates them is forward performance, which starts in September and which we will log here.

Power ratings — 2026 Week 1

Elo scale, 1500 = average. Base is the team rating carried forward and regressed to a market-derived preseason prior; QB is the starter adjustment; Pts is the rating converted to points versus a neutral average opponent.

Week 1 projections

nfelo's own numbers, unmodified. Spread is from the home team's perspective; negative means home favored.

These are preseason numbers with no 2026 snaps behind them and no closing lines to regress toward yet. Treat the confidence column as a ranking, not a probability you'd bet at.

What we changed, and what we didn't

Didn't: the model. Ratings, spreads and probabilities are nfelo's output, copied. No re-weighting, no blending, no "improvements."

Did: the evaluation. The author reports his own accuracy. We joined his projections to nflverse final scores and scored them independently, published the disagreement count, and reported the benchmark comparison that goes against him alongside the one that goes for him.

Known limitations

Snapshot, not a live feed — ratings are frozen at capture and will not move with injuries or transactions until we re-pull. The 2009–2025 backtest overlaps the model's own optimization window, so ATS figures are optimistic by an unknown amount. Ties are dropped from straight-up accuracy (17 games). Devigging is multiplicative, which understates favorite-longshot bias slightly in exactly the heavy-favorite tail survivor picks live in. And the model's edge, such as it is, is measured against the close — survivor picks lock before the close exists, which is a different and softer problem than the one benchmarked here.