Elo-plus-market NFL model, mirrored from the public repository
greerreNFL/nfelo @ 0d3f8418 · model v4.3.0 · snapshot 2026-08-06.
No reusable software license was verified on 2026-08-07, so Data Dawgs consumes output and does not copy its code.
Ratings and projections are the author's. The backtest below is ours.
nfelo takes 538's Elo framework, adapts it for the NFL, adds a QB adjustment, and then — the part that matters — regresses its own answer toward the betting market. It is the best-documented public model that formally blends a power rating with the Vegas line, and its author publishes the whole thing, weights included.
We cloned it rather than building our own rating, for the reason the research kept landing on: a home-built EPA rating is one more correlated echo of a signal the market has already priced. Borrowing a polished, backtested one costs nothing and starts from a higher floor.
nfelo's projections joined to nflverse final scores. — games, —, ties dropped. Benchmark: the closing moneyline, devigged — the hardest number in sports to beat, and the one the model is explicitly trying to beat.
The two models only disagree on — games out of — — — of the sample. On those, nfelo is right — of the time. That is a coin flip, on the only games where the model has anything to say.
Where it does edge ahead is Brier score — probability quality rather than pick quality. — vs — for the market, a — relative improvement. Real, consistent in direction, and very small. If you want a number, take nfelo's. If you want a pick, the close is just as good and free.
Favorite's projected win probability vs how often that favorite actually won. On the diagonal = perfectly calibrated. The shaded band is where survivor picks live.
Straight-up accuracy, nfelo minus market close, percentage points. Above the line = nfelo won the season.
Taking nfelo's side whenever its projected line differs from the close, graded on the closing number. Breakeven at −110 is 52.38%.
Elo scale, 1500 = average. Base is the team rating carried forward and regressed to a market-derived preseason prior; QB is the starter adjustment; Pts is the rating converted to points versus a neutral average opponent.
nfelo's own numbers, unmodified. Spread is from the home team's perspective; negative means home favored.
These are preseason numbers with no 2026 snaps behind them and no closing lines to regress toward yet. Treat the confidence column as a ranking, not a probability you'd bet at.
Didn't: the model. Ratings, spreads and probabilities are nfelo's output, copied. No re-weighting, no blending, no "improvements."
Did: the evaluation. The author reports his own accuracy. We joined his projections to nflverse final scores and scored them independently, published the disagreement count, and reported the benchmark comparison that goes against him alongside the one that goes for him.
Snapshot, not a live feed — ratings are frozen at capture and will not move with injuries or transactions until we re-pull. The 2009–2025 backtest overlaps the model's own optimization window, so ATS figures are optimistic by an unknown amount. Ties are dropped from straight-up accuracy (17 games). Devigging is multiplicative, which understates favorite-longshot bias slightly in exactly the heavy-favorite tail survivor picks live in. And the model's edge, such as it is, is measured against the close — survivor picks lock before the close exists, which is a different and softer problem than the one benchmarked here.