Receipts Pup

Every call gets a date, a number and a benchmark — written down before the games, so neither the result nor the story can be edited afterwards. Locked 2026-08-06 · calls ·

Every receipt this site has registered

This is a count, not a grade. It says what has been written down and locked, how much of it has been settled, and where each ledger lives so you can read it yourself.

There is deliberately no single score here. An NFL win probability, a college-football forecast and a verdict about whether one of our own tools deserves a collar are different kinds of claim with different base rates. A Brier score blended across them would look like a grade and mean nothing, so grading stays inside the per-domain sheets, next to the assumptions it depends on.

forecasts registered
games covered
settled
awaiting a result
Reading the ledgers…

The tier audit is counted separately, on purpose

Loading the audit…

Nothing above is graded yet, and that is the correct state: the season has not started. A receipt written now is worth something precisely because it cannot be edited later. Every number on this sheet is read from the published payloads at page load, so none of it can drift from the files.

Grading

Finals come straight from ESPN's public scoreboard, fetched by this page — not through the Worker, whose proxy route is dead because ESPN blocks Cloudflare egress. Grading is idempotent: a game already scored is never re-scored, a game not yet final is skipped rather than guessed, and a tie stores no straight-up result.

Calibration

Predicted probability against how often those calls actually landed. Fills in as the season resolves.

Week by week

Survivor receipts — what the board said, before kickoff

Every week, the survivor board's number-one recommendation is written down before the week's earliest kickoff and graded afterwards. Nobody else in this space grades survivor recommendations prospectively, which is the only reason this ledger is worth keeping.

receipts registered

Loading the survivor receipt ledger…

What is graded here is calibration, not victory. The score is whether the pick survived and the Brier on the win probability the board stated at the time. Whether an entry would actually have won a pool depends on the rest of the field — that is not ours to score and this ledger does not claim it. And n stays small. A season is at most eighteen scored predictions per entry, so every figure is printed next to its n and there is deliberately no calibration curve: eighteen points cannot support one.

Machine read: /data/survivor-receipts.json. Captured by work/survivor-receipt.mjs, which refuses to write once a week has kicked off — a late receipt is not a receipt.

DFS week objects — locked expectations, hashed lineups

Each row is a week locked before kickoff with lockedSha256 over canonical JSON and realized:null until Monday standings ingest. Lineups are salt-less SHA-256 hashes of sorted DK ids + CPT — never player lists, projections, or ownership (I3). PoC C1–C5 already hashed; this enables the week-object path.

weeks locked

Loading the DFS week-object ledger…

Machine read: /data/dfs-weeks.json. Built on-device by work/dfs-week.js; Kap publishes locked weeks later.

The locked call sheet

All 272 regular-season games, fixed on 2026-08-06. Call is nfelo's win probability for the home team. Benchmark is the devigged closing market price where a line existed at lock — blank where the market had not priced the game yet.

Model scoreboard — what the models believe

One row per registered model line, then the same beliefs game by game, normalized to home win probability. This sheet is what the models believe; the Scoreboard sheet is how our pre-registered calls grade. Keeping those apart is the point of this page. nfelo locked 2026-08-06 · 538 Classic locked 2026-08-08 · DDPR pre-registered 2026-08-10.

games
registered lines
independent lines
largest range

The board

Loading the registered model lines…

How much of this is independent evidence?

Correlation of log-odds between every pair of registered lines, measured on the published prospective forecasts. It needs no resolved games, which is why it is the one honest comparison available before kickoff. A line derived from two others cannot disagree with them by much, and this table is where that shows up rather than being asserted in prose.

Week 1, game by game

Loading the public dated receipts…

Important: registered lines are not independent confirmation. nfelo uses market-informed machinery; 538 Classic deliberately omits QB, injury, weather and market adjustments; DDPR is a pre-registered average of those two and carries no evidence they do not already carry, which is why the mean, range and standard deviation above are computed over the independent lines alone. SRS, Units, WEPA and QB-adjusted columns remain absent until lawful normalized forecasts exist. Machine reads: /data/model-board.json · /data/model-receipts.json · /data/ddpr-nfl.json · dd_model_scoreboard over MCP.

Provenance before claims

A public repository is not automatically reusable software. Integration mode follows the verified license. Captured 2026-08-07 · primary repositories inspected · moved here from The DawgHouse 2026-08-09. Machine read: /data/upstream-models.json.

Loading upstream provenance…

538 Classic — what a deliberately simple benchmark believes

This sheet is a belief, on the same side of the line as Models. It is not a grade: none of these forecasts has been scored, because the 2026 season has not started. 538 Classic is the original FiveThirtyEight Elo mathematics with no QB, injury, weather or market adjustment — a permanent simple benchmark to beat, not a claim that it is any good. Rendered from its own file when you open this sheet.

forecasts
teams rated
weeks covered
0graded
Loading 538 Classic beliefs…

 

    What the reproduction check does and does not say. It says our implementation reproduces the published 538 mathematics. It says nothing about whether those mathematics predict 2026 well; the file sets independent_skill_claim: false for exactly that reason. Machine read: /data/538-classic.json.

    CFB receipts — the contract, and nothing in it yet

    A belief ledger, on the Models side of the line: it is where college-football forecasts get frozen before kickoff so they can be graded later. It is empty. That is the honest state, not a loading failure — no CFB forecast has been frozen before a kickoff yet, and back-filling the completed 2025 season would be inventing a receipt nobody wrote.

    rows in the ledger

    Loading the CFB receipt contract…

      The CFB work that does have published numbers lives on the CFB Lab; none of it is a pre-registered receipt, which is the distinction this sheet exists to hold. Machine read: /data/cfb-model-receipts.json.

      Tier audit — our own tools, against our own gate

      This sheet sits on neither side of the line. Models is what models believe; Scoreboard is how our pre-registered calls grade. This is neither: it is a judgment about whether each of our tools deserves the tier it claims, reasoned from the code and the published data by one reviewer in one day. It grades tools against a governance gate, not predictions against outcomes. It is published rather than held precisely because that single-reviewer weakness is the thing most worth disclosing.

       

      audited
      promotions
      demotions
      in the pound

       

      Loading the tier audit…

      Machine read: /data/tier-audit.json · the gate itself is on the homepage.

      Pre-registration — written 2026-08-06, before any 2026 game

      Predictions are nfelo's, snapshot 0d3f8418, model v4.3.0. For the 16 games nfelo published directly, its own number is used; for the rest, the probability is derived from its power ratings through a margin model fitted on 4,108 historical games (23.58 Elo per point, 2.10 home-field, residual SD 13.18).

      The benchmark is the devigged closing moneyline, taken at lock. It exists for of 272 games; the rest had no line on 2026-08-06 and therefore have no benchmark and never will. Those games test calibration only.

      What is being tested

      Primary, and resolvable this season: are nfelo's stated probabilities calibrated — do the games it calls at 70% land about 70% of the time? 272 games is enough to catch a real miscalibration.

      Secondary, and NOT resolvable this season: does nfelo beat the closing line?

      One season cannot settle the comparison, and we are saying so now rather than in January. Measured on 4,096 historical games, the paired per-game spread between nfelo and the closing line is so wide relative to the edge that a 272-game season gives:
      • a minimum detectable straight-up difference of ±1.60 pp — nfelo's advertised +0.14 pp is 9% of that
      • a minimum detectable Brier gap of 0.0034 — the observed edge is 19% of that
      • only about five games a season where the two even disagree

      At the observed effect sizes this takes roughly 26 seasons on Brier and 130 straight-up. Any 2026 result on this axis is noise, whichever way it falls. It is logged because season one of twenty-six still has to be season one — not because it will mean anything next winter.

      What would change our mind

      On calibration — the resolvable question — a miss of more than about 8 points in any bin holding 30+ resolved games, in the same direction across the 70–90% range, is a real failure and not sampling noise. That is the finding this page is actually built to catch. A single wrong-looking week is not evidence of anything.

      On the benchmark comparison, nothing that happens in 2026 changes our mind, by construction. Stated in advance so it cannot be reinterpreted later.

      Integrity

      The call sheet is hashed. is the SHA-256 over every game_id | probability | benchmark row in the locked order. If a prediction is ever quietly edited, this hash changes and the receipt is void. Verify it against the page's own data at any time:

      Scoring

      Brier score (lower is better) as the headline, because it grades the number rather than the pick. Straight-up accuracy alongside it, paired against the benchmark on the identical game set. Calibration by bin. Every figure recomputed from resolved games only — nothing is projected, nothing is back-filled.

      Tier audit — 2026-08-07

      A receipt is not only a forecast. This is the same question the site asks of a prediction, asked of our own tools: is each one in the right lane? The gate is does it work, and can it be trusted — forecasts need receipts against a benchmark chosen in advance, measurement tools need named sources and reproducible math. Full rationale, tool by tool, in the write-up · machine-readable.

      Result: no promotions, no demotions. That is a suspiciously comfortable outcome for a self-audit and is worth saying so plainly. It survives only because every collar came back with a condition attached, and because the tool closest to promotion is blocked for a specific stated reason rather than a vague one. This is one reviewer, on one day, reasoning from the code and the published data. It is a judgment, not a measurement. It is published so the next person can find where it was wrong.

      The three Dawgs keep their collars, with conditions

      NFL EPA Stats. Sources named, snapshot dated, and the aggregation was independently re-implemented outside the page and reproduces. What was not done: any external cross-check against a published nflfastR table. Re-implementing a page’s own algorithm proves the maths is reproducible, not that it is right — two implementations of the same wrong filter agree perfectly. One spot-check is owed.

      nfelo Power Ratings. The collar is for the mirror, not the forecasts. It certifies that this site faithfully renders nfelo’s published ratings. It does not certify that nfelo predicts well, and the evidence on file says it does not beat the market: nfelo 0.6674 straight-up against the market’s 0.6677 over 4,053 games — behind by 0.02 points against a 0.196-point standard error, and in-sample-ish besides.

      Fantasy Draft Dashboard. It ran a live 14-team draft end to end, which is the right evidence for an operational tool. Three conditions: the Dashboard is a frame over board, dataviz, report and auction, and those are reachable standalone with no tier chip — same code, different verdict depending on how you arrive. The Grades view is a measurement dressed as a judgment: it scores buying value relative to the room, not whether a roster will win. And the whole rig runs on a market snapshot dated 2026-07-29, from a source that publishes through August. The collar is conditional on refreshing it before draft night.

      Receipts is the closest thing to a Dawg, and stays in Labs for one reason

      Everything else clears: sources named, math reproducible, dated, failure threshold pre-registered before a single 2026 game. The blocker is that the trust mechanism is self-anchored. This site hosts the ledger, the hash, and the spec that connects them. Recomputing proves today’s file matches today’s hash — it cannot prove these were the predictions made before the games, because both could have been replaced together. A page whose entire purpose is verifiability cannot earn a collar on verification that is not independently checkable. Anchor the hash externally and timestamp it, and this clears. Nothing else about the page needs to change.

      The empty Pound is the biggest finding

      The Pound exists to keep failures visible next to the wins. It has no residents. Either nothing here has ever failed, or things are being abandoned rather than recorded — and the first is not plausible. An empty Pound is survivorship bias appearing in the one structure built to prevent it. There are real candidates already: a Worker route that died when ESPN blocked Cloudflare egress, two rejected definitions of this very gate, blind Bozo submission, the composite-score tiebreaker. Populate it, or say on the page that nothing has been retired yet and why that is true. Silence reads as “we don’t fail.”

      What would change these verdicts

      EPA Stats — an external cross-check that disagrees with the page. nfelo — nothing this season; that comparison is unresolvable in one year by this site’s own pre-registered arithmetic. The draft rig — a draft night where it fails under load, or a refusal to refresh the snapshot before it. Receipts — an external, timestamped anchor would promote it, and nothing else would. Any Labs tool — receipts against a benchmark declared in advance.

      Material changes

      Changes to how this site names or frames its claims, recorded here rather than applied silently. The forecast rows are covered by the hash above; this list covers what the hash cannot see.

      2026-08-09 — The Pound is renamed The DawgHouse. Human-facing labels only: the shelf tier reads “The DawgHouse” wherever a person sees it, and /pound.html forwards to /dawghouse.html. Stored machine values did not move — tier stays "pound" and /data/pound-tools.json keeps its name — so nothing downstream re-keys. No forecast, probability or benchmark changed. The same day, the shelf was reduced to genuinely blocked work: the complete NFL tools left it, and the full inventory stays machine-readable in /data/pound-tools.json.