Arena · grading the people who grade the players

The Dog Track. Pup

Every Thursday before kickoff, each ranking service hands in its list. After Monday night we check how close its order was to how players actually finished — and publish the score, not the ranks. The method below was written down before the first snapshot, and nothing here claims a winner the math does not support.

Report Card

PLAIN ENGLISH
The season grade, one card per service. The letter blends three things: whole-list order (Spearman), getting the top guys right (weighted τ — misses at the top cost more), and points captured (the money view). Hygiene counts how often a service left a ruled-OUT player ranked high on Thursday — sloppy, so we track it separately. Cards with the flashing lights are tied for the lead: a photo finish, not a champion.
Grade = blended percentile of the three metrics vs. the field, shrunk. Provisional until the declared season gate. Every input snapshot is stamped by the server before Thursday kickoff.

The Race

PLAIN ENGLISH
Same scores, run as a dog race. Each dog's position on the track is its season accuracy so far — further right is better. The checkered post is a perfect score, which nobody reaches. Dogs bunched together are in a photo finish: the season hasn't separated them yet. Hit Re-run to watch them come out of the gate again.
Dog position = season-to-date shrunk mean ρ mapped onto the track (0.30 at the gate, 1.00 at the post). Bunching = overlapping confidence intervals.

The Board

PLAIN ENGLISH
Every Thursday before kickoff, each service hands in its rankings. After Monday night, we check how close each one's order was to how players actually finished. A score of 1.00 means the order was perfect; 0 means it was no better than shuffling names in a hat. The big number is each service's season average so far. If two services are inside each other's error bars, the board calls it a PHOTO FINISH — too close to name a winner yet.
Score = mean weekly Spearman ρ on the shared player pool, shrunk toward the field. Range under the number = bootstrap 95% CI. Everything is provisional until the declared season gate. Method → the drawer below.

Wild Weeks

PLAIN ENGLISH
One chip per week, per service. A chip far right = that service nailed that week's order; far left = a rough week. Notice every service has chips scattered all over — week-to-week luck is huge, and the scatter inside one service is bigger than the gap between services. That's why we won't crown a winner on this much data. The bar is each service's typical (median) week.
Each chip = one week's Spearman ρ. Bar = median week.

The Cage

PLAIN ENGLISH
Forget correlations — this is the money view. Imagine the perfect lineup: the players who actually scored the most that week. Those points are the whole pot. Each service's chip stack shows how much of the pot its top-ranked players grabbed. 100% would mean their top group WAS the perfect group. A couple of points of difference ≈ one busted start a week.
Capture rate = points scored by each service's top-ranked group ÷ points scored by the actual top group, season to date. Differences under ~1.5 pts are noise at this sample.
Methodology — pre-registered before Week 1

The claim being tested. Whether a ranking service's Thursday order predicts the order players actually finish in, measured the same way every week, on rules fixed before the first snapshot.

Amendment rule. This section published before Week 1. Any change after Week 1 carries a dated note here saying what changed and why. If you are reading a number and cannot find the rule that produced it, that is a bug — say so.

What gets graded

An open entrants registry, not a fixed list. Each entrant is {id, name, type, first_week, color}. The Blend is the per-player mean rank across every service registered before Week 1's first kickoff; its membership freezes at Week 1 and is named below. Entrants registered from Week 1 onward never join it — a benchmark that moves corrupts every comparison made against it.

Scoring

Full PPR points for that NFL week. Actual finish = within-position rank by PPR points, mid-ranks for ties. The pool per position per week is the union of the consensus top-N across uploaded services and the actual top-N by PPR, at depths RB 36, WR 48, QB 24, TE 24. A player a service left unranked is slotted at that service's deepest ranked player + 1, ordered by consensus — never dropped. Players who did not play (inactive, bye, ruled out) leave the correlation pool entirely.

The three metrics — all three, no fourth

  1. Spearman ρ (headline) — the service's order against the actual order, across the whole pool.
  2. Weighted Kendall τ — the same idea with hyperbolic weights on the actual finish (w(r)=1/(r+1)), so missing the RB1 costs far more than missing the RB30.
  3. Capture rate — points scored by the service's top-G group divided by points scored by the actual top-G group. G is RB 12, WR 12, QB 6, TE 6.

Hygiene counts players who were officially OUT at capture time yet sat inside the startable range (top 24 RB/WR, top 12 QB/TE). It never touches the correlation.

Season aggregation and uncertainty

Season value = mean of the weekly values per position. The ALL scope is an equal-weight mean across the four positions — a declared choice, stated here because it is a choice. No dropped weeks, no mulligans. Confidence intervals are a bootstrap over weeks (resampled with replacement, 2,000 draws, percentile method) and are seeded, so rebuilding the page does not move a number that nothing changed. Scores are shrunk toward the field: shrunk = field + 0.7 × (raw − field) until Week 10, and an empirical-Bayes weight after. The headline number is the shrunk one; the raw value and its interval sit underneath.

Ties, and what we refuse to claim

Any service whose interval upper bound reaches the leader's lower bound is tied with the leader and renders as a PHOTO FINISH. Head-to-head calls use only the weeks both entrants were graded; under four shared weeks we say "insufficient overlap" rather than pick. An entrant with fewer than four graded weeks is provisional regardless of its score. Single-week views show raw metrics only — no interval, no shrinkage — because one week is one observation.

Promotion gate, declared now

This page is a Pup. It becomes a Working Dawg at season's end only if at least one pair of services separates with non-overlapping shrunk intervals on the ALL scope. If no pair separates, this page will say so plainly and stay a Pup. "We could not tell them apart" is a result, and it is the one this kind of exercise usually produces.

Provenance

Every snapshot is stamped captured_at by the server, before that week's first kickoff, and is immutable once written. A bad upload is corrected by a logged void plus a fresh snapshot before kickoff — never by replacement. Raw third-party ranks are paid content: they never leave the server, and nothing player-level appears in this page or its machine surface. Only derived scores publish.