DraftScout.ai

Ideas of ours that failed.

A tool that only publishes its wins is a tool you cannot calibrate. Here is what we tried, what the gate said, and what happened next.

Pick-edge v1, one-step regret

Tried: a candidate's value to your squad minus the survive- weighted value of the best player likely to be there at your next pick, discounted for positions your shape has already filled.

The gate: a 160-draft tournament against the strategy that actually ships (needs-aware best value). Scored 1435.4 mean best-XI points against 1505.9, behind in all 8 draft slots, winning 1 of the 160 paired comparisons.

What happened next: killed as the live strategy. Its own-position regret overweighted scarce positions against plain value over replacement. Retained in the code as a diagnostic baseline, not as the pick line.

Pick-edge v2, rollout edge

Tried: simulate taking each candidate now, then play the entire rest of the draft out deterministically, and score the resulting squad's best legal starting XI. Rate each candidate by the plan it produces, not by a single number.

The gate: the same tournament. Scored 1464.8 against 1505.9, behind in all 8 slots, 14 of 160 paired wins.

What happened next: killed. Its model of rival behaviour assumed rivals follow the official list order, and value-driven rivals instead took the high-value players the rollout expected to survive. Iteration stopped there, as pre-registered, rather than trying a third variant: a third try against our own simulator is curve-fitting, not evidence.

Two-season prior for the board's points per 90

Tried: blending two seasons of points-per-90 (w = 0.6, 0.7 or 0.8 toward the more recent season) instead of using one season alone, to stop an injury-wrecked season from burying a player the board barely remembers.

The gate: a pre-registered bootstrap on realised next-season points per 90, primary cohort of 242 players with at least 900 realised minutes. Every blend beat the single-season baseline on the point estimate at every weight (best case, w=0.6: +0.0091 pairwise accuracy), but no 95 percent confidence interval on the primary cohort excluded zero.

What happened next: killed at its own pre-registered bar. A smaller, more permissive cohort did clear zero, and promoting on that reading after the primary cohort missed is exactly the cherry-pick the pre-registration exists to forbid. The board's prior stays single-season. The direction still looks real; a future attempt needs more evidence, not a lower bar.

Finishing-luck blending in the weekly model: on probation, not a failure

Not everything here is a kill. One candidate in the same model programme, blending expected goals into the weekly projection for the personal, season-long build, met its own statistical bar for the first time: seven out of seven gameweek clusters agreed on the direction, and the transparent corrected significance is 0.1094. Met the bar is not the same as certified, and the flag ships on probation until it replicates on real 26/27 gameweeks. It is a different codepath from the board's own finishing-luck flag on the board, which stays display-only and never was a candidate for this gate.

Archetype calibration, explored then deferred

Tried: checking whether recommendation outcomes cluster by archetype (position, fixture difficulty, confidence tier, and more), so the model could correct itself per archetype rather than only overall.

What we found: no single dimension explained more than 4.6% of the variance in actual outcomes. A few specific buckets were individually significant, but the dimension-level signal was too weak to build on.

What happened next: deferred, not built. Applying multipliers to dimensions that do not differentiate outcomes would add noise, not accuracy. The simpler global correction already in the model does the same job with the evidence actually available.

A headline we corrected downward, twice

The queue optimiser's cumulative lift across one league's season was first reported at +16 points across 26 historical gameweeks. Our own audit found four bugs in the replay simulator that were inflating it (opponent claims scored as always-successful, a stale historical priority order, and two others), and the corrected figure became +13.

A double-gameweek data bug then came to light and was fixed separately; re-running the corrected replay on the corrected data brought the figure to +10. A later, independent frozen-database backtest found the +10 was itself inflated by phantom availability the simulator was giving optimised queues, and the cross-validated, currently authoritative figure is +7.

Every one of those numbers is kept in the project's own history rather than quietly replaced. The lesson taken was structural, not just numeric: a simulator that scores its own optimiser is exactly the kind of code that should be audited against real transactions before its number is trusted, and this one was, twice.

The method these gates test is on the how-it-works page.

How it works

Not affiliated with, endorsed by, or connected to the Premier League, Fantasy Premier League, or any club. Player data is derived from publicly accessible Fantasy Premier League endpoints and remains the property of the respective rights holders. Analytics, not betting.

This page last changed 2026-08-02.