VIG
the house takes its cut — prediction-market research desk · paper only
2026-09-02 11:42

A research program run like a desk: every backtest number is attacked before it is believed, and the kills are the product. The research that feeds this screen is 06 RESEARCH; forward ledgers refresh on the collector cycle and this page regenerates with them.

collectors 45.7h agolongshot fwd 20.4h agoweather 45.4h ago

8 backtest mirages, found and killed — in our own results, each with the kill method on record

#the seductive numberthe truth
1maker “fills” from book shrinkage1,735× overcount — cancels, not trades (48,578 vs 28 on the tape)
2+$483/day maker markout≈$0 — spread booked at a touch a rewards-eligible maker never rests at
3crypto “reverse bias”, day-clustered t=−4.2one crash month; month-clustered t=−0.3
4+5pp crypto favorites gap, persists OOS+0.6pp bet-weighted — below costs
5−12pp weather “miscalibration”stale last-trade marks — family prices sum to 1.39
6weather model +125%, SR 0.9−79%, SR −4.9 on books fresh enough to trade
7crypto favorites underpriced (2 verdict cells)pinned afterlife prints — 22–28% of marks postdate the close
8short-longshots OOS SR 3.9, PSR 0.99sizing artifact — same stream −3.3c/share equal-weight; reversed in 2026; retracted

Every seductive number the pipeline produced was attacked until it died or survived. A number that hasn't survived an attempt to kill it doesn't get quoted. Receipts (code + writeup per line): RESULTS.md in the desk repo.

What survived every attack

SURVIVOR Polymarket vs the equity yardstick — same Stoll (2000) decomposition, equities from raw millisecond TAQ; $-weighted, 30s horizon, bps of price.
venue / bookeff half-spread (bps)realized (bps)impact (bps)
US equity mega-cap10.80.2
US equity mid-cap3.21.21.9
US equity small-cap8.62.85.8
Polymarket tight (<1c)17.26.310.9
Polymarket wide (>3c)560.8235.5325.3

Tight prediction-market books trade like a somewhat-worse small-cap (17.2 vs 8.6 bps effective half-spread). Wide books book a “realized spread” ~80× a small-cap maker's take — the fill-at-touch mirage in one row. Regenerable: papers/paper0.

SURVIVOR The crowd beats the model — vig-stripped T-24h weather-ladder prices are better calibrated than a bias-corrected D-1 GFS+ECMWF blend: log-loss 0.360 vs 0.388 over 2,972 buckets, walk-forward. Retail prediction markets embed public NWP by the day before.
PENDING Politics favorites are overpriced (sell side) — the one candidate still standing from six pre-registered hunts: test bet-weighted −3.6pp, month-t −2.1, both engine sizing modes agree, survives slippage at 100/200/300bps, held in both 2026 months. Not claimed: PSR 0.88/0.91 < 0.95 and 7 test months < 8. Adjudication is pre-registered and pending the tail-marks crawl — pass the fixed bar and it's claimed, fail and it becomes mirage #9. No parameter may change in between (HUNT_LOG.md).

Status — what is running forward right now

Longshot forward test
379 entries
20 graded · live CLOB asks, CI 2×/day, self-grading — the sole remaining arbiter of the longshot story
Weather collector
3,069 obs
51 cities · ladder prices + GFS/ECMWF/ICON forecasts, accruing
Maker gate (real fills)
$+717 · t=1.54
15 days, rebate-driven · below the t≥2.0 bar — says so here until it clears it
Papers
0 + 1 drafted
measurement artifacts in maker backtests · favorite-longshot structure in 880k resolutions

Concluded experiment — maker v2 naive control vs v3 defended quoter, cumulative · Jul 23 – Aug 15

$-8,187$0$2,349rewardsadverse fillsNET v2NET v3 defended

v2 quotes symmetrically and never moves — it measures the toll. v3 adds the standard defenses (drift skew, circuit-breaker pulls, inventory caps) with parameters fixed a priori. The distance between the two NET lines is what the defenses recover; the distance from v3 to zero is what is left. Verdict (via the real-fill gate above): the only durable maker income is the fee rebate, roughly cancelled by adverse selection at a realistic resting size.