VIG
the house takes its cut — the research record
2026-08-02 17:49
← 06 RESEARCH / deep_research_arb_2026-07-31.md

Prediction-market arbitrage: building the best machine

Deep research report, 2026-07-31. All claims carry their source in parentheses. Items that failed or escaped verification are flagged inline and collected in the "Refuted or unverifiable" section. Verification verdicts (CONFIRMED / PARTLY / unconfirmed) from the independent check pass are weighed throughout.

Verdict

Keep negRisk set rebalancing as the core engine and spend the next build cycle on presence, not cleverness: move detection off GitHub Actions cron onto an always-on free VM running a websocket book mirror, make every threshold fee-aware (Polymarket charges taker fees since Mar 30, 2026), and filter augmentable events before firing. The evidence that justifies going live at $5-20k: within-event rebalancing is where essentially all historical arbitrage dollars were, $10.6M single-condition plus $32.3M set rebalancing against a $39.6M total (components overlap slightly), versus $95.2K for all cross-market relation trades over a full election year (arXiv:2508.03474, AFT 2025, CONFIRMED), and the surviving residue is depth-capped to tens of shares per episode (arXiv:2605.00864, CONFIRMED), which means a $5-20k book is already at or above the size the market can absorb per opportunity and adding capital past that buys nothing. The order: (1) websocket watcher on a free VM with fee-aware, augmentation-filtered set detection; (2) event-window scheduling around FOMC/CPI/elections/live sports; (3) a small, separately risk-scored relations book (deadline chains, duplicates) with resolution-criteria diffing and correct negation handling; (4) sports grids as a special case of (1); (5) cross-venue as measurement only, no capital, because settlement divergence converts those spreads into unhedged bets.

Taxonomy: what is a lock and what is a spread

Two classes are true locks when confined to a single venue and single oracle: single-condition rebalancing (YES+NO priced away from $1) and negRisk mutually-exclusive-set rebalancing (bid-sum of all legs sold above $1, or ask-sum bought below $1). Everything else is resolution-dependent. This classification is my synthesis, not a sourced taxonomy (verification: PARTLY), and even the "locks" retain 50/50-resolution, added-outcome, and post-hoc-clarification edge cases covered in the risk section.

The dollar history ranks the classes in the same order. The best measurement of Polymarket arbitrage, Apr 2024 to Apr 2025 on-chain: total realized profit $39,587,585; single-condition rebalancing $10.6M ($5.9M buying both sides below $1, $4.7M selling above); within-multi-condition set rebalancing $32.3M (buy-YES-set $11.1M, buy-NO-set $17.3M, sell-YES-set $612K); cross-market dependent-pair arbitrage only $95.2K across 5 pairs, best pair $60,237 (Saguillo, Ghafouri, Kiffer, Suarez-Tangil, "Unravelling the Probabilistic Forest", AFT 2025, arXiv:2508.03474, all figures CONFIRMED against full text). Only about 1% of estimated US-election opportunities were exploited (same source; the 1% figure is specific to the US-election subset). Cross-market logical-relation arbitrage was economically negligible even in the best possible year for it. This independently corroborates the desk's own census finding that the standing pool outside event windows is near zero.

The negRisk mechanism itself: a winner-take-all event deploys via the NegRiskAdapter over the Gnosis CTF; converting NO tokens in m of n markets yields 1 YES token in each other market plus collateral proportional to m-1, capping collateral at the $1 max payout instead of the sum of leg prices (docs.polymarket.com/developers/neg-risk/overview plus NegRiskAdapter.sol source, CONFIRMED; the capital-cap consequence is inference from the mechanism, not quoted doc text). One correctness hazard is confirmed and load-bearing: augmented negRisk events contain placeholder outcome slots named later and sometimes an explicit "Other" bucket, and the docs warn to trade only named outcomes (docs.polymarket.com/developers/neg-risk/overview, CONFIRMED). Bid-sum over visible legs of an augmentable event is not the full partition, so "sell all legs > $1" can be a false lock. The negRiskAugmented boolean is not on that docs page but exists in the Gamma markets API per independent API documentation (verification note on claim; treat the flag location as probable, the hazard as confirmed).

Deadline/containment chains: "X by EARLIER" is contained in "X by LATER", so no-arb requires P(YES by-earlier) <= P(YES by-later), and the only lock direction is buy the cheap superset, short the expensive subset. In NO-space the implication reverses: NO(by-later) implies NO(by-earlier), so a bot treating NO legs as monotone in the same direction as YES legs constructs positions with correlated, not offsetting, payoffs. This has no primary literature; it is derived logic (unconfirmed, flagged as such). The nearest measured analogue is the AFT dependent-pair category at $95.2K. Even with the direction right, the pair rests on two separately worded UMA questions and can be split by wording drift (timezone, "announced vs occurred"); the MSTR case in the risk section is the live demonstration.

Duplicates and cross-venue: across ten venues and 100K+ events 2018-2025, ~6% of events are concurrently cross-listed and semantically equivalent markets show persistent execution-aware deviations of 2-4% on average, with deviations persisting >=30 minutes averaging ~$0.03 and reaching ~$0.07, attributed to constrained arbitrage capacity rooted in semantic non-fungibility rather than mispricing (Gebele & Matthes, arXiv:2601.01706, CONFIRMED). The spreads persist precisely because they are not riskless: Kalshi settles by its own rulebook, Polymarket via UMA, and near-identical contracts have resolved oppositely (defirate.com settlement comparison, trade press, probable). Sports cross-venue is the closest to a lock because both legs settle on the same official score; the pre-crypto precedent is bookmaker-vs-exchange arbitrage in 19.2% of top-5-league soccer matches (Franck, Verbeek, Nuesch, Economica 80(318), 2013, verified).

Sports grids: a correct-score grid is structurally a negRisk partition, so grid arbitrage reduces to set rebalancing within the grid (true lock) plus grid-vs-aggregate consistency (also a lock on the same venue when both derive from the same official result). No dedicated paper exists (derived, unconfirmed); the nearest evidence is the NBA study: 173 games, 75M book snapshots, only 7 executable single-market episodes with median lifetime 3.6s, 290 combinatorial episodes with median return 101 bps, but 76.9% depth-capped to an average 14.8 executable shares, concentrated in final minutes of live play (Cheng, Yang, Zou, arXiv:2605.00864, CONFIRMED). Detection is not the constraint; depth is.

Why the leak exists at all: arbitrage-free pricing over logical relations is #P-hard in general (Kroer, Dudik, Lahaie, Balakrishnan, EC'16, arXiv:1606.02825; Dudik/Lahaie/Pennock EC'12; EC'13 election deployment, verified), so venues price related markets independently outside negRisk sets and structurally leak cross-market inconsistency.

Execution mechanics: fees, migration, atomicity

The fee regime changed and it moves the arb threshold. Fee Structure V2, effective Mar 30, 2026: taker fee = shares x feeRate x p x (1-p), with feeRate 0.07 crypto, 0.05 sports (raised from 0.03 on Jul 10, 2026), 0.05 economics/culture/weather/other, 0.04 finance/politics/tech/mentions, 0.00 geopolitics/world events; makers never pay (docs.polymarket.com/trading/fees plus changelog, CONFIRMED). A sell-all-legs set arb crosses the spread on every leg, so the threshold is no longer bid-sum > $1.00 but bid-sum > $1.00 + sum over legs of feeRate x p_i x (1-p_i). The FOMC/rates category assignment is not documented (unconfirmed); the machine must read each market's actual fee fields from the API rather than assume, and geopolitics is the only guaranteed fee-free surface. Maker rebates run 15-25% of taker fees, paid daily per-market (docs.polymarket.com/programs/maker-rebates, CONFIRMED), and the liquidity-rewards program (quadratic scoring against per-market max spread, 10,080 samples/epoch, daily payout, $1 minimum) means resting arb quotes can earn rewards plus rebates simultaneously (docs.polymarket.com/programs/liquidity-rewards, CONFIRMED).

The platform migrated under the desk's feet. CLOB V2 went live Apr 28, 2026: pUSD (1:1 USDC-backed ERC-20, proxy 0xC011a7E12a19f7B1f670d46F03B03f3342E82DFB) replaced USDC.e as collateral; new contract addresses (CTF Exchange 0xE111180000d2663C0091e4f400237545B87B996B, Neg Risk CTF Exchange 0xe2222d279d744050d28e00520010520000310F59, Neg Risk Adapter 0xd91E80cF2E7be2e162c6513ceD06f1dD0dA35296); relayer calls to the v1 Neg Risk Adapter deprecated Jul 14, 2026; and as of Jul 17, 2026 POST /order(s) returns tradeIDs instead of transactionHashes, breaking any pipeline that parses the old field (docs.polymarket.com/resources/contracts and changelog, CONFIRMED). Any code holding v1 addresses needs migration, and the py-clob-client repo is reportedly archived as of May 25, 2026 (unconfirmed); confirm the currently supported client library before building on it.

Multi-leg atomic execution does not exist at the CLOB layer. The batch endpoint takes up to 15 signed orders but processes them in parallel with independent per-order results, explicitly not all-or-nothing (docs.polymarket.com/api-reference/trade/post-multiple-orders, CONFIRMED). The mitigation stack: per-leg FOK orders, batch-post all legs in one request, depth checks immediately before send, and the on-chain NegRiskAdapter convert, which is atomic in a single transaction and serves as a repair path for a stuck leg (docs plus contract source, CONFIRMED). One caveat the verification pass surfaced: the docs are silent on conversion fees, but NegRiskAdapter.sol contains explicit fee logic (amountOut = amount - feeAmount), so a convert fee can exist even though it has historically been zero; check the live feeBps before relying on convert economics (verification of claim 13, PARTLY).

Order plumbing: GTC/GTD limit (GTD minimum 3 minutes out, expires 1 minute early), FOK/FAK market, post-only added Jan 6, 2026 (docs, CONFIRMED). Marketable orders on select crypto/finance markets sit in a 250ms async matching delay and sports markets in configurable in-play delays during which they cannot be cancelled; this delay window is the main legging-risk exposure (docs.polymarket.com/concepts/order-lifecycle, CONFIRMED). Treat a trade as settled only at CONFIRMED status. The heartbeat mechanism (ping every 5s, mass-cancel after 10s silence) is a free dead-man switch for any resting quotes (docs, CONFIRMED). Trading is gasless via Polymarket's relayer for trades, splits, merges, redeems, and approvals, with Builder credentials required for programmatic gasless ops (docs.polymarket.com/trading/gasless, CONFIRMED); my sub-cent self-submitted gas estimate is unconfirmed.

Infrastructure: the 60s cron is the weakest link

Polymarket officially documents a public market websocket (wss://ws-subscriptions-clob.polymarket.com/ws/market: book, price_change, last_trade_price, tick_size_change per token; PING/PONG every 10s) and an authenticated user channel pushing order and trade lifecycle events through settlement finality with dynamic subscribe/unsubscribe (docs.polymarket.com/developers/CLOB/websocket, CONFIRMED). REST limits are generous and throttled rather than rejected (book/price 1,500 req/10s, general 9,000/10s; docs.polymarket.com/api-reference/rate-limits, CONFIRMED). The event-driven book mirror is free to build and strictly dominates polling.

Meanwhile the current scheduler cannot even deliver its nominal cadence: GitHub Actions cron has a 5-minute minimum interval and routine 5-30+ minute scheduling delays (GitHub community discussions and docs, verified for the mechanism), so "60s sweeps" are only achievable inside a long-running looping job capped at 6 hours. An always-on VM running a websocket consumer beats the current setup by 2-4 orders of magnitude in reaction time at zero cost.

Speed is not the edge, presence is. Single-market episodes have ~3.6s median lifetime (an upper bound, polling-limited) and the binding constraint is book depth, not latency (arXiv:2605.00864, CONFIRMED). Sub-second infrastructure does not unlock the $5-20k scale; what matters is being connected during event windows, where a 60s poll statistically misses most fillable episodes (synthesis, probable). Latency tiers for reference: EU VPS ~22-26ms warm order round-trip, US boxes ~70-90ms transatlantic, home US websocket ~100-150ms effective; all adequate (tradoxvps.com vendor benchmark, secondary source, probable).

Free compute at $0 budget: GCP e2-micro remains always-free but only in us-west1/central1/east1 (cloud.google.com/free; region list matches published terms, not re-verified this pass); Oracle Always Free ARM was silently cut from 4 OCPU/24GB to 2 OCPU/12GB effective Jun 15, 2026, still ample, EU regions available, idle-reclaim risk (InfoQ Jul 2026, CONFIRMED); fly.io's permanent free tier is gone (CONFIRMED). Time-sensitive: DigitalOcean's $200 GitHub Student Pack credit reportedly retires Aug 1, 2026, making today the last redemption day (secondary source, probable); Azure for Students $100 remains.

Competition: who is in the pool

Extraction is concentrated in a few bot-like wallets: the top wallet took $2,009,632 over 4,049 fills, top-10 wallets $8.8M total with per-wallet takes of $384K-$2.0M (arXiv:2508.03474 Table 1, CONFIRMED). Single-condition opportunities were "largely captured"; sports were "surprisingly absent" from extraction despite dominating the opportunity space, and the paper attributes the untouched cross-market residue to non-atomic multi-leg execution risk (same source, CONFIRMED; my earlier note that "several pairs had zero executed arb" over-read the table and at most one of five listed pairs showed zero).

Institutional entry is real: Susquehanna became Kalshi's first dedicated institutional market maker in April 2024 (Businesswire via mirrors, CONFIRMED; the specific "reduced fees and higher position limits" terms could not be verified); Wintermute began two-sided liquidity on both Polymarket and Kalshi around May 29, 2026 (decrypt.co/369475, CONFIRMED); trade press reports DRW, Jump, SIG, Flow Traders building prediction-market desks framed around arbitrage of fragmentation (Finance Magnates, Jan 2026, probable). Entry-level competition is commoditized: open-source Polymarket/Kalshi arb bots and retail products like Arbigab exist publicly (GitHub, arbigab.com, probable). Market scale context: combined Kalshi+Polymarket monthly volume grew from under $5B in Sep 2025 to ~$24B in Apr 2026, exceeding the ~$14B/month average US legal sportsbook handle (Pew Research, 2026-05-27, CONFIRMED); category mix is sports-dominated on both venues (Kalshi 80% sports; Polymarket 39% sports / 32% politics / 20% crypto, Pew, CONFIRMED).

Strategic read: the literature reproduces the desk's own census. The residue that survives resident bots and incoming institutions is (a) small-size due to depth, (b) concentrated in high-volatility event windows, and (c) protected by non-atomic execution risk that large players' size floors make uneconomic (synthesis, probable). That residue is exactly the shape a solo desk at $5-20k with breadth and event-window presence can harvest.

Risk: what breaks the locks

Oracle risk is adversarial, not noise. Confirmed incidents: the ~$7M "Ukraine mineral deal before April" market forced to YES in March 2025 by a whale with ~5M UMA across three wallets, ~25% of active voting power, no refunds (CoinDesk, The Block, The Defiant, CONFIRMED); the ~$160-240M Zelenskyy suit market initially resolved YES then flipped to NO after nine days of disputes, capital locked throughout (CoinDesk, Decrypt, Forbes, CONFIRMED); the $16.5M Clavicular pregnancy dispute, with reporting that the ten largest UMA wallets control >50% of votes and ~60% of UMA voters link to Polymarket accounts (Forbes, Apr 2026, CONFIRMED). Galaxy Research counted 1,150+ disputed markets in 2026 YTD by mid-June, already exceeding all of 2025 (cryptodaily.co.uk, probable). Mitigant with a new concentration: Managed Optimistic Oracle V2 (Nov 2025) restricts resolution proposals to 37 vetted addresses while keeping disputes permissionless (The Block, UMA blog, CONFIRMED).

Set-arb-specific hazards, all confirmed at the mechanism level: (1) a leg resolving 50/50 instead of NO pays $0.50 per share against a cents-wide edge; the 50/50 outcome is official (docs.polymarket.com/concepts/resolution, CONFIRMED). (2) If the oracle returns a second YES in a negRisk set, reportOutcome reverts and the whole set can freeze with collateral inside (NegRiskAdapter README, CONFIRMED). (3) Post-hoc rule clarification: the ~$80M MSTR "sold any Bitcoin by May 31" market resolved NO despite a June 1 8-K disclosing May 26-31 sales, after an additional-context clarification traders say was not in the written rules, while the June-30 sibling resolved YES (CoinDesk, crypto.news, PARTLY: core confirmed, exact volume, the $527k wager, and the verbatim clarification quote unverified). This is the direct empirical warning for deadline-chain trades: date-bracket siblings paid inconsistently with the underlying event. (4) Early resolution: UMA supports a "too early" dispute outcome and deadline markets legitimately resolve YES the moment the event occurs, desyncing a sold set (docs, CONFIRMED for the mechanism; the desync consequence is inference).

Settlement manipulation is documented at scale: Dai, Jia, Yu, "Settlement Manipulation in Prediction Markets" (arXiv:2606.31675, June 2026, CONFIRMED real): on ~16,000 5-minute BTC contracts Feb-Apr 2026, settlement-time Binance order flow spikes and reverses, ~821 flagged traders captured ~$8.2M in two months, largely absent in 15-minute contracts. Aftermath: Bloomberg coverage Jul 15, 2026; ~100 wallets referred to law enforcement; CFTC March 2026 guidance; renewed regulator scrutiny reported June 2026 (Bloomberg, Forbes, CONFIRMED). Separately, the CFTC has an ongoing "extensive" investigation touching a staged-trades marketing campaign (~1,105 fake-win videos, creators paid $2-3k/month), with a bipartisan Senate letter to Chairman Selig (senate.gov primary sources, CNBC, CONFIRMED). Regulatory posture otherwise: CFTC-designated US DCM since Nov 25, 2025, US launch Dec 3, 2025, waitlist dropped May 2026 (cftc.gov, PRNewswire, CONFIRMED).

Custody and platform: collateral moved from bridged USDC.e toward native USDC and then pUSD, with an announced further shift toward in-house control of settlement and a planned POLY token (CoinDesk Apr 6, 2026, CONFIRMED); the Polygon PoS chain itself had a same-day emergency hard fork after a faulty milestone delayed finality 10-15 minutes on Sep 10, 2025 (CoinDesk, Metrika post-mortem, CONFIRMED). Position-level implication: any "arbitrage" whose legs can be split by a UMA vote carries tail risk that scales with position size and market prominence; locks held to resolution still carry venue, contract, and chain risk, just not resolution risk.

Refuted or unverifiable

Build order for this desk

Current state: 60s GitHub Actions sweeps (actually 5-min-plus given cron floors), week-1 realized $474 in an FOMC window, census showing ~$0.17 standing pool outside windows.

  1. Websocket watcher on a free always-on VM (build first). Persistent consumer on the market channel for all active negRisk set token IDs, local book mirror, fire full-book REST verification plus batch FOK execution when fee-adjusted bid-sum exceeds threshold. Oracle Free Tier ARM (EU) or GCP e2-micro (US); either latency tier is adequate since depth, not speed, binds (arXiv:2605.00864). Fold in immediately: fee-aware thresholding per market read from the API (bid-sum > $1.00 + sum feeRate x p_i x (1-p_i)); skip or de-weight negRiskAugmented events; parse tradeIDs not transactionHashes; verify current V2 contract addresses and a supported client library; treat CONFIRMED as the only settled state; register heartbeats for anything resting. Keep GitHub Actions as the redundant slow census and health-ledger layer only.
  2. Event-window scheduling. Concentrate the watcher's armed state around scheduled catalysts (FOMC, CPI, elections, debate nights, live sports endgames). This is where both the literature and the desk's own week-1 record locate the supply. Expected shape: episodic bursts, tens of shares deep, not standing inventory.
  3. Relations scanner with negation handling (small, risk-scored, separate book). Deadline chains and same-venue duplicates. Hard rules: only trade buy-cheap-superset / short-expensive-subset; never treat NO legs as monotone in the YES direction; diff resolution criteria text between the two markets before any fill and skip on timezone/source/"announced vs occurred" mismatches; cap the book at a small fraction of the lock book because the measured historical pool is ~$95K/yr venue-wide and the MSTR case shows even correct positions can be split by clarification. Depth-verify before sizing; expect most detections to be untradeable.
  4. Sports grids. Structurally the same engine as (1): treat each correct-score or bracket grid as a negRisk-style partition, add grid-vs-aggregate consistency checks where both settle on the same official score. Expect 14.8-share-scale fills in final minutes (arXiv:2605.00864); sports taker fee is now 0.05, so the fee-adjusted threshold matters most here. The AFT finding that sports arb went persistently unexploited is the one measured hint of un-harvested same-venue supply.
  5. Cross-venue: measure, don't trade. Log Polymarket-vs-Kalshi quote pairs for semantically matched events and accumulate the desk's own deviation and settlement-divergence dataset. The 2-4% deviations are real (arXiv:2601.01706) but are compensation for resolution risk, not free money; sports pairs settling on identical official results are the only candidate for eventual capital, and only after the measurement layer has a season of data. No Kalshi capital until then.

Sizing note for the $5-20k question: capital is not the binding constraint at any step above; depth is. Deploy small per-episode ($50-500 legs), keep most of the bankroll as dry powder for multi-set event windows, and treat any single set exposure above ~10% of bankroll as taking oracle tail risk the edge does not pay for (Ukraine minerals, Zelenskyy, MSTR all sit in the historical record as reminders).