TOTAL VOLUME:
$124b
24H VOL:
$84,547,050
24H TRANSACTIONS:
2,121,338,658
OPEN INTEREST:
$1,287,835,486
364,458
Markets across
33,243
events
MATCHED EVENTS:
3,079
PLATFORM COVERAGE:
5
Polymarket:
41%
VS.
Kalshi:
59%
A backtest is only as good as the price data underneath it: survivorship bias, thin liquidity, and too few independent events can each quietly inflate your results.
Jared Polites
Sep 8, 2026

TL;DR
A backtest against prediction market history is only as good as the price data underneath it, and that data has three problems most people never check for. Resolved markets survive in the record because they resolved, thin markets move on a single position instead of genuine consensus, and a handful of past elections is not a sample size. Fix those three things first. Everything else is arithmetic.
Backtesting a prediction market strategy means replaying your entry and exit rules against historical price series (from Polymarket, Kalshi, Limitless, Predict.Fun, or Opinion) to see what would have happened. The mechanics are borrowed from quantitative finance. The pitfalls are not the same, because prediction markets resolve to $0 or $1 instead of drifting along a price curve, and that binary outcome changes how bias creeps into your results.
In equities, a backtest is testing whether your entries and exits beat a continuous price path. In prediction markets, you are testing whether you can correctly read a probability before it resolves to a certainty. That distinction matters because the failure modes are structural, not just statistical.
Say you backtest a strategy that buys any contract trading below $0.20 and holds to resolution, on the theory that markets underprice long shots. Across a set of resolved markets, this might show a positive return.
But the contracts that make it into your dataset are the ones that already resolved. The ones still open, thinly traded, or delisted before resolution are invisible to you. That gap is survivorship bias, and it is the first thing to check.
Prediction market archives contain resolved markets by definition. A market that never attracted enough volume to resolve cleanly, or that a platform pulled for ambiguous settlement language, typically doesn't show up in a clean historical export. If your backtest only sees markets that resolved smoothly, you are testing a survivor-only universe, and survivors are not representative of the full population you would have traded in real time.
The fix is not complicated, just tedious. Pull the full market list for the period you are testing, including markets that were delisted, disputed, or never reached meaningful volume, and account for what happened to your hypothetical position in each of those cases. If you can't get that list, say so in your methodology rather than presenting a survivor-only result as a full backtest.
A resolved market's historical price series looks like a clean line on a chart. It rarely was. A market with $2,000 in total contracts can show a price swing from $0.30 to $0.55 on a single trade of a few hundred dollars. That swing is real in the historical record and meaningless as a signal, because no strategy could have traded meaningful size at that price without moving it further.
Here's the generic version of the mechanism: liquidity tightens the spread between what a market implies and what it will actually trade at, and a thinly traded market can move sharply on one large position. A backtest that assumes you could fill your full order size at the historical mid-price is backtesting a market that doesn't exist. Before trusting a result, check the volume behind the price points you're relying on, and discount or exclude the markets where your hypothetical position size would have been a meaningful share of total volume.
The scarcest resource in prediction market backtesting is not data, it's independent events. Four US presidential elections since 2012 is four data points, not four hundred. A strategy tuned until it correctly calls all four is not validated, it is fit to four coin flips after the fact.
Sports markets offer more raw volume of resolved contracts, but the same trap applies at a different scale. A rule discovered by scanning thousands of resolved game markets for whatever pattern happened to correlate with a win is likely to be noise wearing the shape of a signal, especially once you've tried enough variations to find one that worked. The more parameters your strategy has (which platform, which market category, what price threshold, what time-to-resolution window), the more ways it has to accidentally fit the past.
Two practical guards. First, hold out a chunk of your historical data and never look at it while you're building the rule. Test only on the held-out set once, at the end. Second, prefer a rule with few parameters that makes sense before you saw the data over a rule with many parameters that only makes sense after.
None of this requires trading on any single venue. Because PredictionHero aggregates historical odds and consensus probability across Polymarket, Kalshi, Limitless, Predict.Fun, and Opinion, you can build a backtest that checks whether a strategy holds up across platforms before you ever risk capital on one of them, which is a stronger test than any single-platform history can give you. A rule that only works on one platform's price history is a rule that found that platform's specific noise, not a rule that found something true about how markets price the event.
If you're pulling raw price history to build this yourself, Polymarket publishes the deepest historical order book for global events, and Kalshi's CFTC-regulated structure gives you a clean US-regulated venue to cross-check against.
Limitless and Predict.Fun both operate on-chain with full on-chain settlement records, and Opinion's macro-focused contracts are useful for testing strategies against economic data releases specifically.
Full historical price series for every market in your test universe, including markets that never resolved cleanly, plus volume or open-interest data at each price point so you can filter out fills you couldn't actually have gotten.
There's no fixed number, but treat correlated events (multiple markets on the same election, the same season, the same central bank cycle) as one data point, not many. A strategy validated on four presidential elections has an effective sample size of four, regardless of how many individual state or county markets you pulled from them.
Yes, but test each platform separately first. Different user bases and collateral types can produce different pricing behavior for the same event, and pooling before you've checked for that difference can hide a platform-specific pattern inside what looks like a general result.
Testing only on markets that resolved cleanly. That silently excludes disputed, delisted, and low-volume markets, which biases the result toward whatever made a market "clean" in the first place.
Ready to check a rule against real historical odds? Polymarket has the deepest historical order books for global events, and Kalshi gives you a CFTC-regulated US venue to cross-check against. Limitless and Predict.Fun both settle on-chain with a full public record, and Opinion is worth a look for macro-driven event history.
PredictionHero aggregates publicly available prediction market data for informational purposes only. This is not financial advice. Prediction markets may not be available in all jurisdictions.
Related

Options and perpetuals hedge price. Prediction markets hedge the event you are actually afraid of. Here is how to build the cleaner hedge.
Jared Polites · Jun 16

A three-candidate market pricing at 105 cents isn't broken. Here's what that extra 5 cents actually means and why sportsbook vig is a different animal.
Jared Polites · Sep 6

Position size alone proves nothing. Here's how to separate real information from ordinary risk-taking and wash trading, and why some platforms let you check a wallet's history while others don't.
Jared Polites · Sep 4