TOTAL VOLUME:
$134.1b
24H VOL:
$113,466,932
24H TRANSACTIONS:
2,388,728,490
OPEN INTEREST:
$1,423,222,590
402,751
Markets across
30,217
events
MATCHED EVENTS:
2,632
PLATFORM COVERAGE:
5
Polymarket:
39%
VS.
Kalshi:
61%
Closed: Feb 27, 7:00 PM EST
Polymarket
This market tracks whether any Anthropic Claude model will achieve at least a 50% score on the FrontierMath Exam, a rigorous benchmark designed to test advanced mathematical reasoning in AI systems. On Polymarket, the leading outcome currently stands at 51.0%. Resolution will be determined by Epoch AI's Frontier Math benchmarking leaderboard, with credible reporting used as a secondary source if needed. Watch for official Claude model releases and their performance announcements on the Epoch AI leaderboard leading up to the June 30, 2026 deadline.
This market will resolve to "Yes" if any Anthropic Claude model achieves the listed score or greater on the FrontierMath Exam by June 30, 2026, 11:59 PM ET. Otherwise, the market will resolve to "No". This market will resolve according to the Epoch AI’s Frontier Math benchmarking leaderboard (https://epoch.ai/frontiermath) for Tier 1-3. Studies which are not included in the leaderboard (e.g. https://x.com/EpochAIResearch/status/1945905796904005720) will not be considered. The primary resolution source will be information from EpochAI; however, a consensus of credible reporting may also be used.
Prediction market odds on Polymarket reflect real-time aggregation of trader beliefs about Claude's FrontierMath performance, often incorporating faster-moving information than traditional analyst reports. While formal analyst forecasts on advanced AI benchmarks remain limited, the market's current pricing suggests strong confidence in Claude achieving the 25% threshold. Prediction markets typically embed more granular, up-to-date assessments of AI model capabilities than static analyst estimates, especially for rapidly evolving benchmarks. Comparing market odds to any published analyst commentary or expert commentary on frontier mathematics performance can help contextualize the market's conviction level.
The market resolves at Jun 30, 2026, marking the deadline for determining whether Anthropic Claude achieves the specified performance threshold on the FrontierMath Benchmark. Resolution depends on official results published by the FrontierMath organizers and Anthropic's public announcements regarding Claude model performance on the exam. The binary outcome—yes if Claude scores at least 25%, no otherwise—is determined by verified benchmark data. Traders should monitor Anthropic's official communications and FrontierMath's published results as the resolution date approaches to understand how the outcome will be confirmed.
Key catalysts include Anthropic's announcements of new Claude model versions, improvements to mathematical reasoning capabilities, or early benchmark results from related mathematics competitions. Public statements from Anthropic executives about frontier math performance, research papers detailing Claude's reasoning enhancements, or leaked benchmark scores could shift market odds significantly. Competitor announcements—such as other AI labs publishing FrontierMath results—may provide context for Claude's expected performance. Media coverage of AI progress on advanced mathematics, updates to the FrontierMath benchmark itself, or third-party evaluations of Claude's mathematical abilities could also influence trader positioning ahead of the Jun 30, 2026 resolution.