TOTAL VOLUME:

$117.1b

24H VOL:

$122,493,166

24H TRANSACTIONS:

1,380,975,298

OPEN INTEREST:

$1,191,792,202

340,492

Markets across

34,140

events

MATCHED EVENTS:

4,396

PLATFORM COVERAGE:

5

Polymarket:

42%

VS.

Kalshi:

58%

Outcome
Trade
Chance %
Price
Spread
Liquidity
Volume
24h
7d
Open Interest
Ends in
Result
Total markets: 14

Description

This event group tracks whether any Google Gemini model will achieve specific accuracy thresholds on Humanity's Last Exam by December 31, 2026. The markets are focused on determining if the model's performance meets or exceeds various percentage benchmarks.

PredictionHero - Resolution Divergence Alerts (RDA)

Unified Resolution Criteria (Consistent across platforms)

All markets consistently define resolution based on whether a Google Gemini model achieves the specified accuracy threshold on Humanity’s Last Exam by December 31, 2026, using the same primary resolution source.Primary resolution logic: Official Humanity’s Last Exam leaderboard at https://agi.safe.ai/

Core resolution logic:

  • Resolution is triggered if any Google Gemini model achieves the specified accuracy threshold (50%, 55%, 60%, 65%, 70%, etc.) on Humanity’s Last Exam by December 31, 2026
  • Accuracy is defined as the value labeled “HLE Accuracy” or an equivalent metric on the official leaderboard
  • If the official source is unavailable after December 31, 2026, and no alternative official source exists, the market resolves to “No”

Edge cases & clarifications:

  • Source Unavailability: If the official leaderboard is temporarily unavailable during the timeframe, the market remains open until data becomes available again. If it remains unavailable after December 31, 2026, the market resolves to “No” if no alternative official source exists.
  • Multiple Thresholds: If a model achieves multiple threshold accuracies (e.g., both 55% and 60%), all corresponding markets resolve to "Yes".
Timing: Resolution occurs after December 31, 2026, once the final results from the official Humanity’s Last Exam leaderboard are confirmed and publicly available.Our PredictionHero Resolution Divergence Alerts (RDA) are there to help users identify potential differences across platforms. They do not replace or supersede the official rules and description of any prediction market. Users are solely responsible for reviewing and understanding the applicable rules and resolution criteria before placing any trade or bet. If you notice a potential inconsistency, discrepancy, or error in an alert, please report it to our team so we can review and improve the accuracy of our data.
Show more

Polymarket

This market will resolve to "Yes" if any Google Gemini model achieves at least the specified accuracy on Humanity’s Last Exam by December 31, 2026, 11:59 PM ET. Otherwise, this market will resolve to "No". For resolution, “accuracy” refers solely to the value labeled “HLE Accuracy” or a clear equivalent metric if the data’s presentation or terminology is restructured, regardless of the model’s Calibration Error or any other displayed metric. The resolution source will be the official Humanity’s Last Exam leaderboard at https://agi.safe.ai/. If the resolution source becomes unavailable during the listed timeframe, this market will remain open to allow the relevant data to become available again. If the source remains unavailable after the end of the listed timeframe or is otherwise confirmed to be permanently unavailable, official Humanity’s Last Exam results published elsewhere may be used. If no official alternative source is available, this market will resolve to "No".

Kalshi

Resolution is determined by whether any language model achieves accuracy at or above specified thresholds on Humanity's Last Exam before December 31, 2026. The event contains nine separate markets corresponding to accuracy levels of 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, and 90%. Each market resolves independently to Yes if any language model reaches or exceeds its designated accuracy threshold on the exam during the specified timeframe. The determination is based on publicly reported results from the exam administrators or credible third-party verification of model performance.

Frequently asked questions

The dashboard for the Gemini Humanity’s Last Exam score market shows real-time odds and volume across platforms. It aggregates data from multiple venues so you can see how traders price the likelihood of Google Gemini achieving specific performance levels on the Humanity’s Last Exam by the end of 2026. The view includes current probabilities, trading volume of $60,440, and 24-hour volume of $2,098, letting participants compare sentiment and positioning across the market.

On Polymarket, traders currently assign a 71.7% probability to the leading outcome, reflecting a strong consensus that Gemini will hit key benchmarks. This contrasts with many traditional analyst forecasts, which often remain more cautious due to differing methodologies and risk assumptions. The gap can provide early signals of market sentiment diverging from institutional expectations, especially as new test results or model updates emerge. Monitoring this market lets you see where crowd wisdom stands against expert opinion.

On Kalshi, pricing dynamics can diverge from Polymarket due to differences in user bases, liquidity, and market design. Polymarket and Kalshi can show different implied probabilities for the same outcome because of liquidity, fee structure, participant mix, and how each venue defines the contract. For example, Polymarket currently favors Will the highest score achieved by a Google Gemini model on Humanity’s Last Exam in 2026 be 50% or higher? at 71.7%, while Kalshi leans toward Will any LLM score at least 50% on Humanity's Last Exam before Dec 31, 2026? at 99.0%. Factors such as varying trader sentiment, regional participation patterns, and distinct order-book depths can all contribute to the spread of 27.3 between the two venues, creating opportunities for arbitrage or insight into shifting expectations.

This market resolves around Dec 31, 2026, with the outcome confirmed once the event is verifiable from credible public reporting. The final result will depend on official announcements or widely accepted test results showing the highest Gemini score achieved on Humanity’s Last Exam during 2026. As long as the data appears in reputable sources, the market will close and payouts will follow based on the recorded performance.

Key events such as major Gemini model updates, published test results from independent researchers, and announcements about new benchmark versions could all shift this market. High-impact signals include surprise performance leaks, regulatory comments affecting AI testing, or competing models posting strong scores that raise the bar. Traders will also watch for shifts in Polymarket and Kalshi odds, as changing sentiment on either venue often precedes broader market movements.