TOTAL VOLUME:
$134.1b
24H VOL:
$141,541,542
24H TRANSACTIONS:
2,388,728,490
OPEN INTEREST:
$1,440,096,988
406,065
Markets across
30,522
events
MATCHED EVENTS:
2,692
PLATFORM COVERAGE:
5
Polymarket:
39%
VS.
Kalshi:
61%
$
This group of markets centers around the performance of language models on the Humanity's Last Exam (HLE) in 2026. The markets aim to predict whether a language model will achieve a certain accuracy score on this exam, with scores ranging from 50% to 90% being considered.
This market will resolve to "Yes" if any model achieves at least the specified accuracy on Humanity’s Last Exam by December 31, 2026, 11:59 PM ET. Otherwise, this market will resolve to "No". For resolution, “accuracy” refers solely to the value labeled “HLE Accuracy” or a clear equivalent metric if the data’s presentation or terminology is restructured, regardless of the model’s Calibration Error or any other displayed metric. The resolution source will be the official Humanity’s Last Exam leaderboard at https://agi.safe.ai/. If the resolution source becomes unavailable during the listed timeframe, this market will remain open to allow the relevant data to become available again. If the source remains unavailable after the end of the listed timeframe or is otherwise confirmed to be permanently unavailable, official Humanity’s Last Exam results published elsewhere may be used. If no official alternative source is available, this market will resolve to "No".
Resolution is determined by whether any language model achieves accuracy at or above specified thresholds on Humanity's Last Exam before December 31, 2026. The event contains nine separate markets corresponding to accuracy levels of 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, and 90%. Each market resolves independently to Yes if any language model reaches or exceeds its designated accuracy threshold on the exam during the specified timeframe. The determination is based on publicly reported results from the exam administrators or credible third-party verification of model performance.
Prediction market odds, like those found in this market, often reflect a 'wisdom of the crowd' that can differ from individual analyst forecasts. Analysts may be subject to biases or have limited access to all relevant information, whereas a prediction market aggregates the opinions of many participants with financial incentives to be accurate. While analyst predictions can provide valuable insights, the collective forecasting power of this market can offer a distinct perspective on the potential outcomes of the Humanity’s Last Exam. It's important to consider both sources of information when forming your own view.
Prices on Polymarket and Kalshi can diverge due to several factors. Different user bases, trading fees, and market-specific rules can all contribute to price discrepancies. Polymarket and Kalshi can show different implied probabilities for the same outcome because of liquidity, fee structure, participant mix, and how each venue defines the contract. Furthermore, liquidity can vary between the platforms, meaning that a large trade on one platform might have a greater price impact than on another. The differing leading outcomes—Will the highest score achieved on Humanity’s Last Exam in 2026 be 60% or higher? on Polymarket at 49.5% and Will any LLM score at least 50% on Humanity's Last Exam before Dec 31, 2026? on Kalshi at 99.0%—also demonstrate differing interpretations of the exam's potential results. This creates arbitrage opportunities for traders.
This market resolves around Jan 1, 2027, with the outcome confirmed once the event is verifiable from credible public reporting. The highest score achieved on the Humanity’s Last Exam in 2026 will be determined by publicly available results released by the exam administrators. The market will settle based on this verified score, representing the collective prediction of traders regarding the exam's difficulty and the capabilities of those taking it. Trading will cease prior to this date, allowing for a clear determination of the winning outcome.
Several signals could significantly impact this market. Major advancements in artificial intelligence, particularly in large language models, could shift expectations about exam performance. Public statements from the exam creators regarding the exam's difficulty or format could also influence trading activity. Unexpected changes in educational trends or the availability of study resources could also play a role. Finally, any news regarding the exam's administration or scoring procedures would likely cause movement in the market, as traders adjust their predictions based on new information.