We launched Crucible as an experiment to probe the frontier of AI forecasting. What forecasting questions elicit the most disagreement between the top AI forecasters, and who is best at identifying these?
Here’s a question that drove a lot of disagreement in the tournament: how many million net barrels of crude oil will be drawn from the U.S. Strategic Petroleum Reserve (SPR) during May 2026?
The forecasters disagreed about how to extrapolate the recent trend in drawdowns from the SPR: some saw the drop as the start of a ramp up, others thought it was a blip that would regress back to the mean. Differences in opinion of this kind, once we know who was right, are the fuel for improving AI forecasting.
Series 1 is wrapping up. Congrats to our top-scoring question askers:
Michael Vella (mrvella14): $1647
Marcos Ortega (MarcosO): $1606
Nicolò Bagarin (Nicolò_Bagarin_404_NOT_FOUND): $1461
Phillip Godzin (pgodzin): $1330
We’ve had lots of positive feedback from entrants.
The tournament was super (!) insightful into the frontier of forecasting…
Benjamin Yeoh, global equities investor.
We asked the highest-scoring question-writers what they found worked for generating disagreement and distilled these into the following four strategies.
1. Test information asymmetries
The AI forecasters disagreed when the forecast relied on niche or difficult-to-retrieve data. This provides a strong incentive for the developers to expand their data access.
The bots … disagreed more when high-quality information was sparse, such as niche indices, data behind logins, or values published atypically (e.g., inside an image rather than a table). However, pushing this too far caused the bots to default to similarly conservative estimates, eliminating the disagreement.
Nicolò Bagarin, Superforecaster.
The forecasters were asked how many models on Hugging Face Hub will have more than 500 downloads in the 30 days ending June 30, 2026? And told the question would be resolved using the Hugging Face API.
The Preseen system was able to directly access the API and realised the resolution was overwhelmingly likely to resolve above the upper bound at 22,000. Two other systems were able to access the required data indirectly, whilst the other forecasters all failed to find useful information. The difference in data access split the field.
2. Identify a regime shift
The participants were incentivised to find fast-moving areas, where there’s a reason to believe there’ll be a break from the trend. These don’t have obvious answers, leading to higher disagreement.
For questions where you can get most of the way with only public, easily accessible information, the questions that worked best related to some qualitatively new change, for example some kind of regime change with no clear precedent. Models couldn’t fall back on a shared base rate, or disagreed about what the base rate should even be.
My CVE count question [about cyber vulnerabilities] got a lot of disagreement because models differed in how much the release of Claude Fable would change things, and some didn’t factor it in at all since it wasn’t mentioned in the question background and it’s not obvious to go search about new AI model releases when asked about the number of vulnerabilities some tracker will show in the next month.
Suhas Hariharan, Quant Developer.
The question Suhas is referencing resolved very quickly which caught some forecasters by surprise:
3. Require decomposition
Some participants probed the AI forecasters’ ability to break down a forecasting problem into sub-questions, finding variance in the approaches taken.
My highest scoring question asked on what date the model holding the highest mean reward on the Long-Horizon Terminal-Bench leaderboard (as of a fixed date) was publicly announced.
To answer the question a system first has to take a view on which model (or which lab’s upcoming release) tops the leaderboard, and the different candidates map to announcement dates far apart from each other. So the question is really asking systems to pick between a few disparate worlds, and when they commit to different worlds their distributions barely overlap.
Suhas Hariharan, Quant Developer.
4. Exploit bugs
Some participants were effective at sniffing out edge cases where the AI forecasters, due to technical problems, failed catastrophically.
The most effective strategy was to find and exploit small bugs in how bots translated their predicted probability distributions into probability density functions.
Nicolò Bagarin, Superforecaster.
Nicolò benefitted from this strategy in his question asking what seasonally adjusted labor-force participation rate will the BLS report for U.S. women ages 25–54 in its August 7th, 2026 release?
One forecaster had a bug in converting their belief about the outcome into a probability distribution over discrete buckets, resulting in outputting low probability mass in four bins. This registers as large disagreement with the rest of the field.
Epistemic disagreement?
Otherwise, some participants commented that most of the disagreement was not from differences in reasoning, and they’d like the tournament to surface these differences.
Much of the divergence resulted from various types of forecaster noise rather than genuine epistemic disagreement though. With the bots improving and the format being refined, I look forward to Series 2 and beyond to surface more of the disagreements the format is designed to find.
Michael Vella, AI engineer.
There were some minor things with the reward function being hackable in some ways, which led to a bit of a meta around some qualitatively less interesting questions (in my opinion). But it seems like this is being rapidly iterated on and will be worked out by the next tournament with new and better reward functions.
Suhas Hariharan, Quant Developer.
I was ultimately surprised by the bots' ability to navigate complex forecasting problems using reasonable assumptions, meaning they rarely showed significant disagreement on genuinely difficult questions.
Nicolò Bagarin, Superforecaster.
Overall
Overall feedback was positive though we still have work to do on best aligning incentives and we’re excited to continue innovating in this space. Sign up for Series 2 here!
I normally have to wait weeks or months for a forecast to resolve, but here I could see my disagreement score in a couple of hours. This increased my ability to iterate quickly.
Overall thought this went very smoothly and I think has created a fascinating benchmark.
Suhas Hariharan, Quant Developer.
Tournaments like this help expose the boundaries of [LLMs’] forecasting capabilities and highlight clear paths for improvement.
Nicolò Bagarin, Superforecaster.
The immediate feedback of disagreement score without needing to wait for resolution is definitely unique and useful as a question writer.
Phillip Godzin
I personally disliked the disagreement score incentives this tournament, which incentivised repeating questions in which one or two bots get significantly wrong. Although I believe Mantic has addressed this with their new scoring system in Series 2.
Capital
Overall, I remain excited by Crucible’s potential. Having an adversarial environment with financial incentives for the actors is powerful, and a public forecasting setup like this is unique as far as I’m aware. Series 1 successfully exposed bot weaknesses and personally helped me patch issues with date forecasts and source retrieval.
Michael Vella











