AI Superforecasters Are Here: FutureSearch CEO Dan Schwarz | Edited Transcript
FutureSearch’s CEO explains how AI forecasting crossed the superforecaster threshold, how it is evaluated, where it creates value, and how linked forecasts could become a consistent world model.
Chapter Timestamps
00:00 AI beats human superforecasters, forecasts compound, and prediction markets fall short
01:53 Why the AI-superforecaster threshold arrived in the last six months
05:43 Pastcasting and Bench to the Future: evaluating forecasts before outcomes resolve
08:30 Why forecasting skill does not automatically produce trading profit
10:35 The frontier labs, nowcasting OpenAI revenue, and where forecasting creates value
12:43 In-distribution versus out-of-distribution reasoning
17:00 AI systems versus individual forecasters, expert teams, and prediction-market crowds
19:46 Simulated markets, hedge funds, and why a true forecasting edge can remain invisible
23:47 How FutureSearch researches the present, decomposes questions, and aggregates forecasts
28:35 Consumer forecasting and decisions that can actually change personal outcomes
32:02 Forecasting as both a frontier-model capability and an ungameable evaluation
35:23 Irreducible uncertainty, intuitive physics, and hidden reasoning
42:32 Dan’s 2029 outlook and the difficulty of forecasting rapid AI takeoff
45:44 Forecast cost, test-time compute, and whether more tokens keep improving accuracy
51:34 A shared world model, causal graphs, and the danger of correlated failures
56:44 Why prediction markets did not create wiser politics and the case for slowing AI development
1:01:15 Closing remarks
Made with: The Transcript Desk Chrome Extension
Full video:
FutureSearch CEO Dan Schwarz explains why AI forecasting crossed a meaningful threshold in the last six months, how pastcasting and ForecastBench measure it, why forecasting skill is not the same as trading profit, how FutureSearch builds multi-agent forecasts, and why linked forecasts may become a consistent AI world model.
Transcript
00:00-01:52
Opening montage: Excerpts from the conversation preview three themes developed later in the interview: AI systems surpassing human superforecasters, forecasts improving one another inside a shared world model, and the limits of prediction markets as a source of political wisdom.
01:53-06:50
Nathan Labenz: So the occasion for this conversation was this post that just came out from Scott Alexander. I know you guys have collaborated on a couple things over time, but that really caught my eye because Scott is not, I would say, an uncritical AI hype man, and yet the title of the post is ‘The AI Superforecasters Are Here.’ So tell us about it.
Dan Schwarz: Scott was actually one of the first people to see the very first AI forecaster — though it depends how you define ‘forecaster.’ Time-series forecasting has existed for decades in AI and machine learning; it all depends how you define those terms. FutureSearch was started in 2023 — we started building a forecaster before GPT-4 Turbo had even come out, using Claude 2. It was very early days, and Scott tried our forecaster then. He wasn’t super impressed, and I don’t think anybody really was back then. But FutureSearch was founded on the premise that eventually we’d reach the superhuman level, and we now claim that level has arrived. It’s all been very sudden — we’ve been operating this company for about three years, and it’s really in the last six months that these things have started to genuinely scare and impress us with their capabilities.
I think the main thing to note about forecasting is that it’s a very important capability, but a somewhat neglected one. People describe it as hyped, but the whole field of forecasting is only about fifteen years old. Prior to Tetlock, nobody was even studying accuracy — nobody was recording predictions in a resolvable way and scoring them to give us even the most basic scientific evidence on what was going on. Even before Tetlock, many of us were enthused by prediction markets, but that was a very human technology. Prediction markets are very big right now — as you said, Prakash, I built the one that’s currently running at Google, and I think that’s great, but it’s a way of augmenting and coordinating human intelligence. We’re more interested in artificial intelligence now, and it’s taken a little bit of time. Midjourney was already very good at generating images — I don’t know exactly how you’d compare image generation, audio generation, essay writing, coding, math, and all the other capabilities we normally track, or say exactly when each became human- or superhuman-level; it depends how you define it. Forecasting, if anything, is a bit late to the game. But now that it’s here, it really changes the nature of what forecasting even is.
Prakash: You mentioned that in the last six months you’ve started to see things that are a step up, or different, from what you were seeing before. Can you elaborate?
Dan Schwarz: A lot of Scott’s post was trying to cover the level of evidence, because the evidence behind this is a bit distributed. The main thing that’s held forecasting back — and this applies to human forecasting too — is that you generally have to wait for the future to happen to find out if you were right. Humans generally do this in year-long tournaments: when the tournament ends, you find out which humans were best a year ago. Humans don’t get much better over the course of a year, so that’s a good indication of who’s best today. That doesn’t work with AI — if you wait a year to find out who was good a year ago, you’re getting a view of something very out of date. One of the things Scott mentioned is that we use our best forecasting to predict stock returns — we published a set of stock rankings in August 2025, basically a simple model for every stock based on forecasting certain fundamentals and extrapolating them out. We put it on the web, paywalled part of it, and waited. It’s been ten months now, and that portfolio looks extremely good — but what does that really tell you? It tells you our forecasting in that particular methodology was good ten months ago, which isn’t something most people care about now.
So we rely on a couple of different forms of evidence. Some are more short-term — there are tournaments running every couple of weeks or months. At FutureSearch we mostly rely on ‘past-casting’: taking a snapshot of the internet from some months ago and using models’ training-cutoff dates to trick them into forecasting without hindsight bias. That’s very useful because we can evaluate immediately. So when Claude Fable first came out, we were able to evaluate it within twenty-four hours, and it was the best single-agent forecaster on our leaderboard — everyone else had to wait weeks or months to find out how good Claude Fable actually was. Internally, using the benchmark we call Bench to the Future, we saw this progression in real time; the rest of the world sees it some months behind. If you read Scott’s article, you’ll see that over the last twelve months the evidence has really come in, and over the last six months, from live forecasting tournaments and performance on actual prediction markets, you can see AI is at least competitive with humans and even teams of humans working together. Whether it’s better requires synthesizing a lot of different, disparate sources of evidence. If you’re curious about this, or have forecasting needs in your life, you should really try it — just go to FutureSearch, you get $20 free so you can try a frontier forecast immediately, and judge for yourself whether you think it’s good.
06:51-10:34
Nathan Labenz: So what are you guys trying to do with this as a business? It strikes me that if you can beat the market, you’ve got a business right there — though I don’t know how durable the moats will prove to be in a space like this. One way to go, if you have the superforecaster, is just to play the markets. How much are you going to do that versus, say, sell $20 worth of forecasting to retail, versus go to enterprise customers? Where do you think the value of forecasting is highest?
Dan Schwarz: It’s a great question, Nathan. FutureSearch is almost three years in, but we’re still dabbling our toes — fingers in pies, so to speak. From the trading perspective it’s quite subtle, and you’ll see this if you look into Scott’s article and the evidence. Most money made in financial markets, trading, and prediction markets isn’t based on having a forecasting edge — it’s based on any number of other strategies. The AIs making money on Polymarket right now, for example, are mostly doing market-making, arbitrage, or front-running news; they’re not actually trying to predict the future any better, they’re taking advantage of certain properties of the markets. Broadly speaking, finance is the same way — there’s the Warren Buffett school, where if you predict future cash flows better you make money, and plenty of investors subscribe to that, but the solid majority of investors aren’t actually trying to predict the future at all; they’re taking advantage of other patterns. So you can ask: if it’s so profitable to be a superforecaster, where was the money before AI? It’s mixed — various hedge funds have hired superforecasters and set up trading operations, and there’s a bit of evidence, but it’s not really clear superforecasters can beat the market, or how. The link between forecasting and making money in finance isn’t that clear. There is some link, and we have some evidence — you can see it on markets.futuresearch.ai.
Probably the bigger question is what the role of forecasting is as an AI capability, and there the main thing to track is what the frontier labs are doing. I’ve written a lot trying to predict what the frontier labs have been up to over the years. In 2024, FutureSearch was, I think, the first to correctly break down OpenAI’s revenue — we figured out where their business was coming from between consumer, enterprise, and API, when the reporting at the time had it totally wrong. That was more nowcasting than forecasting — really just being good about uncertainty. Nobody knew where OpenAI’s revenue was coming from in 2024; there was no report you could just read, you had to piece it together from dozens of data points, leaks, and rumors. That’s the task of a forecaster, and to some degree, trying to predict the future of AI is the most valuable thing you can do with one — we didn’t wait until we had AI superforecasting to start doing it, we did it with humans, then human-plus-AI teams, and now much more just AI.
Ultimately, the future belongs to the frontier labs — what OpenAI and Google do is largely going to determine the outcomes for humanity from an AI perspective. So the real question is: what do the frontier labs think about forecasting? Are they using it to train their models? Is forecasting an important eval? Do they care how good their models are out of the box at predicting the future? I’ll leave it there, but you can guess where the real money in forecasting is.
10:35-13:43
Nathan Labenz: So — don’t leave it there. How would you characterize the companies in terms of their relative positions or outlooks on these forecasts?
Dan Schwarz: I really can’t say — I’m NDA’d by a few of the frontier labs, so I think as more time passes we’ll learn more. The simplest way to think about it is that forecasting is an important AI capability that people haven’t paid as much attention to compared to things like image generation. I think frontier labs are now paying a lot of attention to it, and it’s very relevant to their day-to-day operations. Think about human forecasting as decision support: what would you do if you had a superforecasting human sitting right next to you that you could ask about any problem? You’d probably get some use out of them, though it’s not exactly clear what. But now imagine something much, much better than the best humans at predicting outcomes — good at conditional forecasting, like ‘if I have this guest on my podcast, how many views will I get, versus that guest?’ If something superhuman in that capability were available at your fingertips, you’d start to think about forecasting very differently compared to the last fifteen years’ model of hiring a human, having them spend a couple of days on it, or putting a question on a prediction market and waiting for a couple hundred people to study it. That’s slow and not tailored to your needs. You can imagine what’s going on in the backrooms of these frontier labs as they make strategic decisions. Here’s a public data point: OpenAI hired superforecasters to evaluate the GPT-4 release, and DeepMind later hired superforecasters to look into Gemini. Even before AI superforecasting existed, there was a need for decisions about how to release these models and what policies to have, and the labs found it worthwhile to bring in superforecasters. Now you have to wonder — would they hire human superforecasters again, or use the AI they have in-house, or something like FutureSearch, to make those decisions?
Prakash: Let me ask a question. Some people still hold the idea that LLMs are stochastic parrots — Gary Marcus and others would probably say that. On the other hand, there’s the idea of modern language models as reasoning engines, giving you in-distribution answers drawn from a historical corpus versus out-of-distribution answers arrived at through actual reasoning. To what extent do you see both of these aspects when they forecast? How much are they referring to a historical corpus of things that have happened, versus reasoning through something that hasn’t happened, or may never have happened, before?
13:44-17:38
Dan Schwarz: That’s a great question, Prakash, because one of the concerns a lot of people have about AI superforecasting is that it’s too in-distribution. I actually heard this from one of the best forecasters I’ve ever had the pleasure of working with. He believes a system like FutureSearch would beat him head-to-head in a tournament about near-term, in-distribution outcomes. But in some post-AGI world, some world with transformative AI, he thinks he’d have a huge edge over the AIs, for exactly the reason you gave — they’re trained to predict things that have actually happened, and when things get wonky you need creative, lateral thinking. I think the rate of AI improvement is so astounding that even that kind of lateral thinking — imagining a completely different scenario — will fall to the AIs. Unfortunately, it’s hard to test this. The more AI continues doing strange things to the world, and we wake up to strange things in the news that become Metaculus questions and ForecastBench questions, and teams like mine try to predict them better, the more evidence we’ll get. But if there’s a real step-change in the nature of the world — some AGI or transformative-AI world, ‘geniuses in data centers,’ or something like the AI 2027 scenario — I think it’s going to be the wild west. I’ll say humans aren’t doing particularly great at imagining transformative AI either, so the bar is a bit lower.
When you actually play with these AI forecasters, you’ll find them quite human in how they structure their reasoning — and that’s not an accident, since they’re trained on how humans have structured reasoning before. A human forecaster loves to say, ‘what happened the last ten times something like this occurred, and what were the outcomes?’ — building a distribution from that. Whether an AI does the same thing because it independently arrives at the same conclusion, because it’s trained on humans doing that, or because it just thinks like a human, I don’t think we have the answer yet. Superhuman reasoning is hard to measure — would you even know it if you saw it? Prakash and I were discussing this before the show, using Fable a lot — I’ve found the way Fable explains things is a little alien compared to how Opus or GPT-5.5 explains things. It’s very concise, the sentences are shorter and full of jargon, like it’s compressing more information into a sentence than humans normally do. To me, that’s the shoggoth starting to show from behind the mask — the alien intelligence is a bit more alien now than it was a month ago, and I don’t think it’d be a wild prediction that we’ll see more of that as post-training gets more specialized and models get larger.
How does this manifest from a superforecasting perspective? Maybe superforecasting is actually the way to look at it. If you look at a great codebase, you might say, well, John Carmack could have written this — it looks great, but a great human would have done it too. But if you look at a brilliantly reasoned strategy — if the administration does this, here are the outcomes — you might start to see something that looks a little alien compared to how any human has analyzed it, and that might be a sign the AI is actually starting to surpass humans.
Nathan Labenz: In terms of surpassing humans, can you help me calibrate? There are individual human forecasters, even individual human superforecasters, and then there are aggregate measures — the Metaculus community, or the market-clearing price — and sometimes those are manipulated, or subject to insider biases, which might in fact make them more accurate. When we say we now have an AI superforecaster, does that mean it’s matching individuals? How close is it coming to those aggregates of humans?
17:39-21:41
Dan Schwarz: It’s interesting — it’s progressed so fast that it seems to have blown through individual humans and reached the level of teams of expert humans faster than anybody noticed. I’m not fully confident in this, because we don’t have the right experiments yet, but the two easiest sources of data are: first, ForecastBench, from the Forecast Research Institute — on their leaderboards you’ll see a superforecaster median, and just in the last couple of months various AI systems, including FutureSearch, are now above that median. Second, the Metaculus tournament series comparing bots and humans, which has been running for almost two years — AI systems are doing better and better. Metaculus also runs head-to-head tournaments open to both AIs and humans, comparing AI responses to what Metaculus calls the ‘community prediction’ — a slightly more sophisticated median of Metaculus forecasters — and that’s also showing multiple AI systems ahead of the community prediction.
Prediction markets are another strange source of evidence, because if you make even a single profitable trade, you’re allegedly beating the entire market. FutureSearch has made hundreds of trades — they’re on our website, so you can see them. Our profitability fluctuates wildly at a low sample size, so it’s hard to say much, but if you see an AI that isn’t market-making or arbitraging, but doing what FutureSearch does — once a week going to Polymarket and Kalshi and picking the markets we think have the most volume and are the most geopolitical, taking the ones with the biggest gap between buy and sell — and it’s moving the market at all, even a little, that’s already the claim that an AI system is outperforming an entire crowd of humans working together to produce a price. To stress again: none of these three methods is very credible on its own — they’re all fluctuating wildly and have methodological issues. So we don’t exactly know, but at least on the surface there are profitable AI traders forecasting, and AI systems beating the two most reputable forecasting benchmarks people trust.
Prakash: I’ll ask probably a silly question, but I have to. The story with DeepSeek was that they originally wanted to simulate what consumers would do — they built systems to ask and simulate the market, so to speak, useful for the hedge fund side of the business. That might be an urban legend, but I have to ask: you must get approached by financial firms all the time — has it ever been successful, using models to simulate participants in a market in order to figure out how things will play out?
Dan Schwarz: I’ve only heard about people trying to simulate entire markets. People occasionally approach me with a startup idea to build an AI prediction market where all the AI entrants bet against each other, and I usually tell them that’s a very inefficient way to produce the best forecast. The reason that works for humans is that humans need incentives — if you’re going to research and make a bet, you have to think it’s profitable in expectation. AIs don’t need incentives; you can just ask one to do something, and it’ll do it as well or as badly regardless of any financial incentive behind it. I think the deeper question is what the hedge funds have been up to this whole time — we know they’ve hired superforecasters, we know they all know about Tetlock, so what are they doing, and is it making them money? I have no idea. Over the last three years at FutureSearch I’ve spoken to several hedge funds and gotten a little insight into what they’re doing, and all I can say is it’s quite abstruse. Every hedge fund is different from every other, and they’re all extremely secretive, so we’re basically not learning anything about this. If somebody figured out how to make money this way, I don’t think we’d find out about it.
21:42-28:34
Nathan Labenz: There wouldn’t be any signs. It’s an interesting thought experiment — how would we know if somebody did have a truly differentiated superforecaster? It sounds like your expectation is that individual hedge funds do well all the time, and there’s not really much we could monitor for that would tell us. Is that right?
Dan Schwarz: I’ll give you an example of how little you’d see this in basic indicators. I spoke with a hedge fund a few years ago, which I’ll keep anonymous — their forecasting play was taking breaking news, like a drone strike or an oil embargo, and trying to trade Bitcoin volatility in the first 300 seconds after the news broke, on the theory that whenever something risky happens, Bitcoin volatility goes up. If you could figure out, faster than everyone else, whether a breaking story was actually disruptive or just another day with nothing interesting going on, you could go long volatility and exit, or go short volatility and exit, within five minutes of the news breaking. If people were doing that, would you ever notice by looking at Bitcoin volatility? I don’t think we would. You could ask an even more basic question: how well are LLMs doing in finance? Do we have any idea how useful they are there? We all assume they’re being used heavily, but do we have evidence — is it profitable, are people publishing papers on it? I think finance is just a poor place to look for evidence of how technology is diffusing into the workplace, because there are such strong incentives to keep it secret. Whereas with coding agents, we’re getting tons of evidence, because the way to make money there is to massively distribute them. So I think we have to look where the evidence actually is, and that’s part of why FutureSearch is a public product you can go try right now.
Nathan Labenz: What can you tell us about how you’re building the forecasting agent? In the coding world we have different models critiquing each other — the general sense is that models from different providers have different failure modes, so they might catch each other’s mistakes, and you get rewarded for having models from different lineages review one another’s work. There’s also the general idea that spending more tokens should get you further. Are you basically following the same trends driving coding agents more broadly, or have you found particular things that buck those trends or might surprise people?
Dan Schwarz: We’re generally following those trends, but I don’t think those trends are very well understood, so let me try to answer in a way that’s informative. The very first thing we figured out building a forecaster in late 2023 with GPT-4, and then GPT-4 Turbo — the main way to get a bad forecast is to get bad information about the present. We were throwing more tokens at web research to try to understand what was actually going on, say with the Ukraine war or COVID, and we’d get stuck there. The research agents weren’t even really agents for a while — they were ‘for’ loops, or what people used to call prompt-chaining. You could try fine-tuning models — there were various techniques available in 2023 and early 2024 — but any human expert forecaster looking at the output would say, you clearly have the facts on the ground wrong, so there’s no way you’re going to get a good forecast out of this. So we started studying what later became known as deep research, and we created something called Deep Research Bench, because you have to be able to evaluate good deep research on the present before you have any chance of a good forecast. One of the things that’s really changed in the last twelve months is that getting the facts on the ground right is becoming more of a commodity — though it’s still not fully a commodity; anyone who’s used deep research recently in a field where they’re an expert will still see that getting the basic facts right is beyond the frontier. More than 80% of the work of forecasting, I’d say, is just doing good research and reasoning on verifiable things about the present.
Once you have that, there are a number of techniques to get a good forecast out of it. FutureSearch is a seed-stage startup, so we’re not training $50 million pretrained models — we’ve got fine-tuning, reinforcement learning, agents and multi-agent scaffolds, the same techniques anyone downstream of the frontier labs has. Some of the most interesting papers on this are coming out of the Foresight LLM project from Lightning Rod Labs — they’ve actually been doing reinforcement learning; there’s a link in a comment on Scott’s post to another group doing similar work. We’re starting to learn a bit about what these techniques are, and people are publishing results. FutureSearch put out a paper in April about strategic reasoning failures in frontier forecasters — we looked at our best forecaster, which is proprietary and which we don’t share the internals of, and compared it to a very good Opus 4.6 research agent, pointing out stylistically what the differences were, what the Opus 4.6 agent was getting wrong. The TLDR is that frontier models still struggle to model human dynamics as well as an expert human, or a better-scaffolded forecaster, can. We’d find examples where a human forecaster, asked whether a bill would pass, would say: this president’s top priority is getting this bill passed, otherwise they won’t be able to go to this conference, and they need this as a face-saving mechanism — and the model would completely miss that, wouldn’t understand the deeper motivations behind a politician.
A lot of forecasting is geopolitical forecasting, and a lot of geopolitical forecasting is figuring out the true incentives facing all the players in an elaborate game-theoretic game. That gives you a sense of the frontier on the research side — if you can’t understand those things about the present, that’s where you’ll get stuck in forecasting. Everything about turning the best possible information about the present into a good forecast about the future is mostly FutureSearch’s proprietary edge, so we keep that under wraps. But you can see research capabilities improving for everyone — just go to the free tier of Gemini, ChatGPT, or Claude, and you’re already getting research dramatically better than a year ago. Pay $20 a month and you’re getting close to expert-human level, though still not quite there, for researching something like an economic or technology situation.
28:35-32:01
Prakash: Let’s say AI forecasting gets really good — how does a consumer use that in their life? How do you actually make use of this? A geopolitical question, like whether Iran goes to war or not, isn’t something the average consumer has much use for. So how would this work out in a consumer product?
Dan Schwarz: It’s a great question, and it’s been a central paradox of forecasting since Tetlock started studying it, or at least since I got turned onto prediction markets lurking on LessWrong forums in 2011. Implicitly, everything we do is based on conditional forecasting — every decision we make carries some implicit forecast about causal outcomes conditioned on whether we take certain actions. Every decision theory ever put in an economics paper has this idea that people model the world, and there’s some association between the map and the territory. Predicting those outcomes is kind of what intelligence really is — yet if you handed people an oracle, they’d struggle to use it. How do you explain that? With the relatively poor-accuracy AI forecaster we built in late 2023, I put it in front of dozens and dozens of people — smart consumers, but also middle and senior managers at all kinds of businesses — and I was shocked at how they didn’t seem to have the imagination to figure out what they should even be forecasting in the first place, despite spending all day making decisions that were implicit forecasts about the outcomes of their actions or their teams’. That’s part of my intuition for why forecasting is just a very neglected capability — people don’t know how to use it. Compare that to Midjourney: ‘I’ve got this AI, I can generate amazing images, how should I use it?’ People struggle with that too, but it’s a lot easier — you can immediately see five or ten things you’d do with it, personally or professionally, for fun or to make money. Forecasting still requires some imagination — it requires introspecting about what decisions you’re even making right now. As the CEO of a seed-stage startup, I’ve really benefited from doing my own forecasting, having AIs look at my decisions and reason about what the outcomes would be conditioned on me doing this or that. But even I, having studied this for a long time, have to put in effort to figure out what decisions I’m even making today, what outcomes I’m tracking. One of the most basic things people get stuck on is how you even measure the outcome — forecasting generally requires stipulating something measurable. Are you optimizing for dollars, users, engagement? How would you even know if something improved your relationship with somebody?
32:02-35:22
Nathan Labenz: Going back to how the frontier companies are thinking about forecasting — I know you can’t share how they actually think about it, but could you give us your thoughts on how they should be thinking about it? Is it complicated? The naive take would be, it’d probably help if models are good at forecasting, so labs should try to make them accurate forecasters. Is there a less obvious level of nuance to that idea that you’d advise AI companies take seriously — whether or not they already are? What should they be doing, from your perspective?
Dan Schwarz: There are really two questions here: what should labs do with forecasting as a capability, and what should they do with forecasting as an eval? Forecasting as a capability is kind of a business decision — does OpenAI care whether ChatGPT is a good forecaster? That depends on whether their consumers care about it. If you’re Anthropic, you probably care more about the enterprise case — when people use Claude for white-collar work, do they care how good it is as a forecaster, are people trying to use it to make financial forecasts in an Excel spreadsheet? That’s a business decision, and I can’t really weigh in on it. People will discover over time just how important forecasting is across everything, but it’ll be a slow process for humans to notice.
From the eval side, it’s very different. Forecasting has this beautiful property: you get ground truth just by waiting. If I ask a question about the future that’s basically impossibly hard — a question even an AGI, an oracle, a god, could never really answer because of chaos theory, like predicting the exact weather in a cubic meter three weeks out — you’d never be able to answer it in advance, but if you just wait, you’ll see what that weather actually was. So you have a completely limitless set of extremely hard, essentially impossible questions where you eventually get exact ground truth, and there’s no other eval like that. If you want to improve a coding harness, you need more and more hard coding problems, not in the training data, where you can say ‘this is definitely the correct answer’ so you can train on it — and that’s hard. Human experts — doctors, lawyers, engineers, financiers, whoever — trying to build evals for the frontier labs are finding they’re not smarter than the models being trained anymore. If you can produce something with a correct answer, the model’s probably already going to figure it out; you need something with a correct answer the model can’t figure out. Forecasting, I think, is the only completely renewable source of that. This connects to the Elon Musk-style quip that forecasting is the ultimate measure of intelligence — zoomed out, I think that’s basically right. I wouldn’t say forecasting is truly ultimate intelligence — coding intelligence, AI R&D intelligence, interpersonal intelligence all matter enormously too — but it is, to some degree, the ultimate eval, and I think that’s something the frontier labs should be paying attention to.
35:23-39:00
Nathan Labenz: Do you have a sense of how much irreducible chaos there is? In the limit of superintelligence, how much visibility do we get into the future?
Dan Schwarz: That’s a great question. David Manheim actually just posted a tweet with a graph estimating this that I’ve been meaning to dig into, because I don’t know where those lines are coming from. Without evidence behind it, my sense is very Yudkowsky-ish — I think there’s a lot of detail in reality that’s far beyond the human mind to understand, and as you approach more sophisticated intelligence you’ll start seeing a lot more patterns. Producing voxel-perfect weather three weeks out is further away than people think, but I think there’s quite a lot of room. Human superforecasters don’t tend to agree with me on this — they think what they’re doing is somewhat near-optimal, and any accuracy improvement you’ll get over them will be tiny and hard to understand. I think that’s just because we only really understand human intelligence. Zoomed out, from an information-theory or Kolmogorov-complexity perspective — modeling the world as byte strings — the AIs will eventually figure out things that are totally beyond humans to notice. There’s no way to prove this, but my sense is we’ll start seeing it over the next year, as AIs get more and more accurate compared to humans in ways humans don’t really understand. You’ll look at the rationale for a forecast — five paragraphs of dense reasoning and then a surprising conclusion — and it won’t really make sense, but it’ll turn out to be really accurate, and we’ll understand it less and less as time goes on.
Prakash: Mhmm.
Nathan Labenz: I think that’s related to other modalities. It’s something I’ve been obsessed with for probably the last two years — the fact that models are capable of learning a sort of intuitive physics in a lot of different spaces, from protein folding to materials science, like what the band gap will be for a particular semiconductor recipe. There’s this strange, alien nature to these problem spaces, far from human sensory perception, that models trained on raw data — or in many cases simulation data — just seem to overcome. I’m wondering if you think that’s a similar phenomenon, where they’re pulling the vacuum on reality so tight that they just naturally get it, in a way that maybe isn’t even language-mediated — or whether you think the chain of thought is still where the action is. When you say you see these five paragraphs of reasoning and then a surprising conclusion, it starts to call into question for me whether the chain of thought is really where the action is, or whether it’s a facade — a superficial projection of some deeper, sub-language understanding or reasoning that might be going on. And obviously, if so, that potentially has a lot of consequences for the safety plans of the frontier companies, which seem to revolve mostly around monitoring the chain of thought. Meandering question, but hopefully you get where I’m trying to go with that.
39:01-43:03
Dan Schwarz: In some trivial sense, that’s definitely the case — if you simply ask a human superforecaster to explain their reasoning, they can’t actually make it fully legible. There’s a layer of intuitive judgment that feels like deep learning: they look at a bunch of evidence the way a chess grandmaster looks at a position and just sees the right move, and they can’t explain it — it just popped into their head, the way the grandmaster moves the knight and it lands on the right square somehow. That happens with humans already, and it happens with AI superforecasting systems today. So there’s no a priori reason to think that reasoning would always be legible — there’s going to be some layer of intuitive judgment, to the extent ‘intuitive judgment’ refers to something going on inside a large language model. It just has to be that way.
Whether it’s very that way or only a little that way is really your question, Nathan — if I read the reasoning traces, the rationales, the research it did, is it more or less what a human would have done, where I can see where it’s coming from? Or is it inscrutable, in the way that it discovers some new pattern in the world nobody’s ever seen before — the kind of ‘psychohistory’ from Asimov, where there’s really an underlying structure, like some game theory of international politics, that it discovers from being trained on forecasting questions but can’t really elucidate? Even in that case, it could probably write a program saying, ‘politics follows these game-theoretic properties from some obscure textbook nobody’s been thinking about, and if you use this you get incredibly high accuracy predicting these outcomes’ — and we could look at that and say, that makes perfect sense. The real question is: what’s the level at which it’s doing something we cannot follow down the dark forest of its reasoning? Almost by definition, we can’t know what that would look like.
I’d say the best evidence will come from other domains. To the extent the AI 2027 scenario is starting to happen, where AIs get better at AI research, is that legible? Will the brilliant AI researcher at the company look at it and say, ‘oh, that’s such a good idea, why didn’t I think of that’ — clearly helpful and smart, but something a human could have done on a good day? Or will it just write programs full of weird floating-point numbers nobody can understand anymore, and then just go off to the races? We should look at all the domains where something like that is happening. A very salient example is clinical trial prediction — a lot of people want to know how well various drugs will do in humans, especially as AI produces more drug candidates and we can synthesize more compounds, and the bottleneck is safely running hundreds of thousands of people through trials and measuring the effects. If you could predict drug outcomes even modestly better, that would be worth many billions of dollars. What does it look like to predict drug trials — does it look like a human superforecaster reading the relevant papers and vibing a probability or a numeric range? Does it look like building an elaborate software model? Does it look like quantum chemistry, running enormous numbers of protein-folding simulations? What does it actually look like? If it looks like the human version, fine — but maybe something trained on that task discovers a completely different way of predicting clinical trials that all the world’s biologists and pharma executives are completely blind to.
Prakash: Maybe I’ll ask one last question, since we’ve held you past the time — where do you see things in 2029? We’ve had these six months of crossing the Rubicon. Give us an unhedged prediction, looking forward two and a half to three years — where do you think things end up on the forecasting front in 2029?
43:04-46:59
Dan Schwarz: FutureSearch contributed some forecasts to AI 2027, and we studied that problem pretty seriously with the evidence available a bit over a year ago. We built a model of R&D takeoff speeds under the core AI 2027 scenario, where the main way things get crazy is AI being used more in the development of AI — first hitting the superhuman-coder milestone, then the superhuman-AI-researcher milestone. I’m unhappy to report that I think that story is generally correct. I don’t know if the timelines are exactly right, but my forecast from that process, leading to something that looks like superintelligence around 2031, is roughly stable.
The things that have happened in the year since AI 2027 came out really vindicate the theory that the most important thing going on is how useful AI is at improving the productivity of AI researchers within the frontier labs. I publicly predicted that Anthropic was going to run away with it, because they had the best feedback loop of talent actually using their own AI internally — that’s n-equals-one, but I think it’s been pretty well shown to be happening. So the question is what role forecasting plays in that. You can imagine what was going on with Dario Amodei and Tom Brown going to the White House trying to get Claude Fable unbanned — what moves did they have available for dealing with the US government? Even if you’re incredibly cynical and assume the government is just trying to score political points, what are those points, and what moves does Anthropic have to work with? There’s a whole set of intergovernmental negotiations going on, and to what extent is Anthropic, a company working on AI, able to use that AI to negotiate a better outcome with politicians? That sounds like a conditional forecasting question to me — if I write this letter, if I propose this export control, if I weaken the model, what’s the probability this person relents within the next three days? We don’t have evidence that’s happening — there may not be any evidence of it at all — but I think in a 2029 model, whether AI forecasting is helping the frontier labs make better decisions more effectively is very similar to the way improved coding is helping them develop AI systems faster and more productively. That’s the thing I’d watch.
Nathan Labenz: I’ve got to assume they’re at least talking to Claude a bit along the way.
Dan Schwarz: I’m super curious — Anthropic senior leadership, if you’re watching this and want to comment on how useful Claude Fable was in the Claude Fable situation, I’d love to hear it.
Nathan Labenz: They may not have been able to use it either. Prakash, I have one other question, but go ahead if you’ve got one. Two more practical questions — first, what does it cost to run a forecast if I go in as a consumer? Second, if I had a few million dollars to invest in creating a public good, how would you think about building a public good with forecasting? We might say, what’d be really great is an amazing voter guide to the midterms — this person gets in, that person gets in, here’s what you can expect. Could we help the public make better decisions by giving them visibility into the consequences of the votes they might cast? That’s probably a pretty limited idea, though — I’m interested in farther-out ideas for public goods somebody might be able to provide.
47:00-51:33
Dan Schwarz: Glad you asked that, Nathan.
Nathan Labenz: Go for it. I’ve got one more too, actually — no, that’s enough for now, I’ll follow up with the other one.
Dan Schwarz: It costs about a dollar or two to make a frontier forecast — that number can get a lot higher or a little lower, but that’s a good anchor. If you look at the cost per input and output token for an LLM, that gives you a rough sense of the amount of research being done. One of the core questions FutureSearch tackled — I described earlier how our main frontier for a while was just doing good present-day research, until we got good enough at that to use it to improve forecasting — was: can you just pour more tokens into a question to get a more accurate answer? It doesn’t have to be a forecasting question — if I ask you the current state of some clinical trial and want the most accurate answer you can give, can I just pour more tokens into that and get a better answer? This was studied as deep research — writing these fifteen-page reports with 700 citations. That gave you a longer answer, but was it a better answer? It wasn’t super clear, which is why we studied it.
Forecasting gives us an opportunity to do some world-modeling. FutureSearch talked about this a bit at the Manifest conference a couple of weeks ago, and the feature is rolling out in the product, I think literally today. The idea is that once you have a repository of forecasts, every marginal forecast can draw on the implicit world model in those forecasts to give a better answer. FutureSearch co-founder Lawrence Phillips wrote this up on LessWrong a couple of months ago, and it was a bit neglected — he made the case that if you produce a large body of forecasting questions that feed into each other and remain mutually consistent, you could understand the world dramatically better, as a public good. The main barrier is simply: when you put more tokens into your world model, does it actually get better, or does it get worse? His big insight was that around January or February, around Opus 4.6 or GPT-5.4, for the first time it became possible to put more tokens into a broad research task and actually get a better answer, not one that just trails off into garbage-in-garbage-out nonsense. FutureSearch is building this into its products, and that’s part of why we have a consumer product — the more people forecast, the better the forecasts get for them, and in theory the better they get for everybody as we build this deeper implicit model of the world.
Lots of companies and research labs have had ideas about building world models — though the term ‘world model’ as I’m using it might be a bit misleading; a lot of people mean something like geospatial reasoning, building a robot hand that can pick something up. That’s a world model too, just a basic one for predicting outcomes. The difference with forecasting is that the rationale behind a good forecast is a very compressed, brilliant five-paragraph summary — the things most causally relevant to a hard forecast question are the best summary of information you could reasonably produce, because every consideration in there leads to a different date, number, or probability. When you start recording five paragraphs of the most important considerations for the most important questions, and scale that to thousands or tens of thousands of questions, and then read through them asking — what are the latent variables here, what are the implicit claims, why do we think it’s this way, are these consistent with each other — you start to learn a lot, in a way that’s different from just pretraining on all of Wikipedia, which fills you up with random, uncorrelated facts. With a repository of forecasting research, you might learn things about the world you could never have seen by asking one question and getting one answer one time.
More broadly, the big question is: can you pour more tokens into any kind of research and get better research out of it — AI research, coding, whatever? You’ve burned through your Fable tokens for the day, you top up, and can you just let it cook — let it keep cooking better and better as you put more tokens in, or not? From a forecasting perspective: could you get to the point where every marginal new forecast improves the accuracy of every other forecast, until you’re building a consistent world model? That’s what FutureSearch is after right now. We have some preliminary results indicating this does work, and we’re rolling it out to our users right now.
51:34-56:43
Nathan Labenz: What structure does that ultimately get instantiated into — are we talking about a graph database? I can see those kinds of ideas making sense here, but I can also imagine them introducing some weird failure modes. Here’s a fun fact about me: I was actually on the Good Judgment team, way back in the DARPA forecasting tournament — or was it IARPA that funded it — fifteen-plus years ago. I did well, but not top-tier superforecaster level. Around the same time, I also worked briefly at a financial-services consulting firm that had done a lot of the financial risk modeling for Fannie Mae — I probably don’t have to tell you how that story turned out. There was a lot of expert forecasting instantiated in a very spreadsheet-driven, causal-graph kind of way — you could literally hit the visualization button in Excel and see these colored arrows fanning out from cell to cell — and somehow, in the end, it was all totally off. So I wonder how you think about correlated failures as you build out these world models. Is there any kind of correction mechanism — something to catch an assumption like ‘housing prices never go down nationwide’ lurking in the world model? Obviously humans have this problem too, the financial crisis proves that, but you can imagine the next one being even worse, because we’re so reliant on a small set of AI minds working the problem from a thousand different directions that might all share somewhat consistent flaws in their reasoning as they go. Can we protect ourselves against that in any way?
Dan Schwarz: Definitely — I’ll try to answer that both theoretically and with an anecdote. I tried to world-model the Fable situation when it got banned, partly because I wanted to and partly because it was a good forecasting question with some real money trading on Polymarket. And I made exactly the mistake you’re describing, Nathan. I ran a bunch of FutureSearch forecasts and manually went through them — a couple of scenarios, some conditional forecasts, basically thirty-three load-bearing sub-forecasts, starting from what even happened: why did the government issue this export control? Was it a simple misunderstanding? Political leverage? A genuine foreign-threat concern, or hacking risk? We didn’t know, so I put it all together. When I looked at all the outcomes — and I talked about it with Claude Code a lot — one thing came out: basically every forecast, in every scenario, assumed access would come to Americans first and foreigners only later. That was wrong — when it came out last week, it came back for everybody at once. So clearly there was some faulty assumption baked into one of my scenarios — a correlated failure somewhere. I still haven’t fully figured out where my reasoning went wrong; it’s also possible I just got unlucky and we landed in an unlikely outcome. On an n of one, you can never really know if any single forecast was good — that’s one of the hard things about this. But I think I systematically got it wrong by having a bunch of correlated reasoning failures across my scenarios. This definitely does happen.
Metaculus has a system for this — in the years since I was CTO there, they’ve built an actual causal-graph platform and product; you can go to the Metaculus site, click run, and find it there. I think the field still generally believes something like this will work, but nobody’s made a really good one yet. I tried my best over about twelve to sixteen hours on the Fable situation — I think I built a pretty good model, close to a very accurate forecast, but didn’t quite get there. I don’t think the Metaculus models on their site right now are amazing either, but I do fundamentally believe in the approach. As you’re saying, Nathan, this has been tried for a long time — when I was CTO of Metaculus, honestly, it was the dream, the holy grail: can we tie all these forecasts together into some kind of causal graph? What I can say is that AI makes this tractable. There was just no way that was going to work with a bunch of human economists looking at Freddie Mac or Fannie Mae — I can totally understand why that method didn’t work for them then. Whether AI can make it work right now is unclear, but whether AI will make it work in general feels nearly guaranteed, and I don’t think FutureSearch is the only org working on this right now.
56:44-1:01:14
Nathan Labenz: Well, you’ve been super generous with your time — we only had half an hour on the calendar, and we’ve been at it nearly an hour.
Dan Schwarz: Have to go see — right behind me, it got darker and darker as the conversation went on. Did you—
Nathan Labenz: —guys notice that? I need to— my lighting is a constant source of, let’s say, comedic relief for us if nothing else, so you’re in good company there. Maybe just in closing — sketch out a bit more of the future as you hope it might unfold. Not necessarily the most likely scenario, because maybe the most likely thing is people act foolishly and don’t take advantage of the benefits of forecasting. But if we really do a good job — we’re interested in truth-seeking, and we get the AIs working as well as you think they might — how do you think life kind of feels different?
Dan Schwarz: I have to lead with another example of me being a bad forecaster — I guess everyone who tries forecasting thinks they’re bad at it, because they see themselves getting things wrong. A prediction I made really strongly five or ten years ago has basically been totally falsified: I predicted that if we had highly visible, highly liquid prediction markets covering all the major technological, political, and economic things going on, humanity would be wiser and people would make better decisions in government. Well, here we are — we have Polymarket and Kalshi, and I don’t see any wisdom or better decisions coming out of all the gambling happening on those platforms. So part of thinking about our AI future, for me, is trying to understand the present a little better — why isn’t having thriving prediction markets transforming the news, or how people learn information, or how they plan for their futures? One simple answer is that it does, it just takes a while — we’re only about a year into prediction markets making major headlines and being seen by everybody, and maybe it just takes time for people to change their habits. AIs, if that’s the case, can move much faster as they get better at forecasting.
Ultimately — and you said this, Nathan — we’re after the epistemics. It’s not just about predicting an outcome, we want models that are reasonable. One of the beautiful things about forecasting as a human practice is that it makes you more epistemically virtuous: the more you try to forecast, write down what you got wrong, and do postmortems, the more it humbles you and makes you more open-minded — more of a fox than a hedgehog, just a more reasonable person. Prediction markets, with everyone doing this, should be leading to people being more reasonable — except I don’t think people are doing a lot of actual forecasting on prediction markets, they’re doing a lot of trading and gambling, which is related to forecasting but isn’t forecasting. If AIs get better at forecasting and become more epistemically virtuous, we could be in a world where just talking to a chatbot gets you something so much wiser and more grounded, more honest about its own uncertainty, and better at probing you about your own uncertainties. That could make an enormous difference. But putting my cold-blooded forecasting hat back on, I think the technological outcomes of AGI will arrive before that cultural change happens.
So I’m very much in the AI-safety camp — I really think we should slow things down, give ourselves more time, fund more AI safety research, and do more policy work. Because if we have time for the wisdom of these alien intelligences to actually help us make better decisions before the critical decisions get made — there’s a series of decisions coming in the 21st century that we’ll look back on the way we look back on 20th-century decisions about communism, World War Two, the atom bomb. Those decisions are coming, maybe some have already been made, and right now I don’t think they’re well informed by rigorously accurate forecasting AIs. But give it another couple of years, and we might be in a world where everybody has the same grounding — as smart as Kissinger, but actually trying to help, trying to give us all better outcomes. That could usher us through this crazy phase before the paperclip-type stuff starts to happen. So I feel like I’m racing to make AI forecasting useful and helpful — it’s part of a broader epistemics-and-safety process, because otherwise it’s just going to get away from all of us, and a lot of the work we’re doing won’t matter.
1:01:15-1:02:00
Nathan Labenz: Cool — well, as my dad would say, good luck. We’re all counting on you.
Dan Schwarz: Thank you, Nathan. Thank you guys for having me — I really appreciate it.
Prakash: Thank you, Dan.
Nathan Labenz: Dan Schwarz from FutureSearch — thank you for being with us on AI in the AM.
Made with: The Transcript Desk Chrome Extension

