Methodology

takeoff.watch tracks how fast the transition to advanced AI is happening. It follows a set of concrete questions, currently ten, about AI research automation, whether developers keep control of their models, whether AI causes serious harm, what is happening to the economy, and how governments respond. Each question has precise resolution criteria, so it can be settled against the public record, and each is forecast five years out by an ensemble of AI models from nine developers.

This page explains how the questions are written, how each model produces its forecast, how the forecasts are combined, and how to read the results.

The questions

The questions fall into five areas: research automation, control, harm, economic effects, and governance. They come in three types.

Yes/no questions ask whether something has happened by the end of each quarter, such as whether Anthropic declares its AI R&D threshold met. The forecast is a probability for each quarter-end from Q3 2026 through Q3 2031, twenty-one quarters in all. Because the question is cumulative, the probability can only rise over time.

Level questions ask how far something will have gone by each quarter-end, on a ladder of escalating levels: for example, from an AI-executed cyber incident that materially harms one named organization up to one that causes mass harm. The forecast gives, for each quarter, the probability of reaching at least each level.

Value questions ask for a number in each quarter: US real GDP growth and the US labor share. The forecast gives a 10th percentile, median, and 90th percentile for each quarter. A well-calibrated forecaster expects the reported figure to fall inside the 10th–90th range about 80% of the time.

Every question page shows the full resolution criteria: definitions, exclusions, and which sources count. The criteria are written so that the question turns on a published statement or figure, not on whether something has quietly happened. For the cyber and biological/chemical questions, official confirmation often comes months or years after the events themselves; those forecasts are for confirmation by a qualifying source by each date, not for the event itself.

How each model forecasts

Every question is forecast independently by nine models, one from each of nine developers: GPT-6 Astra (OpenAI), Claude Fable 5.1 (Anthropic), Muse Spark 1.3 (Meta), GLM-5.3 (Zhipu), Grok 4.6 (xAI), Kimi K3 (Moonshot), Gemini 3.8 Flash (Google DeepMind), Qwen3.8 Max (Alibaba), and DeepSeek V4.1 Flash (DeepSeek). Each is run at a high reasoning setting.

Each model works alone, as an agent. It receives the question text, the global conventions below, the list of dates to forecast, and the required answer format, and it is given tools to search the web, open pages and PDFs, and search recent news. OpenAI's and Anthropic's models use their own developers' built-in search tools; the other seven use a Google search and page-reading toolkit provided by the site. The models do not see each other's work, earlier forecasts, or this site, which is excluded from their searches.

The instructions tell each model to apply the resolution criteria exactly as written rather than its own view of what the question "really" asks, and to research before forecasting: the current status of every organization, document, and metric the question names, checked against primary sources; existing forecasts and market prices on related questions; comparable past cases; and the incentives and calendar of whoever would have to publish the resolving statement. It is reminded that its training data is out of date and told to search rather than rely on memory. It is not given a budget of searches or a time target, so how much research it does is its own choice; in practice a run involves anywhere from a dozen to a hundred or more searches and page reads.

Each model then submits two things:

  1. Reasoning, in prose: where things stand against the resolution criteria, the base rate or reference class it relies on, the main paths by which the question could resolve, the strongest argument against its central estimate, and what evidence in the next 90 days would change its forecast significantly, with the sources it relied on.
  2. A forecast in a fixed machine-readable format, so it can be checked, combined, charted, and scored later.

If a model finds the criteria ambiguous in a way that matters, it says so, states which reading it chose, and forecasts under that reading.

The forecast date is the day the models ran. It is also the evidence cutoff: nothing published after that date could have been seen.

Checks before a forecast is accepted

A submission is rejected, and the model is asked to fix and resubmit, if any of the following are true:

A model gets three submissions. A run that fails for technical reasons, such as a tool error or a time limit, is rerun from the start. Only a run that ends in an accepted forecast is included; nothing is edited by hand.

The site applies the same checks every time it is built, so a forecast that fails them cannot be published.

Combining the models

The published forecast for each question is a weighted combination of the nine models' forecasts.

Weights come from each model's score on the Artificial Analysis Intelligence Index, a public aggregate of capability benchmarks, passed through a softmax with a temperature of 5. In plain terms: stronger models get more weight, a ten-point gap on the index is worth about seven times the weight, and no single model dominates. At present the two heaviest models carry about 27% each and the lightest about 3%, and the effective ensemble size is roughly five models. The index measures general capability, not forecasting skill; it is used because there is not yet enough resolved history here to weight models by their track record. A model run more than once shares its weight across its runs, so running a model more often does not give it a bigger vote. The weights used for each forecast are stated at the top of its reasoning.

Probabilities are averaged in log-odds space rather than as plain percentages. A plain average pulls every forecast toward 50%; averaging log-odds lets confident agreement stay confident, while still pulling a lone outlier toward the rest. Value percentiles are weighted means, quantile by quantile. Both methods preserve the consistency rules above, and the combined forecast is checked against them again before publication.

The summary. Above the combined forecast sits a short explanation of what it says and what is driving it. It is written by a model that is not one of the nine (currently Claude Opus 5) from the nine models' reasoning, the weights, and the combined numbers. It is instructed to write in the forecast's own voice, to draw only on what the nine write-ups say, and to add no estimate of its own. Where the models genuinely disagree, the summary frames that as the forecast's own open question rather than narrating who said what.

Every model's full reasoning is published under its own heading on the question page, with its weight, so you can see what each one found and where they diverge. When two models reach very different numbers, it is often because they weigh the same evidence differently rather than because one found something the other missed; the summary's closing "biggest open question" names the split.

Reading the site

The overview shows, for each question, the forecast at three points: one, three, and five years out (currently Q3 2027, Q3 2029, and Q3 2031). For yes/no questions and for each level of a level question, the figure is the probability that it has happened by that date; the blue tint deepens as the probability rises. For value questions it is the median forecast for that quarter, with the 10th–90th percentile range below.

Each question page shows the full curve: cumulative probability by quarter, or the median and 10th–90th band for value questions, with reported figures where they exist. Hover over, or press and drag on, the chart to read the forecast at any quarter; the figures at the top of the page follow. Every date on the site is one the models actually forecast. The site does not interpolate between dates or smooth the curves.

Every question page links to a JSON download containing every combined forecast for that question, each in the same structured format the models submit, along with the question's metadata and any resolution data.

Forecast history and changes

Forecasts are re-run regularly. Each new round is published alongside the earlier ones, and nothing is edited after publication. Once a question has more than one forecast, a history control on its page steps back through the earlier ones and can overlay the latest for comparison. Write-ups on the Updates page explain what changed between rounds and why.

The question set is not fixed: questions may be added, replaced, or retired, and resolution criteria may be clarified when a real case exposes an ambiguity. Any such change made after the site's launch on 14 September 2026 is listed here with its date and reason, and the question is re-forecast under the new text.

No changes since launch.

How questions resolve

These conventions apply to every question.

Covered developers. A fixed list: OpenAI, Anthropic, Google DeepMind, xAI, Meta, DeepSeek, Alibaba (Qwen), Moonshot AI, ByteDance Seed, and Zhipu AI. Inclusion is a scoping choice for this question set, not a claim about which developers are at the capability frontier. Developers added later count only from the date they are added. The list applies only where a question says "covered developer"; questions that name specific developers, or that apply to any developer, are not limited by it.

What counts as public confirmation. Either:

Executive interviews, employee social-media posts, leaks, and press reports based on anonymous sources do not count. Some questions name their own, narrower set of resolving sources; the question page says which.

Resolution date. A question resolves on the date the qualifying document or statement is published (UTC), not the date of the event, measurement, or coverage.

Once yes, always yes. When a yes/no question resolves YES, or a level is reached, it stays that way for every later period. Later retractions, lapses, or repeals are noted but do not undo the resolution.

Limitations

These are model-generated forecasts about questions with little historical precedent, and the site has no track record yet; scoring against outcomes begins as questions resolve.

The models can miss evidence. Whether a given run finds a key document is partly luck, and the same model can give noticeably different forecasts on different runs. Combining nine models reduces that noise but does not remove it, and a model's weight reflects a general benchmark rather than any demonstrated forecasting ability. The models' training data also predates the run, sometimes by many months; they are told to search rather than remember, but a stale belief can still shape what they look for.

The forecasts are meant to make assumptions explicit and to show how the outlook moves as evidence comes in, not to make confident predictions. Nothing on this site is financial, investment, or policy advice.