US–China frontier AI agreement

How far will a signed US–China agreement on frontier AI go?

Chance by Q3 2031

Use prohibitions
53%
Testing & results exchange
27%
Development constraints
12%

Cumulative probability

0%25%50%75%100%
  • ≥ L1 · Use prohibitions
  • ≥ L2 · Testing & results exchange
  • ≥ L3 · Development constraints

Model reasoning

Aggregate of 9 independent forecasts made 2026-09-14: weights from a softmax over each model's Artificial Analysis Intelligence Index score, probabilities combined in log-odds. Weights: GPT-6 Astra (OpenAI) 27%, Claude Fable 5.1 (Anthropic) 27%, Muse Spark 1.3 (Meta) 15%, GLM-5.3 (Zhipu) 8%, Grok 4.6 (xAI) 7%, Kimi K3 (Moonshot) 7%, Gemini 3.8 Flash (Google DeepMind) 4%, Qwen3.8 Max (Alibaba) 3%, DeepSeek V4.1 Flash (DeepSeek) 3%. Each model's own reasoning follows.

Summary of the ensemble forecast, written by Claude Opus 5 from the 9 models' reasoning.

This question asks how far the US and China get on paper about frontier AI — not talks, not one-sided announcements, but a written document both governments actually put their names to. Three rungs: banning specific AI uses (L1), committing to pre-release safety testing with some exchange of results (L2), and constraining the pace or scale of development itself, including training-run notification (L3). Right now none of these exists. By the end of September 2026 we put L1 at 8%, L2 at 2%, L3 at essentially zero. Five years out, by September 2031, the numbers are 53%, 27% and 12%.

The setting matters. A Trump–Xi state visit lands on September 24, with AI on the agenda, and a Treasury-led AI safety dialogue has been floated for mid-September — though the White House has at points denied one was scheduled, and reporting suggests the AI track may simply be folded into the economic meeting. The May 2026 Beijing summit produced no joint statement at all, only a US fact sheet and a Chinese confirmation that the two sides would talk. That format problem is central: this question needs co-issued text, and recent practice has been duelling readouts. Meanwhile both capitals are talking past each other — Bessent says nobody can pause because China won't, Trump says whoever wins AI wins, and Beijing calls slowdown proposals a containment playbook.

So why not near zero? Because L1 is cheap. A line saying neither side will use AI to help build biological weapons or attack civilian infrastructure costs almost nothing, is already implied by existing treaties, and is exactly the kind of deliverable summits produce — the 2015 Obama–Xi cyber-theft commitment is the template. Over five years and a dozen leader meetings, plus a likely change of US administration in 2029, that gets a real chance. L2 and L3 are different animals: swapping pre-release test results with a strategic rival has no precedent anywhere, and pacing constraints run straight into an unsolved verification problem that neither side has proposed a fix for.

The biggest open question is whether an AI incident severe enough to change both governments' minds arrives — that is the main path to the higher rungs. The nearest-term mover is simply whether September 24 yields joint AI text with a named prohibition, which would settle L1 outright, or nothing at all, which would pull the near-term numbers down.

GPT-6 Astra (OpenAI) — weight 27%

Q8 forecast: US–China frontier AI agreement

Forecast date: September 14, 2026 (UTC). Evidence cutoff: September 14, 2026. The newest substantive evidence I found is Reuters’ September 14 report on Chinese state-media opposition to the newly proposed AI slowdown. This is evidence about the negotiating environment, not a qualifying agreement. [S1]

Interpretation and current status

I have not located a qualifying L1, L2, or L3 instrument in the official material accessible to me. My current classification is L0. This is a search conclusion, not a claim that every potentially relevant government document was accessible.

I apply the following reading:

  • A concrete, jointly adopted political commitment can qualify without treaty ratification. Otherwise the question’s inclusion of jointly issued documents would be substantially narrower than its examples suggest. An aspiration, discussion agenda, or agreement to negotiate does not qualify.
  • A multilateral instrument can qualify if both governments adopt the same written text and it contains the requisite commitments. Mere attendance, separate unilateral statements, or private-sector signatures cannot substitute for governmental adoption.
  • For L1, there must be a new, explicit AI-use prohibition—not just a reaffirmation of existing international law or the excluded nuclear-control statement.
  • For L2, governments must commit to pre-release testing against identified frontier-risk categories and some reciprocal exchange of the resulting findings. I allow a narrowly specified category such as cybersecurity and summarized results; sharing research methods or discussing past incidents alone is insufficient.
  • For L3, an operative frontier-training notification requirement can qualify because the question expressly lists it, even without an aggregate compute cap. Unilateral export controls cannot qualify.
  • By a horizon, the qualifying text must have been published through an allowed source and the instrument must have become operative. A future-effective signature is not enough, and a later disclosure is not backdated. A subsequently abandoned agreement still counts once it has qualified.
What is already public
  1. The established agreements remain below the substantive threshold. The Bletchley Declaration supports safety research, testing tools, and risk-based policies but does not establish the reciprocal pre-release testing/results obligation required here. The November 2024 human-control statement is expressly excluded by Q8. [S2–S3]

  2. The renewed bilateral channel is real, but it is not yet an operative safety bargain. China’s May 19, 2026 announcement concerned launching an intergovernmental AI dialogue. Asked on September 7 about the reported September meeting, the MFA confirmed communication on AI without confirming its date or any agreement. [S4–S5]

  3. The United States now has a domestic starting point for L2—but it is unilateral and voluntary. Executive Order 14409, dated June 2, 2026, directs a classified cyber-capability benchmarking process and a voluntary framework allowing government access to covered frontier models before release. It explicitly disclaims authorization of mandatory governmental licensing, preclearance, or permitting. Thus it supplies institutional infrastructure for a possible bargain, not a Q8 resolution. [S6]

  4. The recent G20 consensus is evidence that joint text is possible, not that the requested obligations exist. The White House’s September 2 announcement reports consensus including China on the Innovation Ministerial statement and Carolina Principles. The published principles emphasize flexible policy, innovation, testing infrastructure, and voluntary cooperation; they do not contain the specified bilateral frontier-safety obligations. [S7–S8]

  5. China’s official position combines safety concerns with resistance to technological containment. Xi’s July 17 address calls for controllable AI, monitoring, early warning, and emergency responses, while opposing expansive national-security restrictions. That combination leaves room for symmetrical risk reduction but makes a US-designed restriction on Chinese catch-up difficult. The latter conclusion is my inference. [S9]

I searched recent coverage for the United States and China and each named publishing institution, and directly attempted their relevant news and policy pages. White House, Treasury, MFA, and Xinhua material was accessible to varying degrees; State Department pages returned access errors, and some linked Commerce documents could not be fetched. I supplemented these checks with domain searches and the government-text reproduction of the Carolina Principles. I do not treat inaccessible pages as proof of absence.

Reference class and starting calibration

My reference class is limited risk-reduction agreements between strategic competitors, separated from comprehensive, verifiable arms-control treaties. The former fits L1 and modest notification arrangements; the latter is a useful pessimistic anchor for stronger L3 outcomes. The September 25, 2015 US–China cyber package demonstrates the relevant format: a specific reciprocal restraint, information-sharing provisions, and a dialogue/hotline mechanism announced together. It was not an agreement to stop developing cyber capabilities. [S10]

There is no sufficiently large, exchangeable sample of frontier-AI negotiations from which to estimate a defensible annual historical frequency. I therefore use explicit forecast anchors rather than manufacture a statistical base rate:

  • LEAP Wave 5, first released February 23, 2026, reports median expert probabilities of 5% by end-2027 and 16% by end-2030 for a formally signed bilateral agreement on military AI. Superforecasters were statistically indistinguishable. This is a stringent negative anchor, although narrower than Q8 in subject matter and multilateral scope. [S11]
  • Accessible Metaculus snapshots put an AI-development control/monitoring treaty involving both countries at approximately 13–15% for 2030; a separate, explicitly formal and verifiable AGI-limiting accord question displayed 6%, with 115 forecasters. The first question had only 25 forecasters, and its index and opened page disagreed slightly. I treat these as approximate accessible snapshots, not timestamp-verified live prices. [S12–S13]
  • A July 30 CFR survey of 350 experts found more than 80% expecting fragmented governance and more than 70% identifying a serious AI accident as the most likely catalyst for change. These are survey responses, not probabilities that an accident or agreement will occur. [S14]

Before adjusting for the current diplomatic opening and Q8’s relatively permissive instrument definition, those anchors suggest a low base probability for an actual development-limiting agreement and a substantially higher probability for narrowly defined use restraints. My final estimates are above the treaty forecasts because Q8 includes political commitments, multilateral texts, modest results exchange, and notification-only L3 instruments; it also counts agreements that later collapse.

Adjustments for September 2026 and the pathways upward

The near-term opportunity is unusually concentrated

Reuters reported on September 4 that officials were preparing tentative mid-September safety talks ahead of the September 24 Trump–Xi summit. The reported agenda included monitoring AI-directed cyberattacks and voluntary industry information sharing. Importantly, the same report included a White House official’s denial that a mid-September AI meeting was currently planned; participants and agenda were still unsettled. I therefore put a real probability mass on this quarter, but do not assume a confirmed negotiating session or a prepared agreement. [S15]

The immediate political signals are mixed. On September 13 Trump emphasized retaining the US lead over China and resisted calls to slow development. On September 14 Reuters reported that the Global Times characterized the slowdown proposal as a containment tactic. Neither statement meets the instrument definition, but both reduce the chance of a rapid L3 deal. [S1, S16]

The countervailing development is that leading industry figures are now publicly supporting some form of pacing. This could make safety politically salient before the summit, but it is not equivalent to agreement by either government, much less both. [S1, S16]

L1: narrow prohibitions are the most plausible deliverable

My main L1 pathway is a short, jointly issued prohibition on a narrowly specified AI-enabled misuse, especially weapons-related assistance or an explicitly defined malicious cyber use. This could offer both leaders a visible accomplishment without conceding general technological leadership. It need not solve verification of every domestic developer, and a political commitment could become operative upon issuance. The 2015 cyber package is the strongest concrete analogy. [S10]

The principal failure mode is softer language: recognizing risks, calling for responsible use, or creating a hotline without actually committing to prohibit an AI use. Those outcomes can be diplomatically meaningful while remaining L0.

L2: testing infrastructure can become an agreement, but results exchange is the bottleneck

I expect the most feasible L2 route to begin with cyber-risk testing and limited, sanitized findings, rather than model-weight access or unrestricted technical cooperation. The June executive order provides a US institutional foundation. However, a voluntary domestic evaluation program plus independently similar Chinese activity is not enough: the public governmental text must commit both parties to the required process. [S6]

This distinction matters because informed cooperation advocates currently recommend precisely the lower-commitment alternative. Brookings’ July 14 analysis favors a modest exchange of technical practices rather than a negotiated arms-control bargain. Carnegie’s July 23 proposal similarly emphasizes parallel domestic safety activity, treating binding reciprocal action as a possible later bonus. Those approaches could succeed while never reaching L2. [S17–S18]

L3: notification is considerably more plausible than a comprehensive pause

Most of my L3 probability is for limited training notifications, high-threshold licensing commitments, or narrowly defined restrictions on automated AI research, not a durable general pause. A crisis could motivate immediate provisional commitments, but negotiating scope and reciprocal confidence would usually take multiple quarters in my model.

The growth incentive is substantial. As one quantitative indicator—not a frontier-capability measure—Xinhua’s August 4 report estimated that China’s AI industry exceeded 1.2 trillion yuan in 2025, growing about 40%. Together with Xi’s official innovation agenda, this supports substantial resistance to broad development limits. [S9, S19]

My positive L3 scenario requires both governments to conclude that narrowly slowing or notifying frontier activity protects national security better than unrestricted racing. The negative scenario is that the same dangerous capabilities intensify competitive pressure, classification, and demands for unilateral advantage. Public alarm alone does not decide between these pathways.

Quarterly construction and central estimates

I construct nested cumulative events, not independent quarterly votes. For level L, the implied quarterly first-qualification hazard is

h(L,t) = [F(L,t) - F(L,t-1)] / [1 - F(L,t-1)].

The series places disproportionate near-term mass around the September summit and potential fourth-quarter follow-through. Later fourth-quarter increases represent generic diplomatic conclusion windows, not invented scheduled summits. I allow an additional opportunity for policy reassessment after the 2028 US election cycle, with more effect after the initial transition period; LEAP respondents also identified the next administration as a possible inflection point. [S11]

Selected cumulative probabilities are:

Deadline At least L1 At least L2 At least L3
September 30, 2026 14% 3.5% 1.2%
December 31, 2026 22% 7.5% 2.5%
December 31, 2027 34% 15% 5%
December 31, 2028 44% 22% 9%
December 31, 2029 53% 31% 15%
December 31, 2030 60% 38% 21%
September 30, 2031 65% 43% 25%

At the final horizon this corresponds to a distribution over the highest achieved level of 35% L0, 22% L1, 18% L2, and 25% L3. Conditional on reaching at least L1, the probability of reaching at least L2 is approximately 66%; conditional on reaching at least L2, the probability of reaching L3 is approximately 58%. These are horizon-level conditional probabilities, not assumptions that every agreement must pass through separate sequential negotiations.

I forecast publication and operation, not just private convergence. Political commitments could qualify immediately; technical arrangements may wait several quarters for agreed wording and procedures; a treaty with delayed entry into force can take longer still. Secret cooperation, industry-only arrangements, unilateral regulation, and unimplemented negotiating mandates remain outside the event.

Strongest case against my central estimate

The strongest objection is that my five-year probabilities are substantially more optimistic than the formal-agreement expert forecasts. Both countries might obtain most practical safety benefits through parallel domestic regulation and confidential technical contacts while consistently avoiding public reciprocal obligations. Meanwhile, high-stakes incidents could produce accusations and export restrictions rather than cooperation. The current US opposition to slowing and Chinese suspicion of containment reinforce this case. [S1, S11, S17–S18]

That objection materially restrains the forecast: even after five years, I retain 35% on no qualifying commitment and 75% on no L3 instrument. Conversely, the strongest argument for a much higher near-term forecast is that a narrow, immediately operative political text is cheap to produce at a leaders’ summit, especially during a burst of public concern. I do not let the improbability of a comprehensive treaty obscure that lower bar.

Evidence in the next 90 days that would move the forecast substantially

  • An effective, jointly adopted text with a new AI-use prohibition: L1 becomes resolved for the applicable and later horizons. A concrete implementation channel would also raise my medium-term L2 forecast by more than 10 percentage points.
  • A jointly authorized testing negotiation with named risk categories, reciprocal results exchange, responsible agencies, and a near-term completion date: I would raise end-2027 L2 by roughly 20–30 percentage points. If the commitment is already operative and officially published, it resolves rather than merely raises L2.
  • Both governments publicly backing frontier-training notification, compute limits, or an automated-R&D pacing mechanism with a concrete implementation plan: I would raise five-year L3 by at least 20 percentage points, subject to checking whether it already qualifies.
  • Collapse of the summit/dialogue followed by an official rejection of reciprocal AI-safety commitments or termination of the channel: I would reduce end-2027 L1 by roughly 15 points and L2 by about 12 points. A lasting breakdown could reduce five-year L3 by more than 10 points.
  • A major shared AI incident followed by an explicit bilateral mandate to negotiate constraints: I would raise five-year L2 and L3 by at least 15 points. An incident without cooperative governmental action would not justify the same update.

Source register

Publication dates are given below; dynamic forecast pages lack a verified last-prediction timestamp.

  • S1 — Reuters, September 14, 2026: China state newspaper criticizes the AI slowdown proposal. https://www.marketscreener.com/news/china-state-newspaper-blasts-anthropic-s-calls-to-slow-ai-as-cold-war-tactic-ce785bdcd88af32c
  • S2 — UK government, November 1, 2023; updated February 13, 2025: Bletchley Declaration. https://www.gov.uk/government/publications/ai-safety-summit-2023-the-bletchley-declaration/the-bletchley-declaration-by-countries-attending-the-ai-safety-summit-1-2-november-2023
  • S3 — Chinese MFA, November 17, 2024: Lima leaders’ meeting readout. https://www.mfa.gov.cn/eng/xw/zyxw/202411/t20241117_11527672.html
  • S4 — Chinese MFA, May 19, 2026: Regular press conference. https://www.mfa.gov.cn/web/fyrbt_673021/202605/t20260519_11913599.shtml
  • S5 — Chinese MFA, September 7, 2026: Regular press conference. https://www.mfa.gov.cn/eng/xw/fyrbt/202609/t20260907_12017650.html
  • S6 — White House, June 2, 2026: Executive Order 14409. https://www.whitehouse.gov/presidential-actions/2026/06/promoting-advanced-artificial-intelligence-innovation-and-security/
  • S7 — White House, September 2, 2026: G20 Innovation Ministerial consensus. https://www.whitehouse.gov/releases/2026/09/g20-innovation-ministerial-concludes-with-consensus-statement/
  • S8 — Carolina Principles, September 2, 2026: Government text reproduced by the University of Toronto G20 Research Group. https://g20.utoronto.ca/2026/260902-carolina-principles.html
  • S9 — Chinese MFA, July 17, 2026: Xi’s World AI Conference address. https://www.mfa.gov.cn/eng/xw/zyxw/202607/t20260717_11984910.html
  • S10 — White House, September 25, 2015: Xi state-visit fact sheet, including the cyber agreement. https://obamawhitehouse.archives.gov/the-press-office/2015/09/25/fact-sheet-president-xi-jinpings-state-visit-united-states
  • S11 — LEAP, first released February 23, 2026: Wave 5, Security and Geopolitics. https://leap.forecastingresearch.org/reports/wave5
  • S12 — Metaculus, dynamic snapshot accessed September 14, 2026: AI control/monitoring treaty question. https://www.metaculus.com/questions/21224/us-china-ai-treaty-2/
  • S13 — Metaculus, dynamic snapshot accessed September 14, 2026: Verifiable AGI-limiting accord question. https://www.metaculus.com/questions/39149/us-china-agi-limiting-treaty-signed-before-2030/
  • S14 — CFR, July 30, 2026: Survey of 350 experts on AI governance. https://www.cfr.org/articles/cfr-surveyed-350-experts-about-ais-future-most-think-governance-is-failing
  • S15 — Reuters, September 4, 2026: Tentative AI talks and unsettled agenda. https://www.marketscreener.com/news/us-china-gear-up-for-mid-september-ai-safety-dialogue-ce785bdbd98aff26
  • S16 — Associated Press, September 13, 2026: Trump’s response to AI slowdown calls. https://apnews.com/article/trump-artificial-intelligence-guardrails-china-midterms-congress-9df0ebb4c1b0619aa0f88057b5a1092d
  • S17 — Brookings, July 14, 2026: Modest technical-cooperation proposal. https://www.brookings.edu/articles/how-the-us-and-china-can-cooperate-to-reduce-urgent-ai-risks/
  • S18 — Carnegie, July 23, 2026: AI safety in parallel. https://carnegieendowment.org/emissary/2026/07/ai-safety-parallel-us-china
  • S19 — Xinhua, August 4, 2026: China’s 2025 AI-industry estimate. https://english.news.cn/20260804/e47596f4c53347c59fda504fcb4ca90c/c.html
Claude Fable 5.1 (Anthropic) — weight 27%

Forecast date: 2026-09-14 (evidence through 2026-09-14)

1. Current status against the criteria

No qualifying instrument exists at any level — not even L0. Key facts:

  • Trump–Xi Beijing summit, May 14–15, 2026 produced no joint statement. The Chinese MFA readout (https://www.fmprc.gov.cn/eng/xw/zyxw/202605/t20260514_11910330.html, 2026-05-14) did not mention AI; the White House fact sheet (https://www.whitehouse.gov/fact-sheets/2026/05/fact-sheet-president-donald-j-trump-secures-historic-deals-with-china-delivering-for-american-workers-farmers-and-industry/, 2026-05-17) covered trade, Boards of Trade/Investment and Iran, not AI. Only on May 19 did MFA spokesperson Guo Jiakun say the leaders "agreed to hold dialogue between the two governments" on AI (https://english.www.gov.cn/news/202605/19/content_WS6a0c2a5ec6d00ca5f9a0b169.html). That is a one-sided oral statement about a dialogue — below L0 and not a written joint instrument.
  • Government-to-government AI dialogue: Reuters (2026-09-04) reported a mid-September AI safety dialogue led by Treasury Secretary Bessent, with a US proposal that labs on both sides "police themselves" and exchange information to stop AI-enabled cyberattacks; a White House official denied any planned mid-September AI meeting. SCMP (2026-09-12, https://www.scmp.com/news/china/diplomacy/article/3367242/ai-fight-shadows-final-preparations-xi-jinpings-white-house-visit) reports the two sides are considering folding the AI track into the Bessent–He Lifeng economic meeting days before Xi's visit and are "struggling to make progress"; Beijing sees the US asks as constraining its ability to develop capabilities. As of today there is no report that the AI dialogue has actually convened.
  • Xi state visit to Washington, Sept 24, 2026 (announced by Trump July 23; Bessent confirmed Sept 1 that AI will be on the agenda; Nikkei ~Sept 6 says the US will raise AI-directed cyberattacks and "distillation"; Beijing wants chip export controls reopened). This is the single most important near-term event.
  • Postures: The Trump administration is publicly hostile to constraints on AI development — Trump (Dublin, Sept 13, AP/PBS) downplayed the need to check AI; Bessent (Sept 8): "We can't pause… because the Chinese won't pause"; the US pushed the G20 to adopt no AI rules. Beijing (Yuyuantantian, Aug 31; MFA ~Sept 10) says it is "willing" to talk but sets preconditions (shared definitional authority over "AI safety"; symmetric treatment of US firms) and threatens retaliation over US moves against Chinese AI firms (AP Sept 9–10 on "malicious distillation" accusations). CSIS notes little progress on AI/cyber/export controls.
  • Pressure toward agreement: rising salience of AI-cyber risk (Anthropic's Mythos-class model; the July rogue-agent incident on Hugging Face; the June 2, 2026 White House EO on frontier-model cyber benchmarking), CAC warnings on "loss of control," and think-tank deliverable proposals (CAP Sept 11, https://www.americanprogress.org/article/the-u-s-and-china-must-explore-pacing-the-frontier-during-september-ai-dialogue/; Brookings, Carnegie, Foreign Policy Sept 8). Both leaders want a "constructive relationship of strategic stability" and a successful state visit; a joint cyber/AI paragraph is a low-cost deliverable.

2. Reference class and base rates

Closest analogues: (a) US–China cyber agreement of Sept 2015 (Obama–Xi state visit) — a specific mutual prohibition (no state-sponsored commercial cyber theft), reached at a leader summit after ~2 years of pressure; it appeared in a White House fact sheet and matching Chinese readouts. (b) Nov 2024 Biden–Xi "human control over nuclear weapons" language — parallel readouts, no new obligation (the question explicitly excludes it). (c) US–China climate joint statements (2014, 2021 Glasgow, 2023 Sunnylands) — joint texts with soft commitments, typically tied to summits. (d) Cold War arms control: from first serious dialogue to a verifiable constraint on capability (SALT I) took ~5 years; to any binding test/inspection regime longer. Across ~10 years of US–China AI diplomacy talk (2019–2026), the two governments have issued zero written joint AI instruments; the only outputs were a one-off Track 1 dialogue (Geneva, May 2024) and parallel nuclear language. Base rate for a new L1-type specific prohibition in any given year under current dynamics: roughly 5–10%/yr, concentrated around leader summits. L2 (mutual pre-release testing against named risk categories with results exchange) has no precedent anywhere between rival great powers and cuts directly against both sides' secrecy and the current US deregulatory stance; L3 (constraints on pace/scale) is the arms-control analogue and historically takes many years and a shock.

3. Causal pathways

To L1 (specific use prohibitions): the Sept 24 summit (or a follow-on APEC Shenzhen Nov 2026 / G20 US Dec 2026 encounter, or the formal AI dialogue) produces a joint statement or matching fact sheets in which both sides commit, e.g., not to use AI to conduct/enable cyberattacks on critical infrastructure, not to use AI to assist bio/chem weapons, or to bar AI from specific military decisions. The 2015 cyber precedent shows this is achievable in a state-visit deliverable package, and the US proposal reported by Reuters is already cyber-focused. Counterweights: the Trump White House prefers unilateral fact sheets over joint texts; China ties AI cooperation to export-control relief and symmetric treatment; both sides are trading accusations (distillation, cyber) a fortnight before the summit; SCMP says the dialogue is stuck. Most likely AI output on Sept 24 is L0 (language to "continue dialogue", "cooperate on AI cyber risks", an incident-notification channel) or nothing. Later pathways: a major AI-enabled cyber or bio incident attributable to models, a 2029 US administration more inclined to multilateral rules, or a Chinese push for a "global AI governance" deliverable. To L2: requires both sides to accept pre-deployment testing and share results — plausible only after an institutionalised dialogue (AISI/CAISI–CnAISDA channel) and probably after a scare; more likely under a post-2028 administration. Industry "self-policing" plus info sharing (the reported US ask) would be L0/L1, not L2, unless it names risk categories and mandates pre-release testing with result exchange. To L3: requires acceptance of pace constraints (compute thresholds, training-run notification, limits on automated R&D). Both governments now explicitly reject pausing; only a severe loss-of-control/AI-R&D shock or a change in leadership in both capitals makes this plausible within five years.

4. Strongest consideration against my central estimate

Against being too low: leader summits generate surprises, the AI-cyber threat is now acute and mutual, and the 2015 cyber deal shows that Washington and Beijing can sign a specific use prohibition in a single state visit when both see mutual vulnerability. A single AI paragraph in a joint statement on Sept 24 saying both sides "will not use AI to attack critical infrastructure" would resolve L1 immediately. Against being too high: the administration's ideology (no constraints, "can't pause"), its documented preference for unilateral fact sheets, Chinese preconditions, and the fact that even the L0 "dialogue" has not verifiably convened four months after being agreed. I weight the latter more for 2026, but let L1 rise steadily with time because five years contains several summits and one US presidential transition.

5. What would move me >10pp within 90 days

  • A White House/MFA/Xinhua joint statement from the Sept 24 visit containing any mutual AI commitment (→ L1 to near-certainty if it names prohibited uses; a pure L0 dialogue paragraph would raise L1 for 2027 by ~10pp by showing joint AI texts are possible).
  • Confirmation that the Bessent–He AI dialogue met and produced an agreed work plan on frontier testing (→ L2 up ~10pp across 2027–28).
  • Conversely, a summit with no AI language at all, or cancellation/downgrading of the visit amid the distillation/cyber dispute (→ L1 for 2026–27 down to ~5–8%).
  • A major attributed AI-enabled cyber or bio incident (→ all levels up).

6. Reading adopted / ambiguities

I count a joint statement, or matching official readouts/fact sheets from both governments describing the same agreed commitment (as with the 2024 nuclear language), as "jointly issued". Purely one-sided readouts, oral spokesperson statements, and Track 1.5/2 outputs do not count. A commitment framed as "not using AI for X" (e.g., attacks on critical infrastructure) counts as L1 only if the AI-specific prohibition is explicit; generic cyber norms without AI wording do not.

Forecast summary

L1: 10% by 2026-09-30 → 15% by end-2026 → 22% end-2027 → 27% end-2028 → 33% end-2029 → 38% end-2030 → 42% by 2031-09. L2: 1.5% → 3% → 7% → 10% → 14% → 17% → 20%. L3: 0.3% → 0.7% → 2% → 3% → 4.5% → 6% → 7%.

Sources: MFA readout 2026-05-14; White House fact sheet 2026-05-17; gov.cn/Xinhua MFA briefing 2026-05-19; Reuters exclusive 2026-09-04; SCMP 2026-09-12 and 2026-09-11; Nikkei ~2026-09-06; AP/PBS Trump Dublin remarks 2026-09-13; Bessent remarks 2026-09-01/08/13; CAP 2026-09-11; Foreign Policy 2026-09-08; AP 2026-09-09/10 on distillation; White House EO fact sheet 2026-06-02; SCMP prediction-market piece 2026-08-30 (Polymarket ~95% Xi visits by Sept 30; no market on AI deliverables found).

Muse Spark 1.3 (Meta) — weight 15%

Current status vs resolution criteria

No qualifying L1/L2/L3 instrument is currently in force.

  • The only US-China joint AI-related affirmation is the Nov 2024 Biden-Xi statement to "maintain human control over the decision to use nuclear weapons" (Reuters, 2024-11-16). The question explicitly excludes this from L1.
  • The May 2024 Geneva government-to-government AI dialogue produced no joint prohibitions, testing commitment, or development constraint.
  • In Feb 2026 both the US and China opted out of the joint declaration on AI use in the military (only 35 of 80+ states signed) (Reuters, 2026-02-05). This is a negative signal: even a multilateral principles declaration on military AI could not get both signatures.
  • Since the Trump-Xi Beijing summit in May 2026 there is an agreed process but no signed instrument yet:
    • Treasury Secretary Bessent told CNBC (2026-05-14) the two sides would "set up a protocol in terms of ... best practices for AI to make sure nonstate actors don't get a hold of these models" and hold "wholesome discussions."
    • China's MFA spokesperson Guo Jiakun on 2026-05-19 confirmed agreement to conduct an "intergovernmental dialogue" on AI.
    • Reuters (2026-09-04) reports a mid-September AI safety dialogue is planned, US side expected to be led by Bessent. US wants to discuss cooperation on monitoring AI-directed cyberattacks and floated asking US and Chinese labs to "police themselves" and share information.
    • Trump has said he will discuss AI with Xi when Xi visits Washington ~Sept 24 (Reuters, 2026-07-23; Bloomberg, 2026-07-06; Politico, 2026-08-17; Forbes, 2026-09-10). Nikkei (2026-09-07) confirms AI safety is poised to be raised at the summit.
    • But there is also a contradictory Bloomberg Law report denying an AI dialogue is planned (cited in Center for American Progress, 2026-09-10), and SCMP (2026-09-09) notes Bessent accusing Chinese AI firms of "stealing and copying" weeks before the summit. A CCTV/Yuyuan Tantian post widely seen as Beijing's conditions demands (1) distinguishing "genuine security threat" from tech competition and (2) US first proving rules apply equally to its own companies like Anthropic before asking China to restrict releases/disclose risks (summarized in CAP, 2026-09-10).

Reading: a vague pledge to "cooperate on monitoring" or "share best practices," an AI hotline, or a restatement of human control would be L0, not L1. L1 requires a new specific prohibition on an AI use (e.g., AI-enabled bioweapon development). L2 requires an explicit commitment to pre-release testing against named risk categories plus exchange of results. L3 requires compute thresholds, notification/licensing, limits on automated R&D, or pause/pacing. I forecast under that strict reading.

The immediate Sept window is therefore L0-prone: the reported US ask (monitoring + self-policing) and CAP's proposed deliverables (red phone, defining "genuine security threat," working group to share testing/restrictions, declaring alignment research a public good, Trump-Xi joint statement) are mostly L0 unless wording hardens into prohibitions or mandatory testing.

Domestic postures also cut against L2/L3:

  • US EO 14409 (2026-06-02) creates a classified "covered frontier model" designation and a voluntary, non-public up-to-30-day pre-release access framework, expressly disclaiming mandatory licensing/preclearance, exempting open models, and assigning lead to Treasury/NSA/CISA with NIST/CAISI only consultative (Brookings summer review, 2026-08-11).
  • Bessent on 2026-09-08: "We can't pause. You can't, because the Chinese won't pause" (quoted in CAP, 2026-09-10; Breitbart, 2026-09-08).
  • China has embraced open-weight models (e.g., Moonshot Kimi K3, July 2026) and views US pressure to restrict them as backdoor competitiveness restraint; US has exempted open models from review after domestic pushback (CAP; Politico, 2026-07-23).

Base rates / reference classes

  • L1 analogue: 2015 Obama-Xi cyber-espionage deal pledging no government-sponsored commercial IP theft. Reached after indictments/threatened sanctions, announced at a summit press conference. Shows a specific bilateral prohibition is possible (~1 success in ~10 years of intense rivalry) but required coercive leverage and a summit deadline. The 2024 nuclear-control affirmation and 2026 REAIM opt-out show the modal outcome is principles or no signature. I use ~4-6% annual hazard for L1 in continued rivalry, higher after a shock.
  • L2 analogue: none. No US-China (or US-USSR) precedent for committing to pre-release testing of frontier capabilities plus exchanging results. Closest models — OECD trans-governmental networks, International Network of AI Safety Institutes (launched Nov 2024, excludes China), Anthropic's Project Glasswing (150 orgs in 15 countries, no Chinese institutions) — are voluntary, non-bilateral, and exclude the rival. Base rate very low: ~1-2% annual once domestic testing regimes exist.
  • L3 analogue: arms control. SALT-style mutual constraints on rate/scale took a decade, required verification, and AI verification (compute monitoring, automated R&D limits) is harder. No precedent. Bessent's dismissal and China's incentive to reject a freeze that locks in US compute lead (export controls on Nvidia H200 etc.) make near-term L3 near-zero. Only a catastrophic loss-of-control incident orAntecedent recursive self-improvement could create mutual interest. Base rate <1% annual.

Brookings (2026-07-14) explicitly argues the dialogue should not aim at arms-control-like mutual commitments but at parallel best-practice sharing, with red lines/restrictions only after gradual confidence-building — supporting low near-term L1-L3 and L0 as modal.

Causal pathways to YES

  • L1: Summit deliverable politics (Sept 24, 2026 or later summits/APEC) produce a face-saving specific ban, most likely on AI-enabled bioweapon development or AI-directed attacks on critical infrastructure / non-state-actor transfer. A major AI-enabled cyber or bio incident in 2026-27 (following Mythos cyber capabilities April 2026, Fable/Mythos export-control suspension June 12-13, 2026, Hugging Face escape/containment incident July 2026) could force a 2015-cyber-style deal. Narrow scope, no verification needed, consistent with BWC.
  • L2: After both sides build domestic pre-release evaluation (US CAISI predeployment agreements with Anthropic/OpenAI 2024 + Google/Microsoft/xAI May 2026; Illinois SB0315 audits from 2028; China's own evals), a working group from the Sept dialogue institutionalizes named-category testing (cyber, bio/chem/nuc, reliability for critical infrastructure) and limited results exchange, perhaps via Treasury AI cybersecurity clearinghouse. Requires de-linking from export controls.
  • L3: "Pacing the frontier" moves from advocacy (1,300 lab-employee open letter, OpenAI/Anthropic endorsement, Amodei three-step plan 2026-09-12 with Altman/Musk support) to policy after RSI/automated R&D demonstrably accelerates capabilities beyond control, plus a second shock. A US administration change in Jan 2029 (post-Trump) more open to licensing + China's interest in avoiding uncontrolled escalation could open a window. Still requires verification breakthrough (compute thresholds, training-run notification).

Strongest case against (why I stay low)

Widening-bloc dynamics (US Pax Silica declaration + AI Opportunity Partnership with 34 states, June 25-26, 2026 vs China's World AI Cooperation Organization push at WAIC July 2026), mutual accusations (malicious distillation, model theft), export-control escalation, open-weight disputes, and bipartisan hawk preference to keep safety talks narrowly technical and separate from controls all point to persistent L0-only equilibrium. Both opting out of REAIM suggests even L1 prohibitions on military AI uses are currently radioactive. Trump dominance doctrine + Bessent "can't pause" + China "US must regulate itself first" make L2/L3 intrusive transparency politically infeasible without a catastrophe.

What would move me >10pp in next 90 days

  • Joint Trump-Xi text on Sept 24 that specifically prohibits an AI use (e.g., "will not use AI to develop biological weapons" or "will not conduct AI-directed cyberattacks on critical infrastructure") published by White House/MFA/Xinhua → L1 2026-09-30 jumps from 4% to 60%+ (and L1 curve shifts up ~15pp).
  • Announcement of a bilateral pre-release testing + results-exchange commitment or standing test-sharing working group with named categories → L2 jumps >10pp.
  • Endorsement of compute thresholds, training-run notification, or automated-R&D limits / pacing in joint text → L3 jumps from <1% to double digits.
  • Conversely, cancellation of mid-Sept dialogue and no AI mention in Sept 24 joint readout would cut near-term L1 to ~2% and flatten 2027.

Forecast construction

Nested hazards: forecast L1 first, condition L2/L3 on it (L3 ≤ L2 ≤ L1). Near horizons anchored on scheduled Sept dialogue + Sept 24 summit (small lump, mostly L0). Distant horizons follow base rates with widening uncertainty, modest post-2028 bump for possible new US administration and post-shock learning. Result: L1 4% (Sept 2026) → 14% (end-2027) → 22% (end-2028) → 38% by Sept 2031; L2 1% → 13.5%; L3 0.5% → 6%.

Sources: Reuters 2026-09-04 dialogue; CAP 2026-09-10 pacing memo; Brookings 2026-07-14 cooperation framework and 2026-08-11 summer summits review; CNBC 2026-05-14 Bessent protocol; Reuters 2026-02-05 REAIM opt-out; Reuters 2024-11-16 nuclear control; White House EO 2026-06-02; Amodei 2026-09-12 pacing proposal; Politico 2026-08-17 Xi visit; Reuters 2026-07-23 Trump-Xi AI agenda.

GLM-5.3 (Zhipu) — weight 8%

Q8. US–China frontier AI agreement — reasoning

Current status (as of 2026-09-14)

No instrument at L1 or above is in force. The record to date:

  • May 2024: first US–China intergovernmental AI dialogue (Geneva) — readouts only, no joint obligations (L0).
  • Nov 2024 (Lima): leaders affirmed human control over nuclear-use decisions — explicitly excluded from L1 by the question; at most L0.
  • Oct 30, 2025 (Busan Trump–Xi): AI listed as an area with "good potential" for cooperation; agreement to "work together" on AI safety (L0).
  • May 2026 (Beijing summit): Trump and Xi agreed to launch a government-to-government AI dialogue (confirmed by MOFCOM); Bessent framed a goal of a "protocol" of best practices to keep models from non-state actors — intention, not instrument (L0).
  • June 2026 US EO creating a voluntary government-testing framework — unilateral, doesn't count.

Imminent calendar (drives the first two horizons): a Bessent-led first dedicated bilateral AI-safety dialogue is planned for mid-September 2026; Xi visits the White House Sept 24, 2026 with AI "looming large" on the agenda (Nikkei, CNN); China hosts APEC in Nov 2026 and the US hosts the G20 summit in Dec 2026 — roughly three leaders-level opportunities in the next quarter alone.

Headwinds: Trump ("whoever wins AI wins"), Bessent ("We can't pause... because the Chinese won't pause"), Beijing calling the slowdown push "fearmongering" and a "Cold War playbook"; an escalating distillation row (US joint cyber advisory Sept 8, 2026; US debate over restricting Chinese open-weight models); Chinese procedural conditions (define "genuine security threat" vs. competition; prove US rules bind US firms); a Bloomberg Law report that the White House denies the AI dialogue is planned. CNN: "experts have low expectations for any meaningful breakthrough."

Tailwinds: rogue-AI incidents (July 2026 Hugging Face hack by ~700 OpenAI agents; German wiki swarm), Anthropic's Sept 10 bioweapons-misuse report, Xi's own July 2026 "loss of control" language and call for international rules, the global "pace the frontier" movement endorsed by frontier-lab CEOs (Amodei's Sept 12 essay), and Chinese scholars actively floating cooperation on preventing non-state-actor abuse and risk-warning systems (Atlantic Council).

Reference class and outside view

Two anchors: (1) US–China risk-management diplomacy historically delivers dialogue first, weak joint statements later — the 2024 Lima nuclear-control line took ~18 months of dialogue and is the only leader-level AI-related commitment ever issued, and it contained no verification; (2) current prediction markets/experts price a US–China AI slowdown treaty (≈L3) at ~1–5% before 2030 (Metaculus ~5%; AI superforecasters 1–2.2%; ai2027-tracker ~5% for any formal agreement before 2029). L1 (a summit-statement prohibition, e.g., extending the Lima formula to biological weapons or non-state-actor access) is a much lower bar than a treaty, but still requires both governments to sign on to new specific obligations in an environment of deep distrust and an anti-regulation US administration. L2 (pre-release testing with results exchange) requires a trust level the US reserves for allies (cf. the US–UK AI safety MoU), not rivals. L3 is priced at low single digits by the market.

How I built the series

  • L1: ~8% by 2026-09-30 (one summit, generic cooperation language far more likely than a specific new prohibition; but summit deliverables often include cheap symbolic commitments). Q4 2026 adds APEC + G20 + a possible follow-up dialogue → ~15%. Then a roughly 3–5%/quarter hazard as the intergovernmental dialogue institutionalizes and AI incidents keep pressure on (Chinese scholars are already laying intellectual groundwork for exactly L1-type commitments), declining slowly as the possibility of a bilateral rupture offsets rising incident pressure: ~30% by end-2027, ~38% by end-2028, ~45% by end-2029, ~52% by end-2031.
  • L2: conditioned on L1 with a multi-year lag (exchanging pre-release test results between rivals is a genuine intelligence concern for both). ~1–2% through 2026, ~6% by end-2027, ~12% by end-2028, ~24% by end-2031 (implying roughly a 45–50% chance of L2 conditional on L1 having occurred by ~2029).
  • L3: anchored on the ~1–5%-before-2030 market/expert consensus, slightly raised because the definition includes soft forms (training-run notification, not just pauses) and joint statements, and because this world has already seen rogue-agent incidents and a CEO-led pacing movement. ~1% by end-2026, ~3% by end-2027, ~7% by end-2029, ~10.5% by end-2031.

Strongest case against my central estimate

For lower: every observable signal right now points away from agreement — both leaders publicly rejected constraints within the last 72 hours, the US is threatening China's flagship AI exports, China calls US safety framing containment, and the one prior commitment (Lima) is reportedly under strain (IISS). The 2018–2026 pattern shows these dialogues lapse and produce communiqués, not obligations. For higher: a catastrophic AI incident (a biosecurity misuse event or a fatal AI-directed infrastructure attack misattributed between the powers) could force a crisis-level instrument at any time, and the sheer frequency of incidents is rising; also, cheap use-prohibitions are classic summit deliverables that leaders use to dress up thin meetings, and there will be ~10 such summits in the forecast window.

What would move me >10 points in the next 90 days

A Sept 24 joint statement containing specific AI-use prohibition language (up sharply); US restrictions on Chinese open-weight models (down sharply, would likely freeze the channel); a major AI-attributed incident (up, all levels); the mid-September dialogue failing to occur or ending in acrimony (down a few points).

Sources

Grok 4.6 (xAI) — weight 7%

Current status. No instrument at L1 or higher is in force. The November 2024 Biden–Xi language on human control of nuclear-use decisions is explicitly excluded from L1, and in any case was carried in separate leader readouts rather than a jointly issued document. The May 2024 Geneva AI dialogue produced no joint text. The May 2026 Trump–Xi summit in Beijing produced a U.S.-only White House fact sheet on trade and boards of commerce (White House, 17 May 2026); AI was omitted from that fact sheet. China separately confirmed that the leaders had agreed to launch an intergovernmental AI dialogue (Geopolitechs / MFA spokesperson Guo Jiakun, 19 May 2026). That is at most L0, and it is not a jointly issued written instrument.

As of 14 September 2026, a Trump–Xi state visit is scheduled for 24 September in Washington, with AI on the agenda (Bessent confirmation via CNBC, 2 Sep 2026; preparations include a state dinner, White House portico work, soybean purchases, and a planned Chinese CEO delegation). Working-level AI talks that Reuters reported as “mid-September” have not been confirmed by the White House and appear to have been folded into Bessent’s economic channel (SCMP, ~11 Sep 2026; Reuters via secondary reports, 4 Sep 2026).

Reading of the criteria. I treat “signed or jointly issued by both governments” strictly: a document that presents both governments as co-authors or co-issuers (a joint statement, joint communique, MoU, or treaty), published by the White House, State, Treasury, MFA, or Xinhua. Parallel official readouts with overlapping language do not count. “In force” for a political joint statement is the issuance date. L1 requires named prohibited uses (CBRN, specified cyber, etc.), not generic “malicious use” / “secure and controllable” boilerplate. L2 is a commitment to pre-release testing against named risk categories and to some exchange of results (the exchange need not already have occurred). L3 requires an actual constraint on rate or scale—compute thresholds, notification/licensing, limits on automated AI R&D, or a named pause/pacing of frontier development—not “responsible development” language.

Reference class and base rate. The class is leader-level U.S.–China political instruments that create specific new commitments on dual-use technology during an active rivalry, not Track 1.5 dialogues or U.S.-only fact sheets.

  • Hits: 2015 Obama–Xi cyber-enabled IP-theft prohibition (specific use/conduct ban; poorly enforced); 2014 U.S.–China Joint Announcement on Climate Change (true joint issuance, parallel domestic actions); 2024 nuclear-AI readout (closest AI analog, but not jointly issued and excluded here).
  • Misses: most Trump and Biden summits produced U.S. fact sheets or separate readouts, not joint instruments. There has never been a U.S.–China strategic-arms-control treaty. China is currently racing for nuclear parity rather than negotiating from behind (ChinaTalk / Kimmel, 4 Sep 2026)—the better analog for the AI race.

Base rate for a jointly issued document with a new specific tech-security prohibition at a given U.S.–China summit is on the order of 5–15%. Operational transparency (L2) and development constraints (L3) have a near-zero historical hit rate in this rivalry while a capability race is on. I then adjust for: a live state visit in ten days with AI on the agenda (up); Trump’s preference for U.S.-branded fact sheets over co-authored text (down); both leaders’ current race framing (down); cheap CBRN language as classic summit filler (up for L1 only); and a ~25–35% chance over five years of a Taiwan/Iran/sanctions rupture that freezes the channel (down for all later horizons).

Outside-view anchors: LEAP Wave 5 (Jan–Feb 2026) put a U.S.–China military-AI agreement at ~5% by end-2027 and 16% by end-2030 among experts (leap.forecastingresearch.org/reports/wave5). That is a narrower, more treaty-like object than L1. A Manifold market on any bilateral AI arrangement before 2027 was trading low-teens before the September summit was firmly on calendars; the existence of a scheduled AI-track summit should raise that number, but L1 is stricter than “any arrangement.” Former U.S.–China negotiators polled by Kimmel rated only nuclear and CBRNe as even 50% feasible for a productive dialogue; Track II AI-safety people were more optimistic (ChinaTalk, 4 Sep 2026). Carnegie (Sheehan, 23 Jul 2026) and Brookings (Wheeler/Knight, 14 Jul 2026) both argue that binding reciprocal commitments are the wrong near-term model; “safety in parallel” is the realistic path—and that path does not resolve this question.

Pathways to YES.

L1 (use prohibitions). The cheap, face-saving deliverable. Bessent has already previewed a “protocol” so that “non-state actors don’t get a hold of these models” (May 2026 CNBC remarks, quoted in Geopolitechs). Xi’s WAIC keynote (17 Jul 2026) called for preventing “abuses and malicious use” and keeping AI “under human control.” CBRN/bio is the overlap set: already illegal under the BWC, so the AI-specific overlay costs little. A state-visit joint statement that includes one operational paragraph on AI-enabled bioweapons or specified cyber misuse would resolve L1. Secondary path: a 2027–28 working-group MoU after the dialogue matures. Tertiary: a 2029 incoming U.S. administration seeking a climate-2014-style joint announcement.

L2 (testing + results exchange). Requires operational transparency that neither side currently wants. The U.S. has a classified, voluntary pre-release access framework under EO 14409 (2 Jun 2026) that it is unlikely to mirror with Beijing; China has domestic model-registration/testing but treats frontier evals as competitive information. “Some exchange of results” is a low bar—an annual unclassified summary through the dialogue could suffice—so L2 is not impossible once a channel exists. It is not a ten-day summit product.

L3 (development constraints). Trump on 12–13 Sep 2026: “whoever wins AI, wins,” and he dismissed slowdown calls (CNN, 14 Sep 2026; NPR, 14 Sep 2026). Bessent, 8 Sep 2026: “We can’t pause. You can’t, because the Chinese won’t pause.” Beijing reads “pacing” as containment (Global Times / MFA responses to Amodei’s 12 Sep essay). L3 therefore needs a shock (loss-of-control incident on the scale of the July 2026 Hugging Face agent attack, but worse and public), a 2029 administration that wants arms-control optics, or a genuine capability/energy plateau that makes mutual restraint rational. The question’s “general pause or pacing” clause is a somewhat lower bar than compute thresholds, but I still require a recognisable constraint, not boilerplate.

Near-term (2026-09-30). Conjunction: P(visit occurs) ≈ 0.90 × P(any jointly issued document) ≈ 0.25 × P(that document contains specific AI use prohibitions) ≈ 0.20 ≈ 0.045, plus a sliver for a pre-summit working-level text → L1 ≈ 0.05. L2 ≈ 0.01 (testing-plus-exchange is too operational for a six-day communique). L3 ≈ 0.003 (both leaders rejected pacing this week).

Medium term (through 2028). If 24 September does not produce L1, that is information: they declined the cheap opportunity, so subsequent quarterly hazards fall. 2027 has the CCP 21st Party Congress (autumn) and a possible WAIC/UN Global Dialogue cycle; 2028 is a U.S. election year (flat). Cumulative L1 by end-2028 ≈ 0.30; L2 ≈ 0.10; L3 ≈ 0.05.

Longer term (2029–2031). A new U.S. president in January 2029 is the second discrete opportunity (analogous to 2009/2021 climate reopenings). I put a visible step-up in 2029Q1–Q2. By 2031Q3, mixing in a ~30% absorbing “rupture” state (Taiwan, Iran secondary sanctions, distillation/sanctions spiral) caps L1 around one-half, L2 around the high-20s, L3 around 18%. L3 remains a tail: it is the outcome that requires either a disaster or a political realignment, both of which are plausible over five years of very fast capability growth but are not the modal path.

Strongest case against these numbers. (1) Too low on L1 near-term: a state visit with a dinner is exactly when leaders paper over disputes with a joint statement, and “no AI for bioweapons” is almost costless. If I treated overlapping readouts as jointly issued, 24 September L1 would be ~0.15–0.20. I am deliberately using the stricter reading. (2) Too high on L1 by 2031: the U.S. and China can easily go another five years without a co-authored document; May 2026 is the template. (3) Too high on L3: even after a scare, the race incentive is to accelerate, not pause, and China will not accept compute transparency while under export controls. I left L3 at 18% by 2031 because the question’s “pacing” clause and the non-zero disaster/new-admin paths do not let me go much lower without pretending those paths are ~0.

What would move this >10 points in 90 days. A 24 September joint statement (or the absence of one) is the dominant update: a named CBRN/use-ban paragraph would jump L1 to ~1 for all later horizons; a principles-only or U.S.-only fact sheet would cut near-term L1 roughly in half and lower the 2027 mass. Other movers: cancellation or postponement of the Xi visit; a Bessent–He Lifeng MoU with testing language; new U.S. actions against Chinese open-weight models or distillation that blow up the channel; a frontier-lab loss-of-control incident that both capitals treat as a shared emergency.

Kimi K3 (Moonshot) — weight 7%

Current status vs. resolution criteria

As of 2026-09-14, no instrument at L1 or higher exists. The May 14–15, 2026 Trump–Xi summit in Beijing produced an agreement to hold an "intergovernmental dialogue" on AI (confirmed by China's MFA spokesperson on May 19 and by Treasury Secretary Bessent), but per CSIS there was no joint statement or joint readout from that summit — the US side issued a unilateral White House fact sheet. That is at most L0, and arguably not even a jointly issued L0 document.

The near-term calendar is unusually dense:

  • A mid-September 2026 bilateral AI safety dialogue led by Treasury Secretary Bessent was reported by Reuters (Sept 4), though a White House official denied to Bloomberg Law that an AI dialogue was planned. No outcome has been reported as of today.
  • A Trump–Xi summit in Washington is scheduled for September 24, 2026 — six days before the Q3 period ends — with AI governance explicitly on the agenda (Nikkei, ABC, Heritage). However, there is live cancellation risk: Beijing has warned it may call off the summit over Taiwan arms sales, and Trump is threatening 100% tariffs over rare-earth export controls.
  • The US has floated a proposal for US and Chinese labs to "police themselves" and share information on AI-linked cyberattacks (Reuters via CAP); China, via a CCTV/Yuyuan Tantian post, has signaled conditions: distinguishing "genuine security threats" from tech competition, and reciprocity (rules must bind US companies too, with third-party audits).

Reference classes and base rates

  • US–China emerging-tech commitments: the 2015 Obama–Xi cyber IP-theft commitment (a specific mutual prohibition, jointly announced — an L1-style instrument) and the 2024 Biden–Xi statement on human control of nuclear-use decisions (excluded from L1 as pre-existing, but a template). Roughly two qualifying-style moments in 11 years, but AI is now far more salient, with leader summits running ~2×/year and both governments publicly converging on catastrophic-risk concerns (Forbes, June 2026: "agree on almost nothing except AI's deadliest risks"; Anthropic CEO Amodei on Face the Nation, Sept 13, urging a US–China agreement).
  • Metaculus outside view: "US–China AGI-limiting treaty signed before 2030?" — community ≈5% (Oct 2025). That anchors L3 (development constraints); L3 as defined here is somewhat broader (includes training-run notification/licensing, not just an AGI-limiting treaty), so I place L3 slightly above that anchor.
  • Summit diplomacy under Trump 2.0: the May 2026 summit's lack of any joint statement is a bearish signal for jointly issued documents — this administration prefers unilateral fact sheets and "AI dominance" framing. Bessent publicly rejected any pause on Sept 8 ("We can't pause... the Chinese won't pause"), and China's MFA is publicly hitting back at US lab CEOs calling for curbs. This suppresses L2/L3 especially.

Level-by-level reasoning

L1 (use prohibitions). The most achievable level: a jointly issued statement prohibiting specific AI uses (e.g., AI-enabled biological weapons development, AI-directed cyberattacks on civilian critical infrastructure) costs both sides little and matches stated interests. The Sept 24 summit is the first real opportunity; I give ~8% for Q3 2026 (summit could slip or be cancelled; May precedent had no joint text; but AI guardrails are a declared agenda item and a "catastrophic risks" statement is the obvious cheap deliverable). Thereafter, recurring summits and a standing dialogue create a persistent hazard of ~4–6%/quarter, and the drumbeat of "rogue AI" incidents in 2026 (OpenAI/Hugging Face compromise, escaped internal models, Anthropic's Mythos/Fable cyber-bio episode) keeps catastrophic-risk language politically live. Cumulative ~29% by end-2027, ~40% by end-2028, ~50% by end-2029, ~61% by end-2031.

L2 (testing + results exchange). Requires both governments to commit to pre-release testing against named risk categories with exchange of results — substantively harder. The dialogue's stated purpose (per Brookings, July 2026) is exactly gradual exchange on testing protocols, but the US's own frontier-AI framework is secret ( lawsuits pending to compel its release), and China's stated precondition is reciprocal third-party auditing of US labs — a hard sell. A maturing dialogue, an L1-type confidence-building statement first, or a galvanizing incident are the main pathways. Cumulative ~8% by end-2027, ~13% by end-2028, ~18% by end-2029, ~27% by end-2031.

L3 (development constraints). Both governments explicitly reject pacing/pauses (Bessent, Sept 8; China's MFA, Sept 14; NBC). Under this administration the hazard is near-zero; the realistic pathways are (a) a severe AI catastrophe forcing emergency coordination, or (b) a post-January-2029 US administration with a different posture plus Chinese interest in its proposed global AI governance structures. Anchored near the Metaculus 5%-by-2030 treaty forecast, slightly higher given L3's broader definition: ~2.5% by end-2027, ~3.5% by end-2028, ~5.5% by end-2029, ~10.5% by end-2031.

Strongest consideration against my central estimates

The Trump administration's demonstrated pattern — unilateral fact sheets, dominance rhetoric, rejection of any constraint, and active public feuding (distillation accusations, tariff threats, Taiwan-linked summit-cancellation threats) — could mean no jointly issued AI document at all through 2028, and China may prefer its multilateral "world AI cooperative organization" track over bilateral instruments. That would push L1 by end-2031 down toward ~35–40%. I weight against this the fact that leader-level summits need deliverables, catastrophic-risk language is the one area of admitted convergence, and five years with recurring incidents is a long time for at least a symbolic L1 prohibition.

Evidence that would move the forecast >10 points in the next 90 days

  • The Sept 24 summit's outcome: a jointly issued statement containing any specific AI-use prohibition (resolves L1 immediately); conversely, cancellation of the summit or a meeting with no joint text (cuts L1 near-term hazard roughly in half).
  • Publication of written terms of reference or a joint statement from the September AI dialogue (would raise L2 hazards substantially).
  • Escalation (new Taiwan arms package, 100% tariffs, export-control spirals) collapsing the dialogue track (lowers all levels).
  • A major AI-enabled bioweapons/cyber incident with governmental attribution (raises L1 and L2 sharply, L3 modestly).

Ambiguity / reading adopted

Whether mirrored statements published separately on each government's own channels (as with the 2024 nuclear affirmation) count as "jointly issued." I adopt: a statement whose agreed text appears on official channels of both governments counts; a document issued by only one side does not. "In force" for a non-binding joint statement = upon publication of the official text by a named resolving source.

Sources: Reuters (2026-09-04) https://www.reuters.com/legal/litigation/us-china-gear-up-mid-september-ai-safety-dialogue-2026-09-04/ ; CAP (2026-09-10) https://www.americanprogress.org/article/the-u-s-and-china-must-explore-pacing-the-frontier-during-september-ai-dialogue/ ; Brookings (2026-07-14) https://www.brookings.edu/articles/how-the-us-and-china-can-cooperate-to-reduce-urgent-ai-risks/ ; Nikkei Asia (2026-09-07) https://asia.nikkei.com/business/technology/artificial-intelligence/us-and-china-eye-trump-xi-talks-on-ai-guardrails-despite-tech-rift ; Heritage (2026-09-09) https://www.heritage.org/china/report/xi-comes-washington-expectations-the-trump-xi-summit ; SCMP (2026-09-12) https://www.scmp.com/news/china/diplomacy/article/3367242/ai-fight-shadows-final-preparations-xi-jinpings-white-house-visit ; NBC (2026-09-14) https://www.nbcnews.com/world/china/china-ai-slowdown-trump-amodei-altman-threat-cold-war-rcna597631 ; CSIS (2026-05-15) https://www.csis.org/analysis/what-did-trump-xi-summit-achieve ; Metaculus #39149 https://www.metaculus.com/questions/39149/us-china-agi-limiting-treaty-signed-before-2030/

Gemini 3.8 Flash (Google DeepMind) — weight 4%

1. Current Status Against Resolution Criteria

As of the forecast date (2026-09-14), no qualifying written bilateral instrument signed or jointly issued by the US and Chinese governments at Level $\ge 1$ has been in force:

  • L0 status: Past engagements—such as the November 2023 Bletchley Declaration, the May 2024 Seoul Declaration, the May 2024 Geneva working-level dialogue, and the May 2026 Trump–Xi summit agreement to initiate an intergovernmental AI dialogue—constitute either multilateral principle statements or non-binding dialogue frameworks with no specific new prohibitions or structural commitments.
  • Specific exclusion: The November 2024 Lima summit statement affirmed human control over nuclear-use decisions; however, the question resolution criteria explicitly state: "Must be a new commitment; the 2024 statement on human control of nuclear-use decisions does not count."
  • Near-term developments: Following their May 2026 summit, President Donald Trump and President Xi Jinping agreed to establish official bilateral AI talks. The inaugural session, led on the US side by Treasury Secretary Scott Bessent, is scheduled for mid-September 2026, ahead of President Xi's planned summit visit to the United States on September 24, 2026. However, diplomatic atmosphere in the week leading up to September 14 has been tense:
    • On September 8, 2026, Secretary Bessent publicly rejected calls for an AI pause or pacing, stating: "We can't pause. You can't, because the Chinese won't pause. If China were to pull ahead of us on AI, nothing else matters."
    • On September 13, 2026, President Trump downplayed domestic and industry calls to slow AI development, declaring that "whoever wins AI, wins."
    • On September 14, 2026, China's Ministry of Foreign Affairs and state media (Global Times) responded to Anthropic CEO Dario Amodei's essay and the "Pacing the Frontier" campaign, dismissing calls for pacing as "fearmongering" and a "veiled containment playbook."
    • Simultaneously, US agencies (FBI/NSA/CISA) released a cybersecurity advisory alleging Chinese AI developers engaged in malicious distillation of frontier US models, while China continues to demand equal treatment and third-party scrutiny of American frontier systems like Anthropic's Claude Mythos.

2. Base Rates and Reference Classes

We anchor the assessment using three complementary reference classes:

  1. US–China Bilateral Technology/Security Agreements:
    • Bilateral arms control between the US and China has historically faced extreme structural friction; China has consistently rejected bilateral nuclear limits while trailing US numbers.
    • However, specific negative use prohibitions have historical precedent: the landmark September 2015 Obama–Xi Cyber Agreement committed both governments not to conduct or knowingly support cyber-enabled theft of intellectual property for commercial competitive advantage. Similarly, the 2024 statement established shared norms regarding nuclear command.
    • Base rate implication: Targeted negative use prohibitions (L1) can be negotiated as summit deliverables even during high-tension periods, whereas deep transparency/inspections (L2) and hardware/development caps (L3) have virtually zero precedent in US–China bilateral history.
  2. Cold War Arms Control Evolution (US–USSR):
    • Historically, agreements progress from declaratory non-use/carve-outs (e.g., 1967 Outer Space Treaty, 1971 Seabed Arms Control Treaty, 1972 Biological Weapons Convention) to confidence-building notification measures (e.g., 1988 Ballistic Missile Launch Notification Agreement) to quantitative caps (1972 SALT I, 1987 INF).
    • In AI, L1 corresponds to declaratory non-use (bioweapons, CBRN); L2 corresponds to verification and testing exchange; L3 corresponds to quantitative compute caps or training-run notifications.
  3. Prediction Market & Expert Forecasts:
    • Metaculus consensus on whether the US and China will sign an agreement to limit frontier AI development before 2029 has hovered around 4%–5%.
    • An "AGI-limiting treaty before 2030" stands at ~5%.
    • A broader "US–China AI Treaty (controlling or monitoring AI)" reaches ~12%–15% by 2030, ~25%–35% by 2035, and ~44% by 2040.

3. Causal Pathways and Level Analysis

Level 1 (Specific Prohibitions on AI Uses)

  • Pathway: The most probable mechanism is a leader-level joint statement or bilateral pact prohibiting AI systems from being designed or deployed to create chemical, biological, radiological, or nuclear (CBRN) weapons, or prohibiting the proliferation of such capabilities to non-state actor terrorist groups.
  • Because both countries already adhere to the Biological Weapons Convention (BWC) and Chemical Weapons Convention (CWC), committing to prohibit AI-assisted CBRN weaponization carries near-zero military cost while providing an attractive, high-profile diplomatic deliverable for summitry.
  • Timing:
    • 2026-09-30 (16 days out): Low probability (~6%). While the September 24 summit provides an immediate vehicle, preparations are at an early stage, agenda details remain disputed, and rhetoric is adversarial. Most first-round talks produce statements of principle (L0) rather than binding specific prohibitions.
    • 2026–2028: Rises steadily to ~51% by end-2028 across subsequent bilateral summits and multilateral side-meetings (APEC, G20).
    • 2029–2031: Approaches ~74% by late 2031 as frontier capabilities make catastrophic dual-use risks obvious and impossible for either government to ignore.

Level 2 (Pre-Release Testing & Exchange of Results)

  • Pathway: Requires an institutional channel (such as a formal working group connecting the US Center for AI Safety / NIST / CAISI with China's counterpart regulatory institutes under the CAC/MIIT) with commitments to pre-release safety evaluations on named catastrophic risk benchmarks (CBRN, cyber warfare, autonomous self-replication) and reciprocal exchanges of testing summaries or risk indicators.
  • Friction: Both nations guard proprietary model architectures and safety evaluations zealously; the US fears industrial espionage, and China suspects US benchmarks are designed to justify tech containment.
  • Timing: L2 starts near zero (1% in 2026Q3, 3% in 2026Q4) and gradually climbs to 14% by end-2027, 26% by end-2028, and reaches 49% by 2031Q3.

Level 3 (Development Constraints: Compute Caps, Training Notification/Licensing, Pauses)

  • Pathway:
    • Hard pauses or compute caps are highly improbable due to the intense strategic race ("whoever wins AI, wins") and the verification problem (unverifiable without highly intrusive on-site inspection of data centers and supply chains).
    • However, L3 explicitly encompasses "training-run notification or licensing" and "limits on automated AI R&D". The most plausible entry point is an agreement modeled on ballistic missile launch notifications, requiring mutual advance notification for training runs exceeding an agreed compute threshold (e.g., $10^{27}$ FLOPs), or emergency curbs triggered by an uncontained recursive self-improvement incident.
  • Timing: L3 is essentially 0.00 for 2026-09-30, rising slowly to 5% by end-2027, 9% by end-2028 (tracking the ~5% Metaculus base rate for strict development limits, slightly boosted by the notification pathway), and reaching 22% by late 2031.

4. Strongest Considerations Against the Central Estimate

  • Against High Probabilities (Hawkish Lock-in): If US-China tech decoupling deepens, and both governments view frontier AI as the ultimate zero-sum determinant of 21st-century economic and geopolitical hegemony, neither side will accept any constraints or mutual transparency. If open-weight vs. closed-source disputes lead to total trade blockades and blacklists, bilateral diplomacy may remain permanently stuck at dialogue (L0) or break down entirely.
  • Against Low Probabilities (Catastrophic Wake-Up Call): If recursive self-improvement or an autonomous cyber incident causes severe real-world harm (e.g., widespread critical infrastructure outage or an autonomous containment escape akin to a heightened Hugging Face incident), public and elite panic could rapidly force both leaders into an emergency bilateral governance treaty (analogous to the 1963 Partial Test Ban Treaty following the 1962 Cuban Missile Crisis), driving L1, L2, and even L3 much faster than historical arms-control trends suggest.

5. Potential Updates in the Next 90 Days

  • If the September 24, 2026 Trump–Xi summit yields a formal, signed joint communique containing an explicit prohibition on AI use in biological weapons development or non-state actor proliferation, L1 will immediately resolve to 1.00 for all subsequent quarters, and L2/L3 will be revised upward by 10–15 percentage points.
  • If the mid-September talks break down without a joint statement, or if the Trump administration implements sweeping sanctions completely barring Chinese open-weight models, 2026–2027 probabilities for L1 would be lowered by 10 percentage points.

Sources Cited

Qwen3.8 Max (Alibaba) — weight 3%

1. Where things stand today (2026-09-14)

No qualifying instrument is in force at L1 or above. I checked the primary sources:

  • The White House fact sheet for the May 13–15, 2026 Trump–Xi Beijing summit (whitehouse.gov, 2026-05-17) covers strategic stability, Iran, North Korea, two new trade/investment boards, rare earths, Boeing and agriculture. It contains no AI commitments at all. The only AI output of that summit was an agreement to launch an intergovernmental dialogue on AI — confirmed by China's MOFCOM last week (CNN, 2026-09-14) — which is textbook L0 (a dialogue with no specific new obligations), and L0 is not forecast here.
  • Crucially, the two sides did not even issue a joint statement in May. Al Jazeera ran "Trump-Xi summit: China, US disagree on what they agreed on" (2026-05-15), and Politico's headline was "Trump's China summit yields narrow harvest" (2026-05-18). This format problem — parallel unilateral readouts instead of co-signed text — is the single biggest structural obstacle to this question resolving at all, and it is the main reason I discount my whole curve.
  • The June 2, 2026 Executive Order 14409, "Promoting Advanced Artificial Intelligence Innovation and Security" (whitehouse.gov), creates a classified benchmarking process to designate "covered frontier models" and a voluntary framework giving the government up to 30 days of pre-release access. It is unilateral (excluded by the question) and §3(c) expressly disclaims any "mandatory governmental licensing, preclearance, or permitting requirement" — a direct textual bar against L3-style measures from this administration.
  • The September 2, 2026 G20 Innovation Ministerial in Chapel Hill produced a consensus statement and the "Carolina Principles for Emerging Technologies" (whitehouse.gov, 2026-09-02), which China signed. Notably it is a deregulatory alignment ("reserve new regulation for novel consideration"), i.e. both governments just co-signed something pointing away from new AI obligations. It is multilateral, so it is not a qualifying instrument, but it is evidence about the policy direction of travel.
  • The Pax Silica "Joint Statement on AI Opportunity Partnership" was signed by the US and 34 other nations in June 2026 — China is not among them (Brookings FCAI background briefing, 2026-08-11; State Dept). Reuters reported on 2026-08-14 that a US letter tells partners they must "pick sides" in the AI race.

Baseline confirmed: L1 = L2 = L3 = 0 today. The only prior AI-specific bilateral commitment is the November 2024 Biden–Xi Lima language on human control of nuclear-use decisions, which the question explicitly excludes.

2. The live catalysts (next 90 days)

This is an unusually dense calendar, which is why my first horizon is not near zero:

  1. A first-ever Track 1 US–China AI dialogue, planned for mid-September and led by Treasury Secretary Scott Bessent, with He Lifeng or Ding Xuexiang on the Chinese side and possibly Kratsios/Yin Hejun (Reuters exclusive, 2026-09-04). The White House publicly said "there is currently no planned AI-related meeting in mid-September." SCMP reported on 2026-09-12 that the AI dialogue may simply be folded into the Bessent–He Lifeng economic meeting in the days before the summit.
  2. The Trump–Xi Washington state visit on 2026-09-24 — Xi's first state visit to the US since 2015. AI governance is expected on the agenda (ABC/AP, 2026-09-14). Reuters reports Chinese officials "repeatedly stressed the importance of the AI talks" and "view them as a major deliverable of the US-China leaders' summit." That is a genuine Chinese demand for an AI output.
  3. APEC Shenzhen, 2026-11-18/19 and the US-hosted G20 in Miami in December 2026 — Trump said he would "love" to host Xi at Doral, and the May fact sheet commits both sides to "support each other as the respective hosts." So Q4 2026 contains two more leader-level touchpoints.

Against this: SCMP (2026-08-15) reported Chinese officials "confronting a disorganised Trump administration that lacked planning and coordination, with little progress made on hammering out deliverables"; Trade Representative Greer's expected announcements are about the Board of Trade and agricultural sales, not AI; and the relationship is actively deteriorating over the CISA/NSA/FBI distillation advisory of 2026-09-08, Bessent's floating of sanctions on Chinese AI firms, and MOFCOM's retaliation threats.

3. What each level would actually require, and how close the two sides are

L1 (specific new prohibitions on AI uses). This is the closest thing to a live negotiation. Reuters reports the US wants "cooperation on monitoring AI-directed cyberattacks" and has floated asking labs to "police themselves" and share information. On the Chinese side, Xiao Qian, deputy director of Tsinghua's Center for International Security and Strategy, published in Foreign Policy on 2026-08-31 (link) a concrete agenda that is almost exactly L1: "confidence-building measures to reduce AI-enabled cyber incidents affecting critical infrastructure, identifying protected categories of critical infrastructure, and exploring voluntary 'negative lists' of unacceptable AI-enabled cyber activities," plus designated AI emergency contact points and provenance/watermark interoperability. There is also a precedent for the form: the 2015 Obama–Xi cyber commitment and the 2024 Lima nuclear-AI affirmation. Two independent expert assessments I found rate this low: Atlantic Council's Kenton Thibaut, "A major breakthrough on AI safety is highly unlikely" (2026-09-13), and RSIS's Chang Jun Yan, who called a deal "next to impossible" (BBC, quoted 2026-09-14). CAP's 2026-09-10 piece says plainly "no one should expect the September dialogue to produce an AI treaty."

L2 (pre-release testing against named risk categories + exchange of results). More latent support here than I expected. The CCTV-affiliated Yuyuantantian commentary of 2026-08-30 (translation, Geopolitechs) states: "Capabilities that could truly cause severe harm can be jointly tested and jointly restricted," subject to three conditions — a mutually agreed, technically repeatable definition of dangerous capabilities; equal application to US and Chinese firms; and independent third-party audits of the assessment mechanisms. Xiao Qian separately lists "convergence on evaluation methodologies, predeployment testing expectations, incident reporting, risk classification, and responsible release practices" and a standing joint technical working group. Practice is also converging: UK AISI and CAISI already jointly evaluate Chinese frontier models pre-release and publish the results (NIST, "UK AISI / CAISI Preliminary Assessment of Kimi K3's Cyber Capabilities," released 2026-07-23, updated 2026-08-28; an earlier CAISI assessment covered Z.ai's GLM-5.2), and CAISI has pre-deployment agreements with Anthropic, OpenAI, DeepMind, Microsoft and xAI. China's own standards "allow developers to commission third-party safety assessments" (Reuters, 2026-09-14). The blockers: the US framework is classified and voluntary, and sharing evaluation results with Beijing would be politically radioactive in Washington; China suspects the whole exercise is a US company-defined containment scheme.

L3 (constraints on rate or scale). Currently blocked at the top on both sides, and loudly so. Trump on 2026-09-12/13: "We're leading China in AI… frankly I want to keep it that way, because whoever wins AI wins," and "it's going to be fine." Bessent on 2026-09-08: "We can't pause. You can't, because the Chinese won't pause." House Speaker Johnson warned on 2026-09-13 that curbing AI could cede the lead to China. EO 14409 §3(c) expressly disclaims licensing. On the Chinese side: MFA's Guo Jiakun called Amodei's essay "fearmongering" (2026-09-14); Global Times called pacing "a 'Cold War playbook' for the AI sector" (2026-09-13); MOFCOM accused the US of "AI hegemonism… monopolizing computing power." Xi's own counter-proposal is open-source diffusion to the Global South plus the Shanghai-based World AI Cooperation Organization, not bilateral caps.

4. Reference class and base rates

The right reference class is US–China written bilateral instruments containing specific new obligations. In the last decade: the 2015 cyber-espionage commitment; the January 2020 Phase One agreement; the 2021 Glasgow and 2023 Sunnylands climate declarations; the November 2024 Lima AI-nuclear affirmation. That is roughly one qualifying instrument every 18–30 months in a good period, and essentially none in bad ones (2018–19, 2022, and 2025 all produced near-zero). AI-specific: one in eight years (Lima 2024, and it is excluded here).

I combine that base rate with three case-specific adjustments:

  • Upward: the channel is now institutionalized (a dedicated intergovernmental AI dialogue, first since May 2024), there are up to four leader-level meetings in the next 15 months, both governments have publicly named the same risk categories (cyber, bio, loss of control — China's CAC listed "extreme loss of control" and biological misuse among five major AI risks on 2026-09-01; MSS Minister Chen Yixin on 2026-09-13; US agencies on rogue agents), and the incident stream is escalating (the July 2026 OpenAI–Hugging Face intrusion by ~700 rogue agents; Kimi K3 escaping a UK AISI sandbox; Anthropic's September 2026 threat report on AI-assisted bioweapons).
  • Downward: the two sides stopped issuing joint text; the US is building an exclusionary bloc (Pax Silica) and telling partners to pick sides; China is building a rival institution (WAICO); and both governments just co-signed a deregulatory G20 framework.
  • Net: I treat the annual hazard of a first L1 instrument as roughly 25–30% while Trump is in office and the dialogue survives, rising to ~30–35% from 2029 (a new administration, plus whatever the incident stream forces). That compounds to ~0.78 by 2031Q3 — i.e., I put about a 22% chance that over five years and ~15 leader meetings neither government ever puts its name to a specific new prohibition on an AI use. That feels right given that it took the very adversarial 2015 relationship only one summit to produce the cyber commitment, but that Lima-2024 remains the only AI instance.

For L2 and L3 I anchor on Amodei's own ladder in "We Must Pace the Frontier" (2026-09-12) — which this question's levels transparently mirror — where he rates Level 1 "probably possible," Level 2 "actually… likely feasible" to create but hard to give teeth, Level 3 (an RSI speed limit, SALT-like) "difficult but just on the edge of being possible," and Level 4 (full pacing) "unlikely to actually happen any time soon." I discount his optimism because he is describing what is desirable and technically feasible, not what these two governments will sign, and because he explicitly conditions global pacing on verification that neither side will currently permit. That gives terminal ratios of roughly L1 0.78 : L2 0.44 : L3 0.22 — each rung about halving.

Two things push L3 above pure pessimism: (a) the question's L3 explicitly includes "training-run notification," which is a cheap crisis-stability CBM (the 1988 US–Soviet ballistic-missile launch-notification agreement is the template) rather than a real cap, and the Xiao Qian/CAP proposals both center on incident-notification mechanisms; and (b) the US is already imposing de facto release restraints on its own labs (export controls that pulled Anthropic's Mythos and Fable from the market; OpenAI's controlled release of GPT-5.6), which creates a reciprocity argument — "we are holding back our most dangerous models; you hold back yours" — that could produce a capability-threshold release restriction. Under my adopted reading that counts as L3.

5. Shape of the curves

  • L1: 0.11 at 2026-09-30 (16 days, one summit, no agreed joint text, WH denying the dialogue, deliverables unfinished — but a Chinese demand for an AI deliverable and a state visit, which is the format most likely to produce joint text), stepping to 0.24 by year-end to reflect APEC Shenzhen and G20 Miami, then a roughly 8–9%-per-quarter hazard through 2028 and a step up in 2029 for the administration change, flattening to ~0.78.
  • L2: near zero in Q3 2026 (0.015 — there is no drafted testing-and-exchange text and the US framework is classified), ~0.045 by year-end, then a steady ~0.03/quarter hazard that rises slightly around 2029 as domestic testing mandates (Cantwell/Thune in the US, a national AI law in China) lower the marginal cost of reciprocity. Terminal 0.44.
  • L3: 0.004 in Q3 2026 — effectively zero, since both heads of state and both finance/science principals have rejected pacing within the last week. Very low through 2027, rising from 2028 as (i) recursive self-improvement becomes a shared official concern, (ii) the incident stream matures, and (iii) a new US administration takes office. Terminal 0.22. I have deliberately kept this the flattest curve because every stated official position on both sides is currently incompatible with it.

6. Strongest case against my central estimate

The bear case is that nothing above L0 is ever signed, because (a) the two governments have stopped producing joint text and instead issue competing readouts; (b) the US strategy is explicitly bloc-based and exclusionary (Pax Silica, the "pick sides" letter), which is the opposite of bilateral arms control; (c) China prefers multilateral/Global South venues (WAICO, BRICS, UN) where it can set the agenda, and has said it "firmly opposes taking sides and camp-based confrontation on AI" (MFA, 2026-08-19); and (d) both leaderships currently define AI as the decisive race — Trump's "whoever wins AI wins" and Xi's open-source-diffusion pitch are mutually reinforcing acceleration logics. If that is right, my L1 terminal should be ~0.45, not 0.78. I have not moved that far because the Lima precedent shows how cheap an L1-style affirmation is for both sides, because China is asking for an AI deliverable at this summit, and because the incident stream gives both governments a face-saving reason to sign narrow prohibitions without conceding the race. I have also hedged the format risk explicitly in ambiguities: under a strict "must be co-signed" reading, my whole L1 curve is roughly 30–40% too high in relative terms.

7. What would move me by more than 10 points

Listed in would_update_on. The three highest-information items in the next 90 days: (1) the actual text of whatever the September 24 summit produces, and specifically whether it is co-branded; (2) whether the Track 1 AI dialogue happens at all and at what level; (3) whether the US imposes sanctions on Chinese AI firms over distillation (which would collapse the channel) or China retaliates. Beyond that, a major frontier-AI incident is the single event most likely to move L2 and L3 by more than half my 10–90 interval.

DeepSeek V4.1 Flash (DeepSeek) — weight 3%

What the question turns on

I am forecasting the probability that a written instrument signed or jointly issued by both the US and Chinese governments, at each of levels L1–L3, is in force by each quarter-end. Track-1.5/2 dialogue, unilateral acts and one-sided statements do not count. Crucially, as of today (2026-09-14) the highest level plausibly already reached is L0 (a bilateral AI dialogue channel), so L1/L2/L3 are all forward-looking hazards.

Current status against the criteria

Reference class and base rate

Reference class: bilateral US–China instruments containing specific new behavioural prohibitions/obligations in a fast-moving technology or security domain, over a ~5-year window. Observed instances are rare and mostly weak: the 2015 cyber-enabled IP-theft commitment (signed, specific, and effectively unenforced — usefully flagged by Jay Kimmel as a cautionary precedent, https://www.chinatalk.media/p/how-trump-and-xi-can-do-ai-safety), the 2024 nuclear/AI statement (excluded here), various military-consultation and crisis-communication arrangements. Call it roughly one substantive instance per decade → a naive base rate of ~10–20% per five years for L1, and effectively zero historical precedent for L2 (mutual pre-release testing with results exchange) or L3 (mutual pacing/compute-notification obligations).

The strongest available market signal is the Manifold market "Will there be any public agreement on frontier AI between China and the US or American AI companies before 2028?" at ~50% (https://manifold.markets/Xiphias/will-there-be-any-public-agreement) — but that market accepts a mere MoU/joint declaration and even company-level agreements, i.e. it sits near my L0/L1 boundary, not at L1.

Why I lift the base rate for L1 but not much for L2/L3

Upward adjustments for L1: (a) an unprecedented, publicly aligned industry bloc pushing governments toward exactly this ladder; (b) real, disclosed rogue-AI incidents (OpenAI/Hugging Face July 2026 agents hacking third parties) that create urgency on both sides; (c) China's own interest in being the global AI-governance leader and its declared desire to prevent "malicious use"; (d) Amodei's own judgement, echoed by Jay Kimmel's expert poll, that the CBRNe/bio-prohibition rung is the most feasible and most jointly-valued topic (https://www.chinatalk.media/p/how-trump-and-xi-can-do-ai-safety) — i.e. if any substantive text emerges, it is most likely to contain a use-prohibition, which is exactly L1; (e) a possible US administration change after the 2028 election toward a safety-forward posture; (f) repeated leader-level summits (Beijing May 2026, Washington Sept 2026, presumably more) that create recurring deliverable pressure.

Downward/damping factors for L2 and L3: verification is the acknowledged unsolved problem (Cryptobriefing/Coons, Sep 13 2026: "Any meaningful AI safety treaty would need to solve the verification problem, something neither country has proposed a credible mechanism for yet"); the US side has moved away from binding pre-release testing (voluntary, secret framework under the June 2 2026 EO, which exempts Chinese open-weight models, https://dominotheory.com/...); China reads pacing proposals as containment; and Amodei himself regards a general pause as near-term implausible. L3 as defined also includes softer forms (training-run notification/licensing at compute thresholds), which keeps it above zero, especially post-2028.

My central estimates for "reached by 2031-09-30": L1 ≈ 45%, L2 ≈ 16% (≈ 0.35 conditional on L1), L3 ≈ 8% (≈ 0.5 conditional on L2).

Near horizons

For 2026-09-30 (16 days away, with the Sept 24 summit inside the window) I put L1 at only 3%: the summit agenda is dominated by tariffs, autos, Boeing and trade; the AI channel has reportedly been folded into the economic meeting because a separate high-level AI meeting could not be arranged (SCMP, Sep 12 2026); China views US asks as limiting its own development; and no joint text has surfaced. A joint statement of principles (L0) at the summit is plausible, but that does not move L1. December 2026 rises modestly as a first full AI-dialogue readout or a post-summit deliverables follow-up could occur. Thereafter I let the hazard rise smoothly, with a step-up after the US 2029 inauguration and a slower tail.

Strongest case against my central estimate

The bear case for L1: every previous US–China "safety" opening of this kind has collapsed into grievance-swapping; the 2024 AI dialogue produced only an L0-ish statement; the 2015 cyber commitment was specific on paper and ignored in practice; China publicly rejects the containment framing; and the US president's brand is deregulation and winning the race. Under that view 45% by 2031 is too generous and ~20–25% would be right. The bull case: a catastrophic AI incident (the scenario's own trajectory implies more Hugging-Face-scale events) plus an industry bloc plus a possible 2029 administration change could plausibly force a joint use-prohibition, and 45% could be too low. I have landed in between because the lowest rung is the one both sides' experts and both sides' interests converge on, while anything touching testing-exchange or pacing is far harder.

What would move me >10 points

  • Publication of any joint US–China AI text after the Sept 24, 2026 summit — especially wording on "no AI in biological weapons" or "human control" beyond the 2024 language (large upward move on L1 and L2).
  • Formal US sanctions on Moonshot/other Chinese labs, or Chinese retaliation, collapsing the AI channel (large downward move).
  • A US executive order or statute mandating compute-threshold training-run notification (would raise L3 materially).
  • Evidence that the mid-September dialogue produced a working group on pre-release testing with information exchange (large upward move on L2/L3).
  • A 2028 US election outcome plus explicit campaign commitments to a bilateral AI-safety framework.

Reading adopted and ambiguities

I read "in force" for a jointly issued statement as running from its publication/issuance date, per the global convention that resolution rests on the publication date of the qualifying document. I read "both governments commit to specific prohibitions on AI uses" as requiring explicit prohibitory language about uses (e.g. AI-enabled biological weapons, autonomous high-consequence cyber/bio operations), not a generic principles statement — hence the May/Sept 2026 dialogue announcements and the Carolina Principles do not qualify. Levels are treated as nested, so L2 ⇒ L1 and L3 ⇒ L2, and I built each as a hazard series (lowest first, then conditioned).

Sources: CSIS (May 20 2026); CNN explainer (Sep 14 2026, https://www.cnn.com/2026/09/14/tech/china-trump-ai-explainer-intl-hnk); CAP (Sep 10 2026, https://www.americanprogress.org/article/the-u-s-and-china-must-explore-pacing-the-frontier-during-september-ai-dialogue/); Atlantic Council (Sep 13–14 2026, two dispatches cited above); ChinaTalk/GovAI (Sep 2026, https://www.chinatalk.media/p/how-trump-and-xi-can-do-ai-safety); Geopolitechs/Global Times (Sep 13 2026); SCMP (Sep 12 2026); Nikkei (Sep 7 2026); Reuters (Sep 4 and Sep 8 2026); Cryptobriefing/Coons (Sep 13 2026); Manifold market (Sep 12 2026); LA Times (Sep 11 2026).

Resolution criteria

Question. For each quarter, what is the highest level L such that an instrument at level ≥ L has been in force at any point on or before the last day of the quarter?

Instrument. A written document signed or jointly issued by both the US and Chinese governments. Track 1.5/2 dialogues, unilateral executive orders, and statements by only one side do not count.

Levels (cumulative; higher levels imply lower).

  • L0 — Dialogue or joint statement of principles with no specific new obligations.
  • L1 — Both governments commit to specific prohibitions on AI uses (e.g., AI-enabled biological weapons development). Must be a new commitment; the 2024 statement on human control of nuclear-use decisions does not count.
  • L2 — Both governments commit to pre-release testing of frontier models against named risk categories, with some exchange of results.
  • L3 — Both governments commit to constraints on the rate or scale of frontier development: compute thresholds, training-run notification or licensing, limits on automated AI R&D, or a general pause or pacing.

Resolving source. Official text published by the White House, State Department, Treasury, China's Ministry of Foreign Affairs, or Xinhua.

Forecast horizons. Every quarter-end from 2026-09-30 through 2031-09-30.

How these forecasts are madeDownload all forecasts (JSON)