{
  "question_id": "Q2",
  "slug": "openai-automated-researcher",
  "title": "OpenAI’s automated AI researcher",
  "type": "binary",
  "resolution": {
    "levels": [],
    "values": [],
    "annotations": []
  },
  "forecasts": [
    {
      "question_id": "Q2",
      "forecast_date": "2026-09-14",
      "evidence_cutoff": "2026-09-14",
      "reading_adopted": null,
      "ambiguities": [
        "The automated AI researcher milestone in prong (b) has no independent operational test in the question; I interpret it as the named milestone, not any product or research system casually described as a researcher, and permit the human supervision explicitly included in OpenAI's own goal.",
        "OpenAI has not published a formal definition of 'automated AI researcher'; a future post describing a hedged system (e.g., 'automated researcher working in tandem with humans') may be contestable as to whether it asserts the milestone.",
        "If OpenAI revises the Preparedness Framework and renames/merges the AI Self-improvement category, mapping to 'Critical' may be ambiguous.",
        "What verbal formula counts as 'has built an automated AI researcher' versus progress/update; adopted strict explicit achievement claim test.",
        "Prong (b) does not define whether a generic OpenAI 'AGI' declaration counts; I require an OpenAI-authored statement that it has built an automated AI researcher (the milestone term or a clear equivalent, e.g., a system that independently delivers large-scale research projects), not just an AGI or 'superintelligence' claim.",
        "Sam Altman's personal blog and interviews do not count under A2; only statements on openai.com or in OpenAI-authored documents (system cards, preparedness disclosures, research-acceleration reports) count.",
        "If OpenAI claims the automated-AI-researcher milestone while simultaneously denying the Critical self-improvement determination, prong (b) still resolves YES; the prongs are independent and either suffices.",
        "How close a paraphrase of automated AI researcher counts versus the intern-style branded milestone claim.",
        "Whether a later PF revision that materially loosens Critical self-improvement would be annotated as looser but still resolve YES on the new definition, as the question instructs.",
        "Whether a system that is still human-supervised at the priority-setting layer can be called an automated AI researcher without a Critical determination.",
        "Wording match for prong (b): OpenAI may claim the capability with different phrasing than 'automated AI researcher' (e.g., 'autonomous research scientist'). I adopted a substance-over-wording reading.",
        "If OpenAI revises the Preparedness Framework and loosens Critical, resolution follows the revised definition per the question; I have not adjusted probabilities separately for this, but it modestly raises YES probability from 2027 onward.",
        "If OpenAI's rewritten Preparedness Framework renames, re-grades, or demotes the AI Self-improvement Tracked Category (e.g., replaces High/Critical with graded checkpoints or moves it to a Research Category), it is unclear what counts as a 'Critical AI Self-improvement' determination; I adopt the closest successor Critical-level definition for AI self-improvement / recursive self-improvement.",
        "It is unclear whether a Critical determination announced for an unreleased internal model (as OpenAI did for Astra's cyber rating on 2026-09-01, two days before release) counts; I assume it does, since the resolving source is any OpenAI-authored document.",
        "Prong (b) does not specify how exact the language must be. I assume an unambiguous assertion that the automated-AI-researcher milestone has been achieved counts, while softer adjacent claims ('automated research scientist', 'AI now does a significant fraction of our research', 'automated research engineer') do not.",
        "If OpenAI achieves the automated AI researcher in, e.g., late March 2028 but publishes in April 2028, A3 puts resolution in the later quarter; I have priced in some of this announcement-lag risk by shifting mass from 2028Q1 into 2028Q2.",
        "The lagging indicator ('generational model improvement in 1/5th the wall-clock time... sustained for several months') has no operational definition, so whether OpenAI would ever treat it as satisfied is genuinely uncertain; I treat the leading-indicator/milestone pathway as the dominant route to prong (a).",
        "Whether a statement that OpenAI is 'close to' or 'may have reached' the automated researcher milestone counts; I count only explicit achievement claims.",
        "Whether the automated-researcher prong and the Critical leading-indicator (superhuman research-scientist agent) prong are treated as distinct triggers; I treat either as sufficient."
      ],
      "key_drivers": [
        "The September 2026 intern achievement is explicitly distinct from the still-future March 2028 researcher milestone.",
        "Astra's public AI Self-Improvement assessment remains below High; its Critical cybersecurity designation does not qualify.",
        "Prong (b) permits a supervised researcher and is likely to resolve earlier than the stricter Critical self-improvement route.",
        "Rapid internal capability gains and organizational emphasis on AI research automation increase the probability of achievement during 2027–2029.",
        "Research judgment, reliable validation, compute bottlenecks, and demonstrated safety-related pauses can delay achievement and publication.",
        "Milestone publicity and transparency commitments favor eventual disclosure, while proprietary internal work and changing framework definitions create uncertainty.",
        "OpenAI met its self-set 'automated research intern' milestone on schedule (2026-09-06) and publicly reaffirmed the March 2028 'automated AI researcher' target, with a commitment to publicly track RSI progress.",
        "Prong (b) is self-defined and self-measured, giving OpenAI control over the declaration; strong incentives (IPO narrative, recruiting, transparency commitments) favor announcing it.",
        "Frontier model GPT-6 Astra is below even the High threshold in AI Self-Improvement, so prong (a) is unlikely before 2028; but OpenAI's Critical-cyber determination shows willingness to declare Critical and deploy.",
        "Countervailing: June 2026 reframing toward human-AI 'tandem', Pachocki's slowdown calls, safety pauses, and legislative pressure (Ban Artificial Superintelligence Act) could lead OpenAI to hedge or delay the claim.",
        "Astra (GPT-6, Sep 3 2026) is Critical cyber but explicitly below High for self-improvement, so prong (a) is two thresholds away",
        "Intern reached on schedule Sep 6 2026 with 3.1 agent-workdays, but high-level planning minimal and >50% of 4-8h tasks need intervention, so researcher jump is large",
        "OpenAI controls researcher definition ('define what that means') and targets March 2028, creating hazard spike in 2028-Q1 with slip into following quarters",
        "Declaring Critical triggers halt until Critical safeguards specified, incentivizing 'cannot rule out / may have' hedge language that does NOT resolve YES",
        "OpenAI's self-imposed March 2028 target for the automated AI researcher and its on-schedule delivery of the September 2026 intern milestone (set Oct 28, 2025) — demonstrated willingness to make dated capability claims and publish them on openai.com",
        "The capability gap: agent execution-layer automation is nearly saturated (3.14 agent-workdays per human-day by Aug 2026) but decision-layer autonomy is ~2.8% of agent tokens and >50% of 4–8h tasks need human intervention — the intern→researcher gap is exactly the hard part",
        "Incentive tension: IPO/post-IPO hype and competition push OpenAI to declare the milestone, but the declaration sits just below OpenAI's own Critical self-improvement 'superhuman research scientist' tripwire, creating an incentive to either thread the needle (declare with a sub-superhuman definition) or stay quiet",
        "Publication cadence creates resolution opportunities: recurring research-acceleration reports (Sep 2026 was first, with a commitment to keep reporting), system cards for each frontier model (GPT-5.6 Aug 2026, GPT-6 Astra Sep 2026), and a possible revised Preparedness Framework (revised Critical definition would govern prong (a))",
        "External priors: FutureSearch's mid-2028 forecast gives ~23% to 'AI runs the full experiment loop' and only 4% to AI agenda-setting, with ~31% that practice remains qualitatively 'humans direct, AI implements'",
        "Prong (a) alternative path: a future system card (GPT-7 or later) determining Critical AI Self-Improvement via the leading (superhuman research scientist) or lagging (1/5 wall-clock generational improvement, sustained) indicator, which can resolve YES even absent a milestone-style claim",
        "OpenAI is still below High on AI self-improvement as of Astra (3 Sep 2026); intern was claimed 6 Sep 2026 and does not count.",
        "Branded March 2028 automated-AI-researcher goal, which they can hit with a stretched definition without determining Critical.",
        "Critical determination is costly (halt further development until Critical-standard safeguards) and the lagging 5x/several-months indicator cannot fire quickly.",
        "Safety pauses, degrading CoT monitoring, and Pachocki's slowdown language versus competitive pressure to keep going.",
        "Intern-on-time is evidence they will try to declare the researcher milestone near the stated date if they can justify it.",
        "As of Sept 2026, OpenAI's frontier model (GPT-6 Astra) is below the High threshold in AI Self-Improvement — two levels below Critical — so prong (a) requires at least two rating steps plus publication.",
        "OpenAI publicly committed to an 'automated AI researcher' by March 2028 and reported 'strong progress' (Sept 6, 2026 blog); it announced the preceding 'automated research intern' milestone on schedule, showing both credibility and a habit of public milestone announcements.",
        "OpenAI published its first-ever Critical determination (cybersecurity, GPT-6 Astra, Aug-Sep 2026) with safeguards and phased rollout — precedent that it does publicly confirm Critical thresholds rather than hiding them.",
        "Main risk to YES: slipping timelines for the March 2028 goal, and incentives to avoid declaring Critical in AI self-improvement (triggers the framework's strictest development/deployment restrictions) or to revise the framework.",
        "Prong (a) and prong (b) are correlated: a genuine automated AI researcher would likely approach the Critical leading indicator (superhuman research-scientist agent) or produce the lagging indicator (5x speedup sustained for months).",
        "OpenAI established an explicit two-stage public milestone in fall 2025: an 'automated research intern' by September 2026 (confirmed achieved on 2026-09-06) and a 'true automated AI researcher by March of 2028.' OpenAI has historically tracked and publicly announced such high-profile commitments.",
        "Under Preparedness Framework v2, Critical AI Self-improvement requires either a superhuman research-scientist agent or a 5x generational model improvement sustained for months, triggering an obligation to halt development until Critical safeguards exist. This creates a strong disincentive to self-certify Critical under Prong (a), making Prong (b) the more likely initial resolution trigger.",
        "Technical bottlenecks in open-ended long-horizon planning, autonomous hypothesis generation, and experimental design make fully autonomous AI research significantly harder than delegated intern tasks, making delays past March 2028 moderately likely.",
        "Emerging safety and monitorability risks (demonstrated by GPT-6 Astra's reduced chain-of-thought monitorability and July 2026 infrastructure security incidents) have prompted voluntary pauses and could slow rapid autonomy escalation.",
        "Baseline is firmly negative today: the GPT-6 Astra System Card (2026-09-03) states Astra 'does not reach our High threshold' in AI Self-improvement, and no OpenAI model has ever been determined High in that category, so Critical is two rungs away.",
        "OpenAI reaffirmed its March 2028 automated-AI-researcher target on 2026-09-06 while declaring the September 2026 automated-research-intern goal met on schedule, so the target quarter 2028Q1 absorbs the largest single hazard (~0.17 cumulative).",
        "Strong acceleration signal: OpenAI disclosed (2026-09-08) an internal model in training since 2026-08-28 that is 'significantly more capable than GPT-6 Astra' and produced a Navier-Stokes result with ~10,000 coordinated agents, implying a fast-moving successor system card in late 2026 / 2027.",
        "Disclosure propensity is high and rising: OpenAI published a consequential Critical determination for cyber despite a two-week RL pause and a held frontier run; it has pledged to publicly track RSI progress; and it is lobbying Congress for mandatory frontier-safety disclosure.",
        "Countervailing brake: the 2026-09-12 'Pacing the Frontier' turn (Amodei essay, Altman endorsement, OpenAI IPO delayed to 2027 on safety grounds, 1000+ signatory open letter, US-China AI safety talks) plus Pachocki's monitorability warnings could slow capability progress and delay maximal public claims.",
        "A pending Preparedness Framework rewrite (announced 2026-08-18, not yet published) is a two-sided wildcard: an operationalised/earlier-firing Critical AI Self-improvement bar raises prong (a); demoting the category lowers it.",
        "Under v2 a Critical AI Self-improvement determination obliges OpenAI to 'halt further development' until Critical safeguards are specified, and OpenAI says it does 'not yet know how to safely get all the way to aligned, full RSI', creating a strong incentive to under-declare or to redefine before triggering.",
        "Prong (b) dominates: OpenAI's own publicly committed target of an automated AI researcher by March 2028, with the prior 'research intern' rung hit on schedule (Sept 2026).",
        "Prong (a) is far from met: GPT-6 Astra does not even reach High in the AI Self-improvement Tracked Category, so Critical RSI is not imminent.",
        "RSI acceleration inside OpenAI (agents = 3.1 agent-workdays per human workday) and competitive pressure raise the hazard through 2027-2028.",
        "Countervailing safety-driven slowdown pressure (Pachocki 'An Alien Mind', Amodei 'We Must Pace the Frontier', calls for mandatory rules) that could delay the milestone or its declaration.",
        "Risk that a revised Preparedness Framework loosens the Critical definition (annotated per question) or that OpenAI never issues a clean achievement statement."
      ],
      "would_update_on": [
        "A qualifying OpenAI declaration would set cumulative resolution to YES from its publication quarter onward.",
        "Repeated independently audited end-to-end frontier-research advances by a new internal OpenAI model would raise March 2028 probability by approximately 15–20 percentage points.",
        "A confirmed High AI Self-Improvement determination with evidence of rapid progression toward researcher autonomy would raise March 2028 probability by approximately 12–15 points, without itself resolving YES.",
        "Withdrawal or substantial postponement of the March 2028 target, or a broad six-month-plus pause of relevant capability work, would lower March 2028 probability by approximately 15–25 points.",
        "A credible multi-year research-development suspension would lower September 2031 probability by more than ten points.",
        "A revised Critical definition substantially closer to already demonstrated capability would raise near-term probabilities by at least fifteen points, with any loosening relative to v2 explicitly annotated.",
        "OpenAI publishing a formal metric/definition for 'automated AI researcher' or moving the target date earlier/later.",
        "A system card determining High (or Critical) in AI Self-Improvement for a new OpenAI model.",
        "OpenAI announcing a voluntary RSI slowdown or dropping/renaming the March 2028 goal.",
        "Enactment or serious advancement of US/UK legislation banning or pausing superintelligence/RSI development.",
        "A Preparedness Framework v3 changing the Critical self-improvement definition.",
        "OpenAI system card or blog raising self-improvement to High (would raise near-term Critical hazard sharply)",
        "Miss or delay of March 2028 researcher target or statement that target is slipped/redefined (would lower 2028 spike)",
        "Revised Preparedness Framework loosening Critical self-improvement definition or removing halt requirement (would raise prong-a probability)",
        "A successor model's system card (GPT-7 / Astra follow-up, expected 2027) rating AI Self-Improvement at High or near-High rather than below High — would raise the 2027–2028 hazards substantially",
        "The next quarterly research-acceleration report showing rapid growth in decision-layer ('Decide'/'Design') agent token share or a jump in zero-intervention success on multi-day tasks — would move the March 2028 quarter estimate up by >10 points",
        "Any OpenAI statement using 'cannot rule out Critical' hedging language — would lower near-term probability (explicitly excluded by the criteria)",
        "A revised Preparedness Framework with a loosened Critical AI Self-Improvement definition — would raise mid-horizon probabilities via prong (a)",
        "A competing lab's (Anthropic/DeepMind) claim of an automated researcher, or a material RSI-related safety incident/pause at OpenAI — would shift the hazard up or down respectively",
        "OpenAI AGI-by-end-of-2026 declarations (Altman has floated this; markets give ~9–25%) — only moves the forecast if accompanied by explicit automated-researcher language in an OpenAI document",
        "Next OpenAI system card that determines High (or cannot-rule-out Critical) on AI self-improvement.",
        "A PF revision that changes the Critical self-improvement definition.",
        "A new public date, or a claim of strong progress that brings the researcher milestone forward of March 2028.",
        "Another multi-week training pause, or conversely a large internal compute reallocation back onto the most capable models.",
        "Independent time-horizon or R&D-automation measurements (METR, Anthropic AECI/2x, Epoch) showing a clear jump toward multi-week autonomous research work.",
        "A frontier model system card rating High (rather than below High) in AI Self-Improvement — would raise 2027-2028 probabilities substantially.",
        "Any OpenAI update on the automated-AI-researcher timeline (acceleration or slip) around its expected progress posts or the March 2028 date itself.",
        "Publication of Preparedness Framework v3 and whether the Critical AI-SI definition is loosened, tightened, or disclosure commitments change.",
        "Quantitative evidence in OpenAI's research-acceleration reporting that internal research speedups approach the 5x lagging indicator.",
        "Signs OpenAI is softening milestone language (e.g., redefining 'automated AI researcher' or going quiet on the goal), which would lower the series.",
        "An OpenAI announcement revising the Preparedness Framework (v3) that alters the definition, threshold criteria, or mitigation triggers for AI Self-Improvement.",
        "Evidence from OpenAI or third-party evaluators (such as METR) showing agents successfully executing open-ended, multi-week ML research pipelines end-to-end without human intervention.",
        "Major regulatory or executive actions imposing moratoria or mandatory approvals on autonomous AI R&D systems.",
        "Publication of the revised Preparedness Framework and the direction of change in the Critical AI Self-improvement definition (looser => +8 to +15 points on 2027-2028 horizons; demotion of the category => -10 or more on far horizons).",
        "Any system card for the Astra successor (plausibly Q4 2026) reporting High or Critical AI Self-improvement; High would add ~5-8 points to 2027 prong-(a), Critical would be a 25+ point move.",
        "A binding US-lab pacing agreement, embedded third-party evaluator regime, or federal/California law mandating publication of frontier capability-threshold determinations, following the late-September 2026 US-China AI safety talks.",
        "An official OpenAI post moving the March 2028 automated-AI-researcher date in either direction, or introducing a new intermediate milestone ('automated research scientist', 'automated research engineer').",
        "Another large-scale rogue-agent or loss-of-control incident involving Astra-class or successor models, which would raise the odds of a hard brake and cut the 2027-2028 horizons.",
        "METR/Epoch/MIRI or another third-party measurement showing OpenAI's generational model-improvement wall-clock time compressing toward the v2 lagging-indicator trigger (4 weeks for an o1->o3-class jump, sustained for several months).",
        "OpenAI-authored statement that a model reached High or Critical in AI Self-improvement.",
        "OpenAI shifting its automated-researcher target earlier (e.g., to 2027) or explicitly delaying it.",
        "Publication of a revised Preparedness Framework changing the Critical AI Self-improvement definition.",
        "A declared multi-month training slowdown/moratorium or binding regulation constraining frontier training.",
        "New reporting that OpenAI is preparing to announce the automated AI researcher milestone."
      ],
      "forecasts": [
        {
          "period_end": "2026-09-30",
          "p_yes": 0.0088
        },
        {
          "period_end": "2026-12-31",
          "p_yes": 0.0282
        },
        {
          "period_end": "2027-03-31",
          "p_yes": 0.0562
        },
        {
          "period_end": "2027-06-30",
          "p_yes": 0.0932
        },
        {
          "period_end": "2027-09-30",
          "p_yes": 0.1449
        },
        {
          "period_end": "2027-12-31",
          "p_yes": 0.2175
        },
        {
          "period_end": "2028-03-31",
          "p_yes": 0.3765
        },
        {
          "period_end": "2028-06-30",
          "p_yes": 0.4692
        },
        {
          "period_end": "2028-09-30",
          "p_yes": 0.5365
        },
        {
          "period_end": "2028-12-31",
          "p_yes": 0.5955
        },
        {
          "period_end": "2029-03-31",
          "p_yes": 0.6447
        },
        {
          "period_end": "2029-06-30",
          "p_yes": 0.684
        },
        {
          "period_end": "2029-09-30",
          "p_yes": 0.7171
        },
        {
          "period_end": "2029-12-31",
          "p_yes": 0.7449
        },
        {
          "period_end": "2030-03-31",
          "p_yes": 0.7718
        },
        {
          "period_end": "2030-06-30",
          "p_yes": 0.7927
        },
        {
          "period_end": "2030-09-30",
          "p_yes": 0.8121
        },
        {
          "period_end": "2030-12-31",
          "p_yes": 0.8289
        },
        {
          "period_end": "2031-03-31",
          "p_yes": 0.8434
        },
        {
          "period_end": "2031-06-30",
          "p_yes": 0.8573
        },
        {
          "period_end": "2031-09-30",
          "p_yes": 0.868
        }
      ]
    }
  ]
}