Research worth keeping up with

AI Research Newsletter.

Consequential AI research and releases, with the context that makes them matter.

A curated digest of work from across the field. Prepared with AI assistance and primary-source links; these are summaries of others’ research.

Daily · Up to three updatesSignal over volume

AI for research

← All updates · 4 updates

GPT-6.1 Sol brings near-Astra task performance to a cheaper model with Critical cyber capability

OpenAI released GPT-6.1 Sol on September 29, reporting Astra-matching performance on DeepSWE v1.1 at roughly one-fifth of the task cost and more than double GPT-6 Sol's maximum-effort score on Terminal-Bench Science. Its system card classifies it as Critical for cybersecurity and High for biological and chemical capability, with the same safeguards stack as Astra.

Why it made the cut: Substantially cheaper access to strong coding and scientific agents, accompanied by a consequential capability-risk classification, matters beyond a routine model refresh. The comparisons are company-reported and depend on reasoning effort, tools, and evaluation setup; they do not establish equal real-world research ability. Astra remains the stronger scientific model in the reported tests, and lower token prices do not guarantee proportionately cheaper completed work.

System card addendum · Official release and evaluation summary

AI-assisted proofs reach the optimal secretary guarantee for all linear matroids

Hamed Abdi, Kiarash Banihashem, MohammadTaghi Hajiaghayi, and Danny Mittal posted a proof of the strong matroid secretary conjecture for every linear matroid, establishing the optimal 1/e competitive guarantee for this broad class of online selection problems. The authors say GPT-5.6 Sol and GPT-6 Astra were essential to the construction and that they independently checked the resulting proofs; a second group subsequently posted essentially the same construction, also obtained with Astra.

Why it made the cut: This is a substantive theoretical advance with explicit researcher attribution of AI's role and corroboration from a separate expert group, rather than a model's unsupported claim to solve a problem. Both accounts are preprints, the full strong conjecture for arbitrary matroids remains outside the proved linear-matroid result, and the coincident constructions raise unresolved provenance questions without establishing that private conversations were shared.

First paper · Concurrent paper · Official research-group announcement

OpenAI publishes an AI-generated, Lean-formalized proposed solution to Navier–Stokes

OpenAI released a 166-page proof claiming that smooth, forced three-dimensional Navier–Stokes flow can develop unbounded velocity in finite time while retaining bounded kinetic energy—establishing alternatives C and D in the Clay problem statement. The company says an unreleased model more capable than GPT‑6 Astra found the construction through a roughly 10,000-agent effort, and it released the accompanying Lean 4 formalization; the extraordinary result still needs independent mathematical scrutiny and formal recognition.

Why it made the cut: If validated, this resolves a 90-year-old Millennium Prize problem and is direct evidence of a frontier AI system producing—and formally encoding—a major new mathematical result. The complete paper and checkable Lean repository make the claim unusually concrete, even as questions about priority and training-data provenance remain unresolved.

Paper (PDF) · Official announcement · Lean proof repository · Independent coverage (Nature)

OpenAI says it has reached “automated research intern”–level AI R&D

OpenAI says its internal agents can now complete well-defined AI-research tasks under human direction that would take a skilled researcher several days, meeting a target it set last year. By mid-August, its research organization was consuming 3.1 agent-workdays for every human workday, while experiments per active experimenter reached their highest level since tracking began; however, OpenAI cautions that these internal metrics do not directly measure overall research progress, and more than half of successful four-to-eight-hour tasks still required human intervention.

Why it made the cut: This is unusually concrete operational evidence that a frontier lab is materially automating its own model-development loop, creating the possibility of faster capability and safety research. The self-reported, preliminary nature of the measurements matters, but so does the scale of actual use inside the lab building the frontier models.

Official research report and methods · Companion safety analysis · Independent safety context (TechCrunch)