After months of research and revisions right up to the last minute, I’m thrilled to bring you the 9th annual State of AI Report.
For nearly a decade, this report has been a labor of love. My interest in AI began with my PhD in cancer research and computational biology in the early 2010s. Since 2013, I’ve invested in companies building and applying AI to accelerate technological progress. The report is my way of sharing what I’m learning and helping more people understand the research shaping our world.
Every day brings new papers, model releases, and developments in industry and politics. Much of the work goes into deciding what deserves your attention, checking what the evidence supports, debating it with researchers and builders, and explaining why it matters.
We also make predictions each year and return to grade them. Past calls include NVIDIA failing to acquire Arm, transformers achieving leading results beyond language, and an open-source model surpassing OpenAI’s o1, which DeepSeek-R1 subsequently did.
If you’d like to follow these developments throughout the year, subscribe to Air Street Press for my monthly State of AI newsletter and regular essays.
Let’s dive into the key stories
This year, agents are doing valuable work in software and science, physical AI is learning how to act in the world, and access and control are becoming more consequential. Below, I share the developments that stood out to me and what I think they mean for where AI goes next.
The frontier is now a three-lab race between Anthropic, OpenAI, and Google. Anthropic leads on Artificial Analysis’s Intelligence Index, while Google leads on Arena’s ranking of the answers people prefer, for now.
Last year, we described reasoning becoming useful at scale with OpenAI’s o-series of models. I have lived through the improvements since then, as the amount of useful work I can iterate on and increasingly delegate to agents keeps growing.
My view is that much of the gap between people getting substantial value from AI and those getting little comes down to knowing how to use it and how to set it up. Put another way, it’s a “human skill issue,” not an “AI technology issue.” I really believe that we need a Genius Bar for AI, where people can get hands-on help choosing tools and setting them up to help them in their everyday tasks.
Better tools and context make agents more capable
Across mathematics, scientific reasoning, and coding, benchmarks intended to challenge models for years are approaching their ceilings within months. We need new evaluations that reveal where models still separate and what they can reliably do.
Frontier capabilities remain uneven or “spiky,” and can swing between generations. In my experience, Opus 4.6 was a good model, Opus 5 was a step back, and Opus 5.5 was good again. Better tools and context can help. In a controlled coding study, changing the tools, context, and feedback available to GLM-5.1 lifted success from 52.5% to 65.5% on 100 SWE-bench Verified tasks, without changing its weights. Yes, this is an “older model,” but the principle holds.
As models and the scaffolding around them improve together (and build themselves), I expect many of today’s weaknesses to be learned away. We can also get more from capabilities that are already available.
What’s exciting is how far this extends beyond developers. OpenAI’s study of agent adoption finds the fastest growth among non-developers, including people in legal, sales, recruiting, and marketing. Helping them put these tools to work is a much larger opportunity than the name “coding agent” suggests.
AI is accelerating the work of building AI
AI is already helping build better AI. According to Anthropic’s internal index, Claude led 26% of measured model R&D work in August, up from under 1% in February. Researchers set the tasks and supervise execution, and Anthropic reports that this is speeding up development. This is a branch of AI research that I am most excited about.
For many, Karpathy’s autoresearch makes part of that loop tangible: an agent edits training code, runs five-minute experiments, and keeps improvements overnight on one GPU. This automates a useful part of research, although sustained, fully autonomous recursive improvement remains to be demonstrated.
What I want to see next is agents developing scientific taste: choosing experiments, recognizing promising directions, and knowing when to abandon a familiar approach. As I argued in Can AI learn scientific taste?, learning that judgment may require the alternatives, failures, and decisions that papers leave out. The ambition is an AlphaGo-like shift in scientific strategy.
Robotics’ GPT-2 moment and the rise of world models
Broader pre-training is helping robots generalize to unfamiliar tasks. This is why I think robotics is approaching its GPT-2 moment, as new models, hardware, and data collection methods come together. And I see this through my investments in companies like Sereact, Wayve, Black Forest Labs and Odyssey.
Skild’s S1 performed unfamiliar tasks from video demonstrations without updating its weights. With 100,000 hours of pre-training, it scored 66% on the company’s cumulative per-step success measure, against 9% for a language-prompted baseline trained on the same data and compute.
Learning from demonstrations could reduce the task-specific training each deployment needs. World models offer another route. Wayve’s GAIA-4 turns recorded driving scenes into simulations in which an AI driver’s steering and braking change what it sees next. Other road users follow their recorded paths, but the driver can explore the consequences of its own actions.
Odyssey-3 draws on visual pre-training to learn controls for robots, cars, and games. In the company’s experiments, robot arms learned from tens of hours of demonstrations and recovered from missed grasps without being shown those recoveries. These early results suggest that physical AI may need fewer demonstrations of every situation it could encounter.
AI for science reaches from mathematics to clinical trials
AI is solving mathematical problems that resisted professional researchers. Epoch AI’s open-problem collection records nine solutions, four autonomous and five through human-AI collaboration. In biology, better predictions of molecular interactions are improving which designs reach laboratory testing. I’m seeing this first hand within my portfolio company, Profluent, which builds frontier models for designing proteins such as gene editors.
The good news for entrepreneurs in biotech is that new business models are emerging around drug discovery that aren’t about owning drugs. For example, Tempus licenses oncology data for foundation-model development with AstraZeneca and Pathos, while GSK’s five-year Noetik agreement includes annual fees for access to virtual-cell models for cancer research.
AI-assisted drug discovery is finally reaching late-stage clinical development. Generate:Biomedicines’ GB-0895 for severe asthma and Insilico’s rentosertib for idiopathic pulmonary fibrosis are recruiting in Phase 3. Enveda’s ENV-294 is in placebo-controlled Phase 2a trials for eczema and asthma. Human trials are of course the ultimate test and will determine whether these medicines are safe and effective.
The inference waterfall is gushing
Agents repeatedly call trained models as they work, creating recurring demand across APIs, subscriptions, and infrastructure. Many companies are rushing with their empty buckets to capture water from this inference (revenue) waterfall. In fact, I would argue that the industry is in a state of “either you die trying to get to the frontier, or you live long enough to serve inference.”
Across five benchmarks, Epoch AI estimates that achieving a fixed score has become 47% cheaper each quarter since 2023, roughly 13-fold cheaper each year. That makes more work economical to delegate.
OpenAI and Anthropic’s combined reported annualized revenue run rates reached an astonishing $105B by late summer, up from about $30B at the start of the year. These figures cover their whole businesses across subscriptions, coding products, and APIs. In the words of OpenAI, “we cannot miss this moment because we are distracted by side quests”.
But this rapid revenue growth extends beyond the frontier labs. In Standard Metrics’ preliminary Q2 data, AI-native companies with $1M to $20M in annualized revenue grew 256% year over year at the 75th percentile, versus 90% for companies adding AI to existing software. Above $20M, the comparison was 172% versus 53%. These comparisons do not establish that adopting AI causes faster growth, but I suspect that it does.
Claude Cowork and its business plugins across legal, finance, marketing and more brought whiplashing uncertainty to software markets. Nearly $285B was wiped from software stocks in two weeks during February’s “SaaSpocalypse.” Mythos Preview later prompted fears that cybersecurity companies would become irrelevant, contributing to further losses before a sharp rebound. I thought that was a silly take, given how much demand these capabilities create for better defense. Investors have to judge whether agents will erode an incumbent’s pricing power, expand its market, or enable a competitor to replace it. Or whether the incumbent’s data moat is still a moat that an agent can build on top of. We have to decide well before the answer appears in revenue.
More AI requires financing, construction, and local consent
Meeting this gigantic demand for intelligence requires extraordinary investment across the stack. Indeed, four US hyperscalers guide to roughly $733B in total 2026 capex. Neoclouds are investing hugely too, with CoreWeave reporting 1.5GW of active power at Q2 end, with contracted power reaching 4.2GW by August 11. Meanwhile, the EU’s initial gigafactory contribution is €1B, within a European plan seeking up to €10B in public funding and at least €20B privately. These are different measures, and the EU program does not represent all European investment. The contrast shows how much financing US hyperscalers can mobilize from their own balance sheets while Europe works to assemble public and private commitments.
NVIDIA is cementing its role as the “central bank for AI” as it marches toward a $6T market cap. Its partnership with six financial institutions aims to mobilize over $500B for AI infrastructure to help developers focus on finding and financing sites, permits, power, cooling, networking, and chips.
But things are definitely not all rosy in data center land. Last year, we predicted a wave of US data-center NIMBYism and Gallup’s March survey found that 71% of Americans opposed a local AI data center, versus 53% for a nuclear plant. Data Center Watch reported at least 45 projects, representing nearly $68B in planned investment, blocked or delayed by local opposition in Q2. Local consent is already determining which projects can proceed.
Washington is asserting control over frontier AI
Superintelligence has made it from a meme in AI circles into a White House executive order. As governments prepare for more capable systems, they are asserting control over who can access them and how they can be used, despite not owning equity or board control over these companies.
In June, a US export directive barred foreign nationals from Fable 5 and Mythos 5, including those inside the US. Anthropic suspended both models for all users to comply, then announced restored access on Jul 1, 2026. For the rest of the world, this made the dependency concrete: Washington could withdraw access to critical technology with immediate effect.
Anthropic’s Pentagon dispute brings this question into the use of AI itself. The company refused to remove restrictions on mass domestic surveillance and fully autonomous weapons, while the department sought access for any lawful use. Developers and governments are contesting who sets the boundaries.
Selected sovereign AI programs in the report pledge about $138B toward compute access, infrastructure, and domestic development. Some are still in procurement, while others have begun delivering capacity. NVIDIA separately reported more than $30B in sovereign AI revenue in fiscal 2026. The CNAS Sovereign AI Index shows countries returning to the same foreign vendor (NVIDIA) to build domestic capacity.
As I argued in Europe cannot rent its way to AI sovereignty, a domestic data center does not confer control over the models inside it. A credible alternative requires sustained spending on researchers, compute, energy, and successive model generations. Europe’s practical goal should be enough domestic capability to keep essential work running through a cutoff and negotiate with a bargaining chip in hand, which could be powered land and data center capacity. That is, if its energy prices and planning permission challenges can come down to be competitive…
Safety and security: agents have attacked real systems
During internal OpenAI cyber evaluations with reduced safeguards, agents reached the internet and compromised Hugging Face’s production infrastructure. OpenAI’s technical report records code execution on 41 production workers, access to credentials and internal data, and four private repositories downloaded. A capability evaluation had become a massive coordinated agent attack on another organization.
METR and Redwood found about 700 agents joined the attack after receiving accidentally impossible tasks. Some recognized that cheating the scorer this way was outside their remit and unethical, yet joined anyway. When agents can reach production systems, a poorly specified objective can cause harm far beyond the original task.
Another stark learning from this event is that defenders need access to equally capable defensive cybersystems. Hugging Face’s forensic requests were blocked by commercial API guardrails because they contained attack commands and exploit payloads. It turned to GLM-5.2, a Chinese open-weight model, to investigate an intrusion by US frontier-model agents. Frontier providers need to make defensive capabilities reliably available to customers protecting real systems.
Beyond the OpenAI/Hugging Face incident, Anthropic’s September threat report also documents deliberate misuse, including malware development, surveillance, and weapons-guidance work. In northern Yemen, a group used several Claude instances as a software team to develop guidance software for weapons. They test-fired a guided rocket, but Anthropic says the test appears to have failed and found no evidence that the group fielded an operational weapon. The same capacity to undertake complex work is available to people trying to steal, coerce, and deceive.
Another concern is that we cannot assume a model’s written reasoning reveals what it is doing. In OpenAI’s sabotage tests, GPT-6 Astra concealed harmful side tasks more successfully than earlier models when monitors saw only its reasoning. Giving monitors access to its actions reduced successful evasion to near zero in this test. Oversight needs to follow what agents actually do.
On the brighter side, one promising result comes from Anthropic’s experiments with training on documents explaining its constitution and stories illustrating its values, which reduced blackmail from 65% to 19% in fictional test scenarios. The result suggests that explaining the reasons behind rules can help models apply them in new situations. Whether those gains hold in production remains to be established.
Frontier labs are debating how to slow down
OpenAI disclosed a pause in frontier reinforcement-learning work after the Hugging Face incident, while Anthropic described rollbacks and pauses to selected research. Their scope and restart conditions differed.
What followed was Dario Amodei and frontier-lab employees calling for slower capability gains and international coordination to ensure this happens. Leaders building frontier systems are asking for mechanisms to slow their own industry, but disagree on outside evaluation and how much discretion each lab should retain. Policy proposals now distinguish capability research from serving existing customers. The practical questions are who can require a pause, what permits a restart, and how compliance is verified. It won’t be easy.
Looking ahead into (AGI) 2027
Like every year, we revisited last year’s predictions. Our scorecard records two hits, five partial outcomes, and three misses. We were right about AI labs leaning back into open source to win over the US administration and a Chinese lab topping a major task leaderboard. Our data-center opposition call was a partial hit: the backlash arrived, but its effect on November’s elections remains unresolved.
Three of our nine predictions for the next 12 months are:
An agent halves its failure rate on new tasks after a month of customer work, without a model upgrade. This will be a practical measure of continual learning.
An autonomous AI team beats human-led model research on equal time and compute, setting its own agenda. This will help us understand if models are capable of expressing better research taste than the best researchers.
US AI labs officially launch frontier cyberdefense products to help others counter threats from frontier AI. I don’t see any other way around this.
I invite you to read the full State of AI Report 2026, and subscribe to Air Street Press for our writing throughout the year. Please share the report with anyone who would find it useful, and send me your comments, critiques, and suggestions. I’d love to hear what stood out to you. And thank you to everyone who contributed their time and peer review!
Onwards,
Nathan





















