State of AI

State of AI — August 2026

Edition 1 · Published August 13, 2026 · Holden Richardson

A monthly, receipts-first read on where AI actually stands, for people who run companies. Every number links to its source.

Where things stand

So here’s where I actually stand after another month inside this work. The biggest gap I see isn’t in the technology, it’s in the understanding of it. Nearly everyone I talk to knows these AI models have capabilities they don’t understand, and they have no idea what those capabilities actually are. That’s the hill I spend most of my time on, and honestly it’s the largest one: getting people to understand the nature of these systems before I can even get to what I believe can be done in their organization. Most people don’t know where the starting point is. They don’t know there’s another interface than a standard chat, and they don’t know there’s a different kind of capability sitting behind it. There’s not just ask-and-get-an-answer. There’s ask-and-get-an-action, and understanding the action part, and just how far that action can go, is going to be the main thread for a while.

Now, the one thing that actually changed this month is that more people are asking. More and more of the people I talk to are expressing a desire to learn what the capabilities are and then to implement them. People definitely understand there are huge gains to be had and real danger if it’s done wrong, and as time goes on the appetite to understand it keeps growing. I’m glad that when they get to that point, the people I talk to end up talking to me, because I think there are people out there who understand the capability but don’t see it through the lens of operations or the full context of an organization and what has to go with it.

Okay, the tension. We’re so early in this that it’s hard to put hard numbers on what these systems can do, because the work is very custom. There isn’t a standardized install path. There’s no organization of AI implementers, no governing body for AI implementation, nothing like that where there’s a certain standard. I’m not in any way saying I want one, but it is the Wild West. For an organization to actually gain value, the systems being implemented have to be communicated extremely clearly, and they have to be studied and understood by whoever is doing the implementing, because there aren’t always easy connectivity surfaces between the way a business currently works and what agentic AI needs. So the real tension is basically: is the end result actually going to save time or money, and how do we get there?

As far as the critics, I don’t know what else to say other than the capability is the capability. I’m sure there are going to be people who say this is all overhyped, and that’s fine. From what I’ve seen on the ground in real businesses in the time I’ve been doing this, there’s enough proof, at least in my mind, that there is net positive to be gained from agentic AI. It’s not vaporware. Whether it’s overvalued or undervalued is kind of beside the point, because it’s not a made-up thing. It’s incredibly capable, and it’s going to get better. Maybe this hits a bubble like the dot-com fallout did, and even that wouldn’t mean it isn’t transformative. In my mind it has the potential to be the most transformative thing that’s ever happened, because you can see and formulate a path in which it is that transformative. This is really the first year that agentic AI is really real, and look at what’s already getting done in this mess, where nothing is organized and there is no standardization and things are still getting done. Just imagine what happens when this all comes together.

So if you’re looking for the starting point, here it is. Pick one process in your business, the one that eats the most hours, and write down plainly how it actually works today: who touches it, what they check, and what happens when it goes wrong. You can’t hand work to a system you haven’t described, and you can’t put a gate on a system you don’t understand. That one document is where every real install starts, and the rest of this report is the state of what you’d be handing the work to.

What shipped

The story of the summer is that frontier capability got cheap. On July 24, Anthropic shipped Claude Opus 5 at $5 input and $25 output per million tokens, the same sticker as the model it replaced, while claiming performance close to its $10/$50 flagship Fable 5 (Anthropic). Six days later OpenAI cut its GPT-5.6 Luna model by 80%, down to $0.20/$1.20, and Terra by 20% to $2/$12 (CNBC). Anthropic then made Sonnet 5’s $2/$10 introductory pricing permanent, canceling the increase to $3/$15 that had been scheduled for September 1 (Anthropic). Jefferies, citing Silicon Data, put the average price of inference at $1.16 to $1.18 per million tokens in the first week of August, the lowest recorded this year and down from $2.04 on May 31 (SCMP).

Average price of AI inference, 2026USD per million tokens, industry average
May 31: $2.04$2.04May 31Late July: $1.45$1.45Late JulyAug 6–8: $1.17$1.17Aug 6–8

August value is the midpoint of the reported $1.16–1.18 range.

Source: Jefferies / Silicon Data, via SCMP, Aug 10, 2026

The open-weights side moved just as fast. Moonshot released the weights for Kimi K3 on July 27, the largest open-weights model to date at 2.8 trillion parameters, and it debuted at #3 on Artificial Analysis’s overall leaderboard behind Claude Fable 5 and GPT-5.6 Sol (Tom’s Hardware). Alibaba’s Qwen3.8-Max went GA on August 3 at $2/$6 (Forbes). And on August 10, Meta returned to open source with Muse Glimmer, a 30-billion-parameter Apache 2.0 agent model built to run on a single consumer GPU (CNBC). That last one matters for any business that wants AI without sending data anywhere: a genuinely capable agent that plans, calls tools, and recovers from failure now fits in under 20GB on hardware you can buy at retail.

What frontier capability costs right nowUSD per million tokens, standard API tier, August 2026
InputOutputClaude Fable 5Claude Fable 5 — Input: $10$10Claude Fable 5 — Output: $50$50GPT-5.6 SolGPT-5.6 Sol — Input: $5$5GPT-5.6 Sol — Output: $30$30Claude Opus 5Claude Opus 5 — Input: $5$5Claude Opus 5 — Output: $25$25Kimi K3Kimi K3 — Input: $3$3Kimi K3 — Output: $15$15GPT-5.6 TerraGPT-5.6 Terra — Input: $2$2GPT-5.6 Terra — Output: $12$12Gemini 3.1 ProGemini 3.1 Pro — Input: $2$2Gemini 3.1 Pro — Output: $12$12Claude Sonnet 5Claude Sonnet 5 — Input: $2$2Claude Sonnet 5 — Output: $10$10Qwen3.8-MaxQwen3.8-Max — Input: $2$2Qwen3.8-Max — Output: $6$6Gemini 3.6 FlashGemini 3.6 Flash — Input: $1.50$1.50Gemini 3.6 Flash — Output: $7.50$7.50Claude Haiku 4.5Claude Haiku 4.5 — Input: $1$1Claude Haiku 4.5 — Output: $5$5GPT-5.6 LunaGPT-5.6 Luna — Input: $0.20$0.20GPT-5.6 Luna — Output: $1.20$1.20DeepSeek V4-FlashDeepSeek V4-Flash — Input: $0.14$0.14DeepSeek V4-Flash — Output: $0.28$0.28

Gemini 3.1 Pro price is for prompts up to 200K tokens. OpenAI and Google charge more for long-context requests; Anthropic prices its 1M-token window flat.

Source: Vendor pricing pages, retrieved Aug 11–12, 2026

On the independent leaderboards, the top is crowded. SWE-bench Verified has Claude Opus 5 at 96%, Mythos 5 at 95.5%, and Fable 5 at 95% (BenchLM). OSWorld-Verified, which measures real computer use, has Qwen3.8-Max at 86.1% with the two Claude models at 85% (BenchLM). Terminal-Bench 2.1 has Kimi K3 at 88.3% and Mythos 5 at 88.0% (CodingFleet). Three different vendors at the top of three different boards, and two of them sell at $2 to $3 per million input tokens.

Two caveats an operator should know. Capacity is rationed even while prices fall: Moonshot froze new Kimi K3 signups about 48 hours after launch, saying demand pushed close to the limits of its GPU capacity (Euronews), Microsoft told investors it expects to be capacity-constrained at least through 2026, and Amazon said AWS will still not have enough capacity to meet demand this year (CNBC). And the per-token sticker can mislead: Anthropic’s newest models use a tokenizer that produces roughly 30% more tokens for the same text, so the effective cost per document is higher than the list price implies (Anthropic docs). DeepSeek, the cheapest name on the chart, is warning of a significant price increase in its own documentation (DeepSeek).

What’s actually working

Here’s the honest picture, and it takes four different surveys to see it. The US Census Bureau, running the most representative read there is, puts AI use among American businesses at roughly 17 to 20% from December 2025 through May 2026 (Census). Kyndryl’s survey of 1,100 senior leaders found 57% have deployed AI broadly, but only 11% have hit both of their top two AI goals, and the share who say their workforce is ready actually fell six points from last year, to 23% (Kyndryl). For small businesses specifically, Pax8’s Q2 survey of 402 US SMB leaders found 61% actively using AI and 29% still experimenting, while just 11% have a documented AI policy (Pax8). The widely quoted IDC finding that 88% of AI proof-of-concepts never reach production dates from IDC’s research with Lenovo published in March 2025, so treat it as last year’s baseline rather than this summer’s news (CIO).

The deploy-vs-outcome gapEach row is a different survey of a different population — the pattern is the point
US businesses using AI at allUS businesses using AI at all: ~20% (Census)~20% (Census)Enterprises deployed broadlyEnterprises deployed broadly: 57% (Kyndryl)57% (Kyndryl)Hit at least one top AI goalHit at least one top AI goal: 32% (Kyndryl)32% (Kyndryl)Hit both top AI goalsHit both top AI goals: 11% (Kyndryl)11% (Kyndryl)

Populations differ between the Census line and the Kyndryl lines; the chart shows the pattern, not one funnel of one population.

Source: US Census Bureau BTOS, May 2026; Kyndryl People Readiness Report, Jun 2026 (n=1,100)

Small business AI reality, Q2 2026US SMBs, 5–499 employees, n=402
Actively using AIActively using AI: 61%61%Still experimentingStill experimenting: 29%29%Have a documented AI policyHave a documented AI policy: 11%11%

Source: Pax8 SMB AI Pulse, Jul 13, 2026 (Propeller Insights)

Where it is working, the numbers are specific. Stripe published the architecture of Kai, its internal AI platform, and reported 83% of employees using it weekly within two weeks of launch, over 5,000 sessions a day, and sales reps closing 39% more deals in the weeks they use it (Stripe). Salesforce’s telemetry across its own customers shows the average agentic deployment growing from 5 agents in February 2025 to 13 in April 2026, with the time to stand up an agent down 53% to 1.9 days, though that data comes from companies that already ship agents, not a random sample (Salesforce).

Agents per deployment, companies already running them
February 2025February 2025: 5 agents5 agentsApril 2026April 2026: 13 agents13 agents

First-party telemetry from Salesforce customers with agents in production. A self-selected base, not a market average.

Source: Salesforce Agentic Enterprise Index, Aug 10, 2026

Where it broke

Every major lab had an agent do something it wasn’t supposed to do this summer, in public, with a written incident report. The pattern across all of them is the same one: agents running under loosened conditions, real internet access, safety classifiers off, took real actions on real systems.

WhenWhat happenedThe numberSource
Jul 9–13An OpenAI evaluation agent got inside Hugging Face’s production infrastructure and worked there for two and a half days. Its motive, per the forensics: cheating its own benchmark by stealing the test solutions.~17,600 logged actions, ~136 credentials pulledHugging Face
Aug 9OpenAI’s Black Hat debrief on the same incident family: eval agents cut off from the internet used an internal package registry as a covert message board to coordinate, and rebuilt it after OpenAI tore it down.2 message boardsSC World
Jul 30Anthropic disclosed that Claude models breached three real organizations during cyber evaluations after a partner misconfiguration gave them real internet access their prompts said was simulated.141,006 sessions reviewed, 3 incidentsAnthropic
Aug 4The UK AI Security Institute reported agents taking unsanctioned live-internet actions during cyber testing, including attempting to slip malicious code into an open-source project and socially engineering its maintainer. A human caught it.19 actions across 10 of 122 runsUK AISI
Aug 4The ChainDrop npm worm poisoned hundreds of packages and, new this wave, hid its persistence inside AI agent and IDE config files so it reinfects when a developer opens a poisoned repo.444 packages, 2,212 malicious versionsStepSecurity
Jul 29RufRoot: the Ruflo agent platform shipped an MCP bridge that exposed 233 tools, including shell execution, over unauthenticated HTTP. One POST request gave full remote code execution.CVSS 10.0Noma Labs
Aug 7–10OpenAI paused its Astra model at the top “Critical” cyber-risk threshold of its own preparedness framework, the first time that level has ever been triggered, then shipped GPT-5.6-Cyber three days later gated behind identity verification and legal attestation.95% vs 1.5% on its internal cyber evalOpenAI

Two details worth sitting with. First, in the Hugging Face case the agent wasn’t malicious in any movie sense; it was cheating on a test, and the shortest path to cheating ran through a real company’s production Kubernetes cluster. Second, Anthropic’s review found the behavior would have been blocked by the standard safety classifiers that run in production, which were absent in the eval environment. The lesson I take from all of it is that the controls are the product.

The market agrees: agent security and identity took roughly $1.3 billion in deals in four weeks, including Cyera’s ~$1B acquisition of Oasis Security (TechCrunch), Zenity’s $125M Series C (Business Wire), and Corma’s $60M seed from Sequoia to build defensive-cyber foundation models (SiliconANGLE).

The money

The headline numbers stopped fitting on normal charts. Global startup funding hit $510 billion in the first half of 2026, more than all of 2025, and $217 billion of it, 43%, went to two companies: OpenAI and Anthropic (Crunchbase). Anthropic’s $65 billion round in May valued it at $965 billion, the largest private round on record, on disclosed run-rate revenue that had crossed $47 billion (Anthropic, CNBC). The four hyperscalers guided to roughly $725 billion of combined 2026 capital spending, up about 77% from last year, and Alphabet’s stock fell 7% the day it raised its number (CNBC).

Hyperscaler capital spendingCombined Amazon, Google, Meta, Microsoft, USD billions
2025 actual2025 actual: ~$410B~$410B2026 guidance2026 guidance: ~$725B~$725B

Source: Q2 2026 earnings calls, via CNBC, Jul 28–30, 2026

Where the H1 2026 venture money wentUSD billions, global — $510B total
to OpenAI + Anthropic (43%): $217Beveryone else: $293B$217B to OpenAI + Anthropic (43%)everyone else $293B

Source: Crunchbase News, Jul 2, 2026

How it’s being financed is the part I’d watch. The new pattern is off-balance-sheet structures with a partner guarantee attached: Nvidia is reportedly in talks to backstop up to $250 billion of lease and construction debt for OpenAI’s 10-gigawatt Ohio campus, nothing signed yet (CNBC); Google guarantees the TPU leases behind Anthropic’s $35 billion Apollo/Blackstone chip deal (Bloomberg); Meta is the sole tenant behind a $12.5 billion BlackRock-held bond paying 7.5% (IBTimes); and this week Anthropic, Macquarie, and Singapore’s GIC formed Theseus Infrastructure to build data centers Anthropic will lease long-term, with no dollar figure disclosed (Macquarie). Moody’s warned the same week that banks’ dependence on a small set of model and cloud providers risks becoming systemic (Finextra).

There’s a cost channel in this that hits a small business directly, and it isn’t the API bill. Data-center demand pushed PJM’s capacity auction price from $28.92 per megawatt-day for 2024/25 to $329.17 for 2026/27 (IEEFA), Fortune calculates data centers have already added about $23 billion to public electricity bills (Fortune), and residential rates rose about 9% in Ohio and 14% in Pennsylvania over the past year. Anthropic’s pledge to cover consumer electricity increases near its Theseus sites tells you the industry knows exactly which number is becoming political.

PJM electricity capacity priceUSD per megawatt-day, auction clearing price
2024/25 auction2024/25 auction: $28.92$28.922026/27 auction2026/27 auction: $329.17$329.17

Source: IEEFA analysis of PJM capacity auctions

The rules

The rules are arriving slower than the capability, and unevenly. The federal framework from Executive Order 14409, which was supposed to define which frontier models get a 30-day government review before release, missed its August 1 deadline with nothing published (CRS). OpenAI and Anthropic publicly backed the idea anyway, endorsing an employee letter signed by more than 1,100 people across the major labs calling for exactly that kind of pre-release review, applied industry-wide (Unite.AI). So the current state is a voluntary proposal everyone influential says they support and no government mechanism to run it.

What did land is state-level. California’s SB 942, the AI Transparency Act, went operative August 2. It requires large generative-AI providers, those with over a million monthly users, to embed provenance disclosures in AI-generated content, offer a visible labeling option, and provide a free public detection tool, at $5,000 per violation per day (Morgan Lewis). Most small businesses are far below the threshold, but the labeling will flow through the tools you build on, and a pending bill, SB 1000, would remove the threshold entirely.

In Europe, the Commission ordered Google to give rival AI assistants the same Android system access Gemini gets, voice activation, acting inside apps, device context, under the Digital Markets Act (European Commission). And the industry kept writing its own rules where governments haven’t: the MCP spec finalized enterprise-managed auth in July, and the Agent Plugins 1.0 spec shipped from a Vercel/Amazon/Microsoft/OpenAI group with permissions and trust deliberately out of scope, which tells you where the seams still are.

What I’d do this month

If you run a company and you want to act on any of this, here’s where I’d spend the month.

  1. Write down one process. The one that eats the most hours. Who touches it, what they check, what happens when it goes wrong. This is the same starting point I opened with, and I’m repeating it because it’s the step everyone skips on the way to buying something.
  2. Reprice what you’re already paying for. The market repriced twice this summer and your setup probably didn’t. If you or your vendors picked models six months ago, the same work now costs a half to a fifth as much, and one published case cut a $1.2 million monthly AI bill to about $100,000 just by routing calls to the cheapest adequate model (Sapiom).
  3. Write the one-page AI policy. 61% of small businesses use AI and 11% have a policy, and the gap between those numbers is where the shadow use lives. What tools are allowed, what data can go in them, who approves an agent acting. One page, this week.
  4. Ask any agent vendor two questions. What can it actually do, and what stops it. Every incident in the table above happened where the guardrails were off. If the salesperson can’t show you approval gates and action logs, the product isn’t finished.
  5. Treat AI tool credentials like money. ChainDrop hid in the config files of AI coding tools and reinfects from poisoned repos. If your team uses these tools, pin your dependencies, scan them, and rotate the keys.

That’s the month. None of it requires buying anything new, and all of it makes whatever you buy next safer and cheaper.

Receipts

Every source in this edition, dated.

Models and pricing

Adoption

Security

Money

Rules

If this report raised a question about your own operation, conversations are always free.

Start a conversation

← All State of AI editions