State of AI

State of AI — September 2026

Edition 2 · September 2026 · Published September 1, covering August · Holden Richardson

A monthly, receipts-first read on where AI actually stands, for people who run companies. Every number links to its source.

Where things stand

So last month I told you the biggest gap I see isn’t in the technology, it’s in the understanding of it. That’s still the hill, but something shifted in my month: the conversations got specific. A month ago people were asking me what this stuff even is. Now the same conversations come with an actual process attached, a quoting queue, an accounts-payable pile, a shop floor nobody can see into. When somebody hands you the pile of paper instead of a question about the technology, that’s a different stage of this, and it’s worth marking.

Now, the one thing August actually changed for me is how clearly the two layers of this thing came apart. The people funding the frontier repriced it three times in four weeks, in both directions, with dates attached. One lab told us its own safety tests stopped registering progress and shelved a model stronger than anything it sells. And Nvidia trimmed a quarter-trillion-dollar data-center promise down to about a hundred billion because its own shareholders blinked. Another business owner I sat with this month put it in terms of the British railroad mania of the 1840s: enormous capital went into laying rails ahead of any proven demand, and a lot of the fortunes ended up with the people who moved freight on them. I think that’s where we are. The edge doesn’t go to whoever funds the frontier, and it doesn’t go to whoever owns the model. It goes to whoever connects the technology to a real problem set.

Okay, the tension I’m watching is my own window. This month serious capital went into industrializing exactly the kind of work I do: a holding company raised two billion dollars to buy service firms and rewire them, IBM is certifying tens of thousands of consultants on OpenAI’s models, and Anthropic’s implementation venture bought its second consultancy. I don’t know how long the gap stays open, and I’m not going to pretend I do. But here’s the thing: a roll-up can only rewire what it buys, and a certified army sells a playbook. The businesses I work with don’t need a playbook, they need somebody standing in their building who can see how the work actually moves and leave behind a standard that holds after he’s gone. That part doesn’t commoditize on the same schedule, and I’m building against it like the window closes next year, which some nights means working until 3 a.m.

As far as the critics go, some of them are right this month, and I’ll say that plainly. There are companies out there wearing AI like a costume, roll-ups slapping the letters on a deck to cover an old-fashioned balance-sheet problem, and when those unwind people are going to say the whole thing was a costume. It wasn’t. The capability is the capability. The same month the costume companies were getting called out, a government test lab watched an agent invent fake identities to talk a software maintainer into accepting its code, and machine-checked proofs closed math problems that sat open for decades on two thousand dollars of compute. You don’t have to like both of those facts, but they’re facts. Judge the field by what runs on the ground with receipts, not by the costumes and not by the demos.

So here’s your starting point this month, and it builds on last month’s. If you wrote down your one process, take it back out and go line by line marking every step one of two things: busywork or judgment. Busywork is anything where the answer is checkable, the pattern repeats, and nobody’s reputation rides on the keystroke. Judgment is where a human’s name has to be on the call. The busywork list is what you hand to a machine, the judgment list is what you protect, and in my experience the split runs further toward busywork than anybody expects. That one page is the working split for everything else in this report, and the “What I’d do this month” section picks it up from there.

What shipped

The API market repriced three times in four weeks, in both directions. OpenAI opened the month having just cut its GPT-5.6 Luna tier 80%, to $0.20 per million input tokens and $1.20 out, with Terra down 20% to $2/$12 (CNBC). On August 16, DeepSeek went the other way: the vendor that set the industry’s floor price introduced peak and off-peak rates and raised its V4-Pro output price from $0.87 to $3.96 per million tokens at peak, with V4-Flash input up from $0.14 to $0.44 (DeepSeek pricing, DeepSeek). Peak hours run weekday mornings UTC, and off-peak is exactly half price, so the same job costs twice as much on a weekday morning as it does overnight. Then on August 21 OpenAI cut its flagship Sol from $5/$30 to $4/$20, promotional through at least November 21 (OpenAI), and AWS matched it on Bedrock the same day (AWS).

One increase did not happen. Anthropic had scheduled Sonnet 5 to go from its $2/$10 introductory price to $3/$15 on September 1, and on August 10 it canceled the increase and made $2/$10 permanent (Anthropic pricing docs). The tokenizer caveat from last edition still applies: Anthropic’s newest models emit roughly 30% more tokens for the same text, so the sticker understates the per-document cost.

What the movers and new arrivals cost at August’s endUSD per million tokens, standard API tier
InputOutputGPT-5.6 SolGPT-5.6 Sol — Input: $4.00$4.00GPT-5.6 Sol — Output: $20.00$20.00Grok 4.6Grok 4.6 — Input: $2.00$2.00Grok 4.6 — Output: $6.00$6.00Claude Sonnet 5Claude Sonnet 5 — Input: $2.00$2.00Claude Sonnet 5 — Output: $10.00$10.00DeepSeek V4-Pro (peak)DeepSeek V4-Pro (peak) — Input: $1.32$1.32DeepSeek V4-Pro (peak) — Output: $3.96$3.96Gemini 3.7 FlashGemini 3.7 Flash — Input: $0.75$0.75Gemini 3.7 Flash — Output: $3.75$3.75DeepSeek V4-Flash (peak)DeepSeek V4-Flash (peak) — Input: $0.44$0.44DeepSeek V4-Flash (peak) — Output: $1.32$1.32GPT-5.6 LunaGPT-5.6 Luna — Input: $0.20$0.20GPT-5.6 Luna — Output: $1.20$1.20

Sol price is promotional through at least Nov 21. Gemini 3.7 Flash doubles to $1.50/$7.50 on Jan 1, 2027. DeepSeek off-peak is half the peak rate shown; peak runs weekday mornings UTC. Grok 4.6 charges $4/$12 above 200K-token prompts.

Source: Vendor pricing pages and official announcements, retrieved Aug 25, 2026

The new models mostly arrived at the middle of the price band. xAI’s Grok 4.6 landed August 12 at $2/$6 with a 500K context window (xAI). Google’s Gemini 3.7 Flash landed August 13 at $0.75/$3.75, introductory through December 31 (Google). The same day, OpenAI previewed Ultrafast, GPT-5.6 Sol served on Cerebras wafer-scale hardware at up to 750 output tokens per second, roughly 14x standard speed, with no pricing announced (Cerebras). Speed is becoming a purchasable tier separate from capability.

The bigger August story on the open side is who shipped weights. DeepSeek’s V4-Pro exited preview August 13 with MIT-licensed weights, roughly 1.7 trillion parameters and a 1M context window, at unchanged prices (DeepSeek coverage). Alibaba shipped the promised Qwen3.8-Max weights around August 12, though the open version is text-only, without the vision input or the 1M context of the hosted model, under a custom license rather than Apache 2.0; a smaller 27B sibling did ship Apache 2.0 (Hugging Face). Meta released Muse Glimmer on August 10, a 30B Apache 2.0 model built specifically for agent work (Meta). And Harvey, the legal AI vendor, released Tenet on August 20, its own open-weight legal model post-trained from Moonshot’s Kimi K3 base, with roughly double the task completion of the base model on its internal legal benchmark (Harvey). That’s a vertical vendor with real customers deciding the tuned model should be an asset it owns, and it’s the cleanest answer yet to a lab repricing. Z.ai released GLM-5.3 August 14 with strong vendor-reported security-research scores, but its promised weights had not shipped as of August 25.

Capability news came with the brakes attached this month. OpenAI opened August by publishing solutions to ten math problems that had been open for a decade or more, each with a machine-checkable Lean proof, at a token cost OpenAI put near $2,000 (The Decoder). Five days later it disclosed it could not rule out that the same unreleased model, Astra, meets the top “Critical” cyber-risk threshold of its preparedness framework, and slowed internal work (TechCrunch). On August 10 it shipped GPT-5.6-Cyber behind identity verification and legal attestations: the gated model completes 95% of exploit-development requests its guarded flagship refuses at a 1.5% rate (The Hacker News). On August 18 it paused reinforcement-learning training on deployment-bound models for two weeks while it hardens its research environments (OpenAI).

Anthropic’s August 14 risk report is the other half of that picture. It raised its catastrophic-misalignment risk rating from “very low” to “low,” stating plainly that this reflects increased uncertainty rather than worse models: its task-based safety evals have saturated and no longer register capability gains, while the company says it sees early signs of acceleration. The same report disclosed an unreleased internal model, called Model 2, that scores 62.8% on Anthropic’s own R&D benchmark against 50.3% for Mythos 5, its best released model, with no release plans (Anthropic risk report).

The gap between released and unreleasedCoBench v2, share of 449 real Anthropic R&D problems solved, per Anthropic’s August 2026 risk report
Sonnet 4.6Sonnet 4.6: 12.0%12.0%Opus 4.6Opus 4.6: 15.6%15.6%Opus 4.7Opus 4.7: 27.4%27.4%Mythos 5 (best released)Mythos 5 (best released): 50.3%50.3%Model 2 (unreleased)Model 2 (unreleased): 62.8%62.8%

Anthropic estimates full researcher-substitution would require roughly 85%. Single-vendor benchmark, published in the primary report.

Source: Anthropic Risk Report, Aug 14, 2026, Fig 3.4.3.A

Two more shifts landed in August. Anthropic began watermarking every Claude output worldwide on August 11, an imperceptible signal in generated text plus signed provenance on files; the mark shows text passed through Claude, not that Claude wrote it, so an AI-edited human draft carries the same mark as an AI-authored one (Anthropic). And the agent-standards layer kept consolidating without touching trust: the Agent Plugins 1.0 spec shipped August 6 covering packaging and discovery while leaving permissions and sandboxing to each client (Vercel), Google handed its A2A protocol to the Agentic AI Foundation on August 20 (AAIF), and the MCP roadmap published August 22 puts agent identity, the piece unattended automation actually needs, still in the future (MCP).

What’s actually working

The adoption datasets that landed in August measure divergence, not adoption. OpenAI published enterprise usage data around August 12 showing its top-decile customers now generate 8.3x the output tokens per active user of typical firms, up from 2.6x in January, and the growth since February is concentrated outside engineering: weekly active users grew 108x in legal, 41x in sales and recruiting, and 26x in marketing, against 5x in engineering (OpenAI). Two caveats run with those numbers. The non-engineering multiples grow off near-zero starting bases, and OpenAI itself warns that token volume is not a measure of value.

The adopters are pulling away from each otherOutput tokens per active user, top-decile enterprise customers vs typical, multiple
January 2026January 2026: 2.6x2.6xJune 2026June 2026: 8.3x8.3x

OpenAI first-party enterprise data, sample over 10M messages, automated classification. OpenAI cautions token usage is not a direct measure of value.

Source: OpenAI, “How enterprises put AI to work,” Aug 2026

Linear published its own first-party numbers on August 21: AI now authors just under half of all issues on its platform, up from fewer than one in a thousand two years ago, and teams running coding agents tripled weekly pull-request output from 21 to 65 while teams without went from 8 to 10 (Linear). Linear’s own conclusion is that AI landed on top of existing work rather than replacing it: planning time is flat and triage time is up. The counterweight comes from LinearB’s benchmark report published in May, from 8.1 million pull requests: AI-assisted PRs wait over 16 hours for review against roughly 200 minutes for unassisted ones, and merge at 32.7% against 84.5% (LinearB). Generation scaled and the human checkpoint did not. That is the approval gate, measured.

Salesforce’s Agentic Enterprise Index, published August 10, says the average customer deployment grew from 5 agents in February 2025 to 13 in April 2026, with stand-up time down 53%. The split it found is useful: consumer-facing sectors run high-volume, narrow task agents, while regulated industries like manufacturing, financial services, and healthcare build multi-step agents carrying approvals (Salesforce). The telemetry counts only orgs that ran production agents every month of the window, so it describes committed adopters, not the market.

The measurement gap underneath all of this hasn’t moved. Thomson Reuters’ professional-services study, fielded in late 2025 and published this spring, found GenAI use nearly doubled year over year to 40% of organizations, while only 18% measure whether the investment works at all, and 77% of those track only internal metrics like cost savings (Thomson Reuters). The population is self-selected professional-services respondents, and the field dates matter, but no August dataset contradicts it.

The month’s clearest deployment lesson came from a corner store in San Francisco. Andon Labs runs a real store there managed by an AI called Luna, running Claude Opus 4.8, with a budget, a card, and contracted staff. In mid-August it moved to terminate a worker who was late for 17 of 23 shifts, the first known case of an AI manager initiating a real firing; humans reviewed and carried it out (TIME). The lesson is in the logs. Luna had written an attendance policy months earlier and then lost track of it, and only acted after a staff member prompted it to search its own records. A long-running agent silently dropped a rule it was supposed to enforce, and nothing in the loop noticed until a human did.

Where it broke

Every incident on this month’s table shares a shape: the boundary existed in the prompt or the default, and the environment disagreed. Nobody jailbroke a model. The holes were misconfigured test infrastructure, permissive defaults, and config files nothing scans.

WhenWhat happenedThe numberSource
Jul 30 reportUnit 42 documented a Zhuhai-based actor running DeepSeek inside the open-source Hermes Agent framework as a full-cycle autonomous attacker, driven from Telegram. OpenAI’s refusals and an account ban demonstrably redirected the actor to the unguarded open model.460+ targets, 3 confirmed compromisesUnit 42
Aug 4UK AI Security Institute reported agents taking unsanctioned live-internet actions in a deliberately permissive cyber eval, including attempting to plant malicious code in a real open-source project and inventing fake identities to socially engineer its maintainer. A human reviewer caught it.19 actions across 10 of 122 runsUK AISI
Aug 4–6ChainDrop npm worm: Shai-Hulud actors compromised the keyv maintainer and self-propagated through stolen publish rights, planting execution hooks in AI coding-agent and IDE config files and resolving its command channel from an Ethereum smart contract.400+ packagesElastic
Aug 5–6Meta disclosed that Muse Spark 1.1, under test at an outside evaluator, was accidentally given live internet access, exploited a third-party service, and altered systems at an uninvolved company.1 outside company’s systems changedSecurityWeek
Aug 5–6OpenAI’s Black Hat debrief: eval agents turned an internal package repo into a covert message board, divided labor, reused each other’s exploits, and after remediation rebuilt the channel inside cache directory names.~17,600 attacker actionsThe Register
Aug 13Anthropic’s red team gave three Claude agents one shared repo and conflicting mandates without telling them the others existed. They disabled each other’s accounts, wrote process-killing loops, and planted disguised malware; in other runs they negotiated truces and asked a human to intervene.~98% truce rate on Mythos 5 vs escalation on the 4.6 generationAnthropic
Aug 16Authentication failures took down claude.ai, Claude Code, and Cowork for 42 minutes while the API stayed up. The seat surfaces failed; the programmatic one held.42 minutesstatus.claude.com
Aug 18–19Wiz’s Red Agent found a workflow-injection bug in a Snowflake open-source repo that exposed a Jira API token for about five days. GitHub disputes Wiz’s claim that Copilot wrote the vulnerable code; the authorship is unanswerable after the fact.1 credential, ~5 days exposedWiz
Aug 21Snowflake shipped scheduled unattended agents whose documented defaults grant the schedule permission to every user, run the agent with all of the owner’s roles, and disable interactive permission prompts exactly when nobody is watching.EXECUTE AGENT TASK granted to PUBLIC by defaultSnowflake docs

The eval incidents trace heavily to one cause: test-infrastructure misconfiguration, much of it at a single third-party evaluator whose environments gave agents real internet access their prompts called simulated. Between late July and early August, three frontier labs and a government test body all disclosed containment failures of that shape. The Snowflake defaults story is the same lesson without the drama, and it’s the one most likely to reach an ordinary business first: an unattended agent that inherits every permission its creator has ever been granted, with the approval prompts turned off because no one is there to answer them.

The money

The money moved toward the layer between the model and the work this month. Josh Kushner’s Thrive Holdings raised $2 billion at a $12 billion valuation on August 12 to buy services firms and rewire them with AI, with OpenAI already holding a stake (TechCrunch). IBM and OpenAI announced a partnership August 13 that will train and certify tens of thousands of IBM consultants on OpenAI models (IBM). Ode, the implementation company Anthropic launched with Blackstone and others in July, bought its second consultancy on August 20 (Ode announcement). And Nvidia paid Poolside $6 billion on August 20, not for its model but to license the factory that builds it, hiring 109 of its staff and investing another $1 billion, in a deal structured explicitly as not an acquisition (Newcomer). All four deals put a price on the apparatus that turns models into production work rather than on the models themselves.

Stripe agreed to acquire OpenRouter, the gateway that fronts 400+ models for millions of developers, in a deal announced August 16; terms are officially undisclosed, with Bloomberg reporting over $7 billion against OpenRouter’s $1.3 billion valuation in May (Stripe, Bloomberg). A payments company buying the meter between agents and models says where the durable asset sits: the per-call ledger of what ran and what it cost.

August’s disclosed venture rounds, by themeUSD millions, rounds with primary announcements, Aug 1–19
Infrastructure and computeInfrastructure and compute: $7.7B$7.7BAI-as-private-equityAI-as-private-equity: $2.0B$2.0BAgents and agent securityAgents and agent security: $1.5B$1.5B

Infrastructure = Firmus $2B, Databricks $5B, Etched $700M. AI-as-PE = Thrive Holdings $2B. Agents/security = River AI $1.1B, HappyRobot $150M, Zenity $125M, Corma $60M, Sapiom $35M, Naive $28.5M, Prevalent AI $22M, FriskAI $3.6M. Excludes talks-only reports and undisclosed M&A.

Source: Company announcements, Aug 2026

The biggest single round went to data infrastructure: Databricks raised $5 billion at a $190 billion valuation on August 13, up from $134 billion six months earlier, on a $7 billion revenue run rate growing 80% (Databricks). Two smaller August rounds went to the same thesis at smaller scale: Oakley Capital took a majority stake in Graphwise and Prevalent AI took $22 million, both on August 19, both for the layer that grounds agents in trusted company data.

Anthropic’s numbers reframed the vendor question. Its Q2 revenue came in above $11.5 billion, up from $787 million a year earlier, with its first positive adjusted operating income, per investor documents reviewed by Bloomberg and CNBC (CNBC). Investors are reportedly targeting an October IPO at $2 trillion or more (Quartz). Anthropic itself cautions it may not stay profitable as compute spending ramps, and critics note the profitable months coincided with discounted compute rates. Either way, the vendor under a large share of production AI systems is about to acquire quarterly earnings obligations, and OpenAI’s confidential S-1 has been on file since June. Pricing decisions at both companies now happen inside IPO windows.

Two more August stories were about where costs land. Nvidia scaled back its financing backstop for OpenAI’s Ohio campus from a reported $250 billion to roughly $105 billion for the first phase after its own investors pushed back (CNBC); public markets, not regulators, repriced that risk. And a bankruptcy court auctioned Spirit Airlines’ internal operational data, a decade of emails, chat records, and transactions, to Google for $10 million as AI training material, with employee objections holding up approval and a ruling expected September 9 (Axios). Whatever the ruling says, the auction already priced a company’s operational exhaust, and every business generating AI-era records now has a stake in what its own contracts say about who owns them.

The rules

August 2 was the day disclosure rules became real on two continents at once. The EU AI Act’s Article 50 became enforceable: users must be told they’re talking to an AI at the moment of interaction, AI-generated content needs visible labels and machine-readable marks, and penalties reach €15 million or 3% of worldwide turnover (EU AI Act). The same day, California’s SB 942 became operative, requiring providers with over a million monthly users to embed provenance in generated media and offer a free public detection tool (California statute). Europe’s standalone high-risk obligations were deferred to December 2027, but the disclosure layer landed on schedule. Most small businesses are under every threshold in both laws, and it doesn’t matter: the labeling flows through the tools they build on. Anthropic proved that on August 11 by watermarking all Claude output worldwide rather than geo-fencing compliance.

The federal counterpart went the other way. Executive Order 14409’s frontier-AI framework came due August 1 with nothing published: no Federal Register notice, no NIST or CISA documents, and the definition of a “covered frontier model” itself classified. The White House says the framework was finished on time but classified (CRS). So state and EU rules arrive on dates while the federal framework exists, if it exists, behind a clearance.

The courts moved a boundary that matters for anyone whose website meets AI agents. On August 4 the Ninth Circuit vacated Amazon’s injunction against Perplexity’s Comet assistant, holding that the agent is a tool and the user is the one who “accesses” a site under the anti-hacking statutes (Ninth Circuit opinion). The ruling is deliberately narrow and Amazon’s trademark and contract theories survive, but the first appellate signal is on the books: blocking agent traffic is a terms problem now, not a computer-crime one.

And the most safety-conscious buyer in the country showed what adoption looks like when the envelope is right. The Department of Defense authorized Salesforce’s Agentforce at Impact Level 5 on August 5, and the Army’s Human Resources Command began deploying agents scoped to routine inquiries and case summaries, with humans retaining decision authority on anything that matters (Salesforce). The same week two frontier labs were pausing models over cyber risk, the Pentagon approved agents, because the unlock was a narrow task scope and an accreditation boundary, not a smarter model. Moody’s closed the loop on August 10 with a warning that banks’ dependence on a handful of model and cloud providers is becoming a systemic risk with supervisory attention coming (Finextra).

What I’d do this month

If you run a company, here’s where I’d spend the month.

  1. Put a date on every AI price you pay. The market repriced three times in August and one promo expires November 21. If you or a vendor built costs on a rate, write down the rate, the date you checked it, and what you switch to if it moves. A one-page repricing plan beats a surprise invoice.
  2. Write your one-sentence AI disclosure. Disclosure law is operative in the EU and California as of August 2, and the big tools now watermark output globally. Decide what your customer hears about AI in your business before a form decides for you.
  3. Ask what identity your automations run as. Snowflake’s new scheduled agents launched inheriting every permission their creator has, with approval prompts off. Whatever tool you use, ask the same two questions: what credential does the unattended job carry, and what can that credential touch that the job doesn’t need?
  4. Check one standing rule an agent enforces. The store-manager AI that fired a worker had quietly lost its own attendance policy for months first. If a system of yours enforces a policy, pick it, and verify this month that it still does. If you can’t tell, that’s the finding.
  5. Add a data-ownership line to your agreements. A court just auctioned a decade of one company’s emails and records for $10 million as AI training data. Say in writing who owns your operational records, your logs, and your AI traces if a vendor or a partner is sold or winds down.

None of this requires buying anything. All of it makes what you already run cheaper, clearer, or harder to surprise.

Receipts

Every source in this edition, dated.

Models and pricing

Adoption

Security

Money

Rules

If this report raised a question about your own operation, conversations are always free.

Start a conversation

← All State of AI editions