← All articles
Sep 26, 2026

The AI Model Price War Arrived in a Two-Hour Window

Claude Opus 5.5 and GPT-6 Sol and Luna both launched September 22 with steep cuts. Here's what the AI model price war actually costs and what it means for your bill.

On September 22, Anthropic published Claude Opus 5.5 at 4:31pm. Two hours later, OpenAI followed with GPT-6 Sol and GPT-6 Luna. Both announcements led with a price cut, not a benchmark. That ordering is the story. When two labs that spent the last three years selling on capability suddenly lead with cost, the AI model price war has stopped being a metaphor and started showing up as a line item on real invoices.

What actually changed

Anthropic priced Opus 5.5 at $4 per million input tokens and $20 per million output tokens, a 20 percent cut from Opus 5. Cache reads dropped further, from $0.50 to $0.20 per million tokens, a 60 percent reduction, and cache writes fell from $6.25 to $5. Anthropic says the model performs at roughly the level of Claude Fable 5.1 on most work while costing 40 percent less to run on typical workloads and generating output more than 30 percent faster, according to Anthropic's own announcement. Subscription users on Pro, Max, Team, and Enterprise plans also got a bump to their five-hour usage limits, with a rate-limit reset they can save and use later.

OpenAI's cuts went deeper on the workhorse tier. GPT-6 Sol dropped to $2 per million input tokens and $10 per million output, down from $4 and $20 for GPT-5.6 Sol, a straight 50 percent cut. GPT-6 Luna, the smaller clerical-work model, fell from $0.20/$1.20 to $0.10/$0.50 per million tokens. Both are permanent prices, not an introductory rate, according to OpenAI's own announcement. OpenAI also claims GPT-6 Sol makes about half as many factual mistakes as its predecessor on its internal evaluation, which is built from de-identified real conversations where users flagged model errors, putting Sol at what OpenAI calls "Astra-level reliability at much lower cost."

Both companies shipped these as their second release in about three weeks. GPT-6 Astra launched September 3. Opus 5.5 is the first model since Anthropic said in an earlier release it would start pacing the frontier rather than racing every competitor announcement, and it still landed two updates deep into September.

Why price became the lever

For most of the last two years, frontier labs competed on a fairly narrow set of benchmarks: coding tasks, reasoning evals, long-context handling. That gap has been closing. Anthropic's own claim for Opus 5.5, that it matches Fable 5.1 on most work while using fewer tokens per task, is itself an admission that the ceiling on raw task performance is getting harder to move in a single release. OpenAI's framing for Sol and Luna makes the same point from the other side: both models exist specifically to bring GPT-6 Astra's improvements down to a cheaper tier rather than to push the top of the range further out.

This is the same dynamic that showed up in smaller ways earlier this month. Grok 4.7 shipped without the usual parameter-count bragging that used to accompany a flagship launch. A 2.7-trillion-parameter model from an independent developer, circulated under the name DataChaz, was built to stream its weights off an 8GB laptop's disk rather than chase a benchmark leaderboard. When the top of the market stops moving as fast, the fight shifts to who can deliver comparable output at a lower cost per token, and that is exactly what both companies did in the same two-hour window on September 22.

What it means for your token bill

If you're running AI into a marketing workflow, a content pipeline, or an agent that calls a model dozens or hundreds of times a day, the practical takeaway isn't which lab "won" this round. It's that the unit economics underneath your current setup likely just shifted, and probably without you noticing yet.

A few concrete places to check this week:

  • Re-run your actual cost per task, not the sticker price. A 40 percent per-workload savings claim from Anthropic and a 50 percent per-token cut from OpenAI don't automatically translate to the same savings on your specific pipeline, since your ratio of input to output tokens and your cache hit rate both change the math.
  • Look at whether you're over-provisioned on model tier. If GPT-6 Luna now handles clerical, high-volume tasks like summarizing and information extraction at a quarter of what you were paying for a heavier model, work you're currently routing to a bigger model for simple jobs is a place to save immediately.
  • Recheck cache configuration. Opus 5.5's cache-read price fell 60 percent on its own, separate from the base token price cut. If your workflow reuses long system prompts or repeated context, that's a larger effective saving than the headline number suggests.

The pattern to watch

Neither of these releases is really about a single model. They're a signal that as coding, reasoning, and factuality scores from different labs keep landing closer together, the practical decision for anyone building on top of these models is shifting from "which model is smartest" to "which model is cheapest for the specific job I'm actually running." That's a genuinely good problem to have if you're paying the bill, and a reason to actually re-check your setup this week rather than assume the numbers you optimized for last quarter still hold.

Sources: Anthropic, "Introducing Claude Opus 5.5"; OpenAI, "Introducing GPT-6 Sol and Luna".

Join the newsletter

AI workflows and systems, straight to your inbox.

No spam. Unsubscribe anytime.