Noise Ops Engineering 4 min read

OpenAI Cuts GPT-5.6 Prices, Updates ChatGPT

OpenAI Cuts GPT-5.6 Prices, Updates ChatGPT
Why we're watching this

An 80% price cut funded partly by GPT-5.6 Sol rewriting its own serving code is a genuine efficiency milestone, landing the same week OpenAI folds its discontinued Atlas browser into the Chrome extension instead.

Key Takeaways
  • OpenAI cut API prices for GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20%, effective Thursday.
  • GPT-5.6 Sol, the flagship model, kept its price but gained a new Fast mode running up to 2.5 times faster at double the cost.
  • OpenAI said GPT-5.6 Sol itself rewrote parts of its own production serving infrastructure, cutting serving costs 20% and improving token efficiency more than 15%.
  • ChatGPT’s Chrome extension can now reference open tabs, summarize YouTube videos, and answer questions from highlighted text, as OpenAI folds in features from its soon-to-be-discontinued Atlas browser.
  • Codex’s ImageGen tool got a new lightbox and canvas for exploring and refining generated visuals directly in a coding workflow.

What Happened

OpenAI cut API prices for two of its GPT-5.6 models on Thursday, lowering GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20%, while adding a faster processing option for flagship model GPT-5.6 Sol.

Luna now costs $0.20 per million input tokens and $1.20 per million output tokens, down from $1.00 and $6.00. Terra falls to $2.00 and $12.00 per million tokens, down from $2.50 and $15.00, while Sol’s pricing stays at $5.00 and $30.00.

Sol’s new Fast mode in the API runs up to 2.5 times faster than standard processing, at double the price, replacing OpenAI’s previous Priority Processing tier. OpenAI said the price cuts were partly funded by GPT-5.6 Sol itself, which autonomously rewrote parts of its own production serving code after launch.

The changes cut serving costs by 20% and improved token-generation efficiency by more than 15%, according to OpenAI. GPT-5.6 Luna and Terra launched roughly three weeks earlier, and the new prices also apply to usage through ChatGPT Work and Codex.

Separately, OpenAI updated its ChatGPT Chrome extension this week, letting users ask about a YouTube video, reference open tabs, or highlight text on a page from a side chat panel. The update folds in features from OpenAI’s Atlas browser, which the company is discontinuing on August 9.

The desktop app also now suggests URLs as users type. OpenAI also added a new lightbox and canvas to ImageGen inside Codex this week, making it easier to explore and refine generated visuals without leaving a coding workflow.

Why It Matters

The price cuts make GPT-5.6 Terra roughly competitive with Anthropic’s introductory Claude Sonnet 5 pricing through the end of August, intensifying the price war among frontier labs on mid-tier models used for everyday production workloads. Auto-review inside ChatGPT and Codex is also moving to Luna, which OpenAI says should cut that feature’s cost roughly tenfold.

OpenAI’s claim that Sol optimized its own serving code is a notable milestone, but it is self-reported with no independent verification of how much of the price cut it actually funded versus routine infrastructure improvements. Anthropic’s comparable pricing is also introductory and scheduled to rise after August 31, so today’s parity is temporary by design.

Making advanced intelligence more abundant and affordable is central to our mission. OpenAI

Bottom Line

Watch whether Anthropic or Google respond with their own mid-tier price cuts once Claude Sonnet 5’s introductory rate expires August 31, and whether OpenAI publishes more detail on exactly how much Sol’s self-optimization contributed versus other infrastructure work. Also watch adoption of the Chrome extension now that Atlas is being wound down.

For engineering and ops teams running production workloads on GPT-5.6, per Relve, an AI tools intelligence platform, the near-term move is to shift high-volume classification and routing tasks to Luna now, and reserve Sol’s Fast mode for the small share of requests where latency actually matters.

Neelam Khan

Neelam Khan

Verified

Lead Editor

Neelam Khan is a Lead Editor at Relve, covering AI news, tools, product updates, search trends, and business use cases. She filters noise from useful signals for founders and teams, drawing on her previous work in AI SEO, content strategy, and tool research with Wellows and AllAboutAI.

Read Full Bio →