OpenAI cut the API price of GPT-5.6 Luna, its cheapest and fastest model tier, by 80 percent, and cut GPT-5.6 Terra, its mid-tier model for everyday work, by 20 percent. As of July 30, Luna’s input-token price fell to $0.20 per million, with output tokens now priced at $1.20 per million. Terra’s input rate dropped to $2 per million, with output tokens priced at $12 per million. OpenAI also introduced Fast mode, a faster processing option for GPT-5.6 Sol available through the API, replacing the old Priority Processing tier and running up to 2.5 times faster than Standard processing at double the price while intelligence stays the same.

An 80 percent cut on an already-cheap model is not a discount for existing spenders. It is a bid to pull new categories of work into OpenAI’s addressable market. At $0.20 per million input tokens, tasks that made no economic sense at the old rate now clear the bar: classifying every inbound support message instead of a sample, running a compliance check on every contract clause instead of a random batch, or tagging every transaction record before a human analyst ever opens the file. Work that previously got routed to a cheaper, dumber model or skipped automatically can now run on a model OpenAI says handles multi-step tool use at high quality.

Part of that pricing room traces back to a disclosure OpenAI made a day earlier: GPT-5.6 Sol had helped optimize its own inference stack, lowering the cost of serving the model family. This announcement is where that engineering work shows up on a customer’s invoice.

OpenAI backs Luna’s economics with a specific claim: on a benchmark it calls Agents’ Last Exam, Luna matches or beats a rival model, Fable 5, on professional work while its cost per completed task runs an estimated 99 percent below Fable 5’s, by OpenAI’s own math. That comparison is OpenAI’s own framing, run on OpenAI’s chosen benchmark, and the company has not published an independent replication.

Fast mode reads differently than the Luna and Terra cuts. Paying double for 2.5 times the speed, with identical intelligence, is not a capability upgrade. It is a latency tax for teams that need an answer sooner and are willing to pay for the wait to shrink. That distinction should shape how a team budgets a Sol-heavy pipeline: Fast mode buys time, not better output, and belongs on the specific steps where a slow response has a real cost, not applied as the default for every call.

OpenAI’s own example workflow spells out the intended split: use Sol to resolve ambiguity and set the plan, then hand well-specified implementation, testing, and evaluation to Luna. That two-tier pattern, expensive reasoning up front and cheap execution after, got meaningfully cheaper to run at the execution stage the moment Luna’s price dropped.

The subscription side does not move. ChatGPT Work and Codex list prices, and their quota allotments, stay the same. But each Luna or Terra call inside those products now draws down a smaller share of that quota, which functions as a quiet capacity increase for subscribers who will not see a line-item change on their bill.

For any team currently paying full API rates on high-volume classification, extraction, or first-pass drafting, the math changed on July 30. Rerun the cost model against Luna’s new $0.20 input rate before renewing a contract priced on last month’s numbers, and audit which Sol calls actually require Fast mode’s premium rather than defaulting every request to it.

OpenAI announced the GPT-5.6 Luna and Terra price cuts and the new Fast mode for Sol in a company post published July 30, 2026.