GPT-5.6 has launched a substantial price reduction initiative effective immediately, with reductions of up to 80%. The input and output costs for the entry-level model, Luna, have been significantly slashed from $1 and $6 per million tokens to just $0.2 and $1.2, respectively. Additionally, Luna now supports tool invocation, long-context processing, and multi-step workflow execution, enhancing its functionality. The Terra model, which serves as the primary offering, has also seen a 20% price drop, with input and output prices now set at $2 and $12, respectively. The flagship model, Sol, maintains its original pricing but introduces a new Fast mode. This mode operates up to 2.5 times faster than the standard mode, albeit at twice the cost, effectively replacing the former Priority Processing option. These price adjustments also apply to Codex and ChatGPT Work. While subscription fees remain unchanged, the token consumption for invoking Terra and Luna models has been reduced. Furthermore, the Auto-review model in both ChatGPT and Codex CLI will be replaced with GPT-5.6 Luna, with costs expected to plummet to roughly one-tenth of the original.
The driving force behind these price reductions is GPT-5.6 Sol's involvement in optimizing its own production system. Under human supervision, by rewriting and optimizing GPU Kernels and enhancing speculative decoding techniques, the end-to-end service costs have been reduced by 20%, and token generation efficiency has surged by over 15%. The Luna price reduction is particularly aimed at lowering the entry barrier for long-term Agent operation. OpenAI's strategic pricing adjustments are already yielding positive results, making advanced AI capabilities more accessible and cost-effective for a broader range of users.
