OpenAI cuts GPT 5.6 Luna price by 80%; achieves 13x cost drop in 4 months
OpenAI announced aggressive price cuts for GPT-5.6 models: Luna dropped 80% to $0.20 per million input tokens (from ~$1.20), Terra fell 20% to $2/$12 input/output pricing, and GPT-5.6 Sol introduced a new 2.5× faster mode at 2× standard price with no intelligence degradation. The cuts are tied to systems-level efficiency improvements across the model, inference stack, and agentic harness layer, including autonomous kernel optimization and speculative decoding gains.
The most striking metric: GPT-5.4 full (OpenAI's March flagship) at its peak benchmark score (51 on Arena Elo equivalent) now costs 13× more than GPT-5.6 Luna at the same performance level. Over just four months, OpenAI has compressed the cost curve by an annualized rate of ~2000× when holding intelligence constant. Downstream, ChatGPT's auto-review and Codex CLI are migrating from GPT-5.4 to Luna, yielding roughly 10× cost savings per task.
For AI infrastructure buyers and LLM platform teams: this pricing shock signals a fundamental decomposition in what 'model cost' means. Distillation, speculative decoding, inference optimization, and prompt caching are driving the marginal unit economics below what system-level efficiency alone would predict. As open models (Poolside's Laguna, Thinky's Inkling, DeepSeek) improve, the proprietar advantage shifts from raw capability to orchestration—harness design, tool routing, and context compaction that OpenAI can tune end-to-end.
Sources
- Primary source
- Latent Space AINews: GPT 5.6 Price Cuts
“Sam Altman@samamajor price cuts today: *80% drop for GPT-5.6 Luna, now $0.20 per million input tokens and $1.20 per million output *20% drop for GPT-5.6 Terra, to $2/$12 *GPT-5.6 Sol gets Fast mode in the API, up to 2.5x the speed for 2x the price”
- Latent Space: Price Curve Compression
“GPT-5.4 full at xhigh scored 51, exactly where Luna max sits today. GPT-5.4 costs $2.50/$15; Luna now costs $0.20/$1.20. In other words, roughly four months later, OpenAI is selling March's full flagship intelligence at about one-thirteenth the token price.”
- Latent Space: Optimization Details
“Self-Optimization: GPT-5.6 Sol was actively used to analyze production traffic, tune load balancing, and autonomously rewrite production kernels... This autonomous kernel optimization reduced end-to-end serving costs by 20%. Speculative Decoding: increased token-generation efficiency by over 15%.”