GPT-5.6 Luna 가격 80% 인하와 GPT의 역습
Key Points
- 1OpenAI recently reduced the API price of its lightweight model, GPT-5.6 Luna, by 80% to $0.20 per million input tokens.
- 2This significant cost reduction was achieved through technical optimizations, including GPU kernel rewrites and improvements to speculative decoding efficiency.
- 3The aggressive pricing strategy positions Luna as a highly competitive option against rivals like Claude Haiku and Gemini Flash, specifically targeting the high-volume enterprise AI market.
On July 30, 2026, OpenAI implemented a significant 80% price reduction for its GPT-5.6 Luna model, just three weeks after its initial release. This strategic adjustment brings the cost of the lightweight model to \$0.20 per 1 million input tokens and \$1.20 per 1 million output tokens. Concurrently, the Terra model received a 20% price cut, while the flagship Sol model remained at its original price point.
Core Methodology and Technical Optimization
OpenAI justified this aggressive pricing strategy by citing internal technical advancements that drastically improved inference efficiency:- GPU Kernel Rewriting: By re-engineering the GPU kernels responsible for inference serving, OpenAI reduced overall operational server overhead by 20%.
- Speculative Decoding Improvements: The company improved token generation efficiency by over 15% through enhanced speculative decoding. In this architecture, a smaller, faster model generates a draft sequence of tokens, which is then validated in parallel by the larger model. This process is represented by the relationship:
where represents the optimized generation time, is the time taken by the lightweight model, and is the probability of the draft being correct.
- Model-Driven Optimization: A critical innovation was the utilization of the GPT-5.6 model itself to refactor and optimize its own serving code, creating a feedback loop where the AI contributes to the reduction of its own computational costs.
Market Strategy and Competitive Positioning
The pricing shift signifies a transition in the AI industry from a pure focus on benchmark performance to aggressive cost competition, particularly targeting enterprise-level high-volume usage.- Competitive Landscape: In the lightweight category, Luna’s input cost (\$0.20/1M tokens) is approximately 5x cheaper than Anthropic’s Claude Haiku 4.5 (\$1.00/1M tokens) and significantly lower than Gemini 3.6 Flash (\$1.50/1M tokens).
- Segmented Strategy: While OpenAI is competing aggressively on price for lightweight and mid-tier models (Luna and Terra), it continues to prioritize performance over price for the flagship Sol model, which remains more expensive than comparable top-tier offerings like Claude Opus 5.
- Infrastructure Features: The introduction of a "Fast mode" offers up to 2.5x speed improvements at 2x the cost, while a new cached input price of \$0.02 per 1M tokens provides an additional 90% discount for services relying on repeated system prompts.
This move is viewed as a calculated attempt to recapture market share in the enterprise sector, where cost-sensitivity is becoming a primary factor in model selection, effectively forcing competitors to respond to a new price-performance frontier.