OpenAI released two new models, GPT-6 Sol and GPT-6 Luna, on Tuesday, just minutes after competitor Anthropic launched its flagship Claude Opus 5.5. These new entries sit below GPT-6 Astra in OpenAI’s hierarchy, designed as faster, more affordable options for everyday workloads rather than the most complex tasks. The release mirrors Anthropic’s tiered strategy, with Sol comparable to Opus and Luna to Haiku. To maintain competitiveness, OpenAI reduced API costs by 50% relative to GPT-5.6’s promotional rates. Specifically, Sol now costs $2 per million input tokens and $10 per million output tokens, down from $4 and $20 respectively. Luna pricing dropped to $0.10 for input and $0.50 for output per million tokens, down from $0.20 and $1.20.
Performance metrics highlight significant efficiency gains and cost advantages over Anthropic’s offerings. On AutomationBench, a workflow benchmark using 47 business tools, GPT-6 Sol scored a 33.2% pass rate at its highest reasoning setting for $0.27 per task. In contrast, Claude Opus 5 scored 26.9% but cost more than 11 times as much per task according to OpenAI’s data. On Agents' Last Exam, which evaluates long-term economic work across 55 sub-industries, Sol achieved a 56.4% completion score, beating Claude Opus 5’s best result at 60% lower cost. For computer use tasks on OSWorld 2.0, Sol nearly matched Claude Opus 5’s medium-effort score of 60.3% with a 60.5% result, while costing approximately 80% less. Internal coding-deception rates also improved, falling to 1.3% for Sol and 2.8% for Luna, far below the 10.4% seen in GPT-5.6 Sol. The models are currently available in ChatGPT Work and Codex for paid subscribers, with Luna accessible to free users via the desktop app.
The simultaneous launch of GPT-6 Sol and Luna immediately following Anthropic’s Claude Opus 5.5 release signals an intensifying price war focused on enterprise adoption and high-volume inference. By halving token costs and positioning these models as efficient alternatives to their premium Astra tier, OpenAI is aggressively targeting the mid-market segment where cost-per-task efficiency often outweighs marginal performance differences. The specific reduction in coding-deception rates to single digits addresses a critical pain point for developers integrating AI into production workflows, potentially lowering the operational risk associated with automated code generation and review processes.
From a market structure perspective, the direct comparison of benchmarks like AutomationBench and OSWorld 2.0 suggests that vendors are increasingly competing on verified utility rather than raw parameter counts or theoretical capabilities. The substantial cost differentials—particularly the claim that Sol offers better performance at 60% lower cost than Claude Opus 5 on certain tasks—may force other providers to accelerate their own efficiency optimizations or risk losing institutional clients who prioritize predictable expenditure. Stakeholders should monitor whether this pricing pressure leads to further consolidation among smaller model providers unable to match these infrastructure efficiencies.


