OpenAI Unveils GPT-6 Sol and Luna With Lower API Prices and Mixed Benchmark Results

OpenAI Unveils GPT-6 Sol and Luna With Lower API Prices and Mixed Benchmark Results

Image to help understand the article

OpenAI unveiled two new artificial intelligence models, GPT-6 Sol and GPT-6 Luna, on Sept. 22, 2026, positioning them as faster and less expensive options for software development and office work.

The company said it used training methods similar to those behind GPT-6 Astra to transfer improvements in professional tasks, factual accuracy, coding, computer use and model alignment to the new systems. Sol is designed for complex coding assignments, while Luna focuses on office tasks, summarization and extracting information.

Lower API prices and wider availability

OpenAI priced GPT-6 Sol at $2 per 1 million input tokens and $10 per 1 million output tokens through its application programming interface, or API. The company said that was 50% below promotional pricing for GPT-5.6.

For GPT-6 Luna, OpenAI cut the input price from 20 cents to 10 cents per 1 million tokens and the output price from $1.20 to 50 cents. The models became available through the API under the identifiers gpt-6-sol and gpt-6-luna on the day of the announcement.

OpenAI also began offering both models through ChatGPT Work and Codex for Plus, Pro, Business, Enterprise and Edu customers. Free and Go users can access Luna through the desktop app. MacRumors reported that neither model was available in Chat mode at launch.

GitHub said it would gradually add the models to GitHub Copilot across VS Code, Visual Studio, Copilot CLI, cloud agents, the Copilot app, github.com, GitHub Mobile, JetBrains, Xcode and Eclipse. Sol is slated for Copilot Pro+, Max, Business and Enterprise plans, while Luna will also be available to Pro subscribers.

OpenAI emphasizes cost and coding gains

OpenAI reported that GPT-6 Sol scored 33.2% on AutomationBench at the xhigh effort setting, at a cost of 27 cents per task. It scored 56.4% on Agents' Last Exam at maximum effort. On DeepSWE v1.1, a software engineering benchmark, Sol scored 68.8% and Luna scored 66.6% at maximum effort.

Those results did not lead every comparison. Claude Fable 5 scored 69.9% on DeepSWE v1.1 at xhigh effort, exceeding Sol's result.

On the offline version of OSWorld 2.0, OpenAI said Sol scored 60.5% at xhigh effort, compared with 60.3% for Claude Opus 5 at medium effort. OpenAI said Sol achieved a similar score at roughly 80% lower cost per task. The company also said Luna at its maximum setting could outperform GPT-5.6 Sol at medium effort while costing one-tenth as much.

In an internal factuality evaluation using anonymized real-world conversations, OpenAI said Sol produced about half as many errors as GPT-5.6 Sol. The company did not include third-party verification of that claim. It also cautioned that the conversations were selected because they were likely to prompt errors and therefore did not represent typical use, where factual mistakes are less common. OpenAI issued a similar warning for its agent and coding evaluations, saying they test deliberately difficult situations rather than measuring failure rates in ordinary use.

Independent tests show uneven progress

Artificial Analysis reported mixed results. In its Coding Agent Index, GPT-6 Sol scored 57 at the maximum setting, up from 55 for GPT-5.6 Sol. Luna fell to 41 from 43 for its predecessor.

Both models improved on the firm's AA-Omniscience Index. Sol rose from 22 to 27, while Luna moved from minus 10 to 1. The hallucination rate measured by that index declined from 92% to 60% for Sol and from 93% to 77% for Luna.

Sol also improved from 40% to 44% on Terminal-Bench 4.0, while Luna edged up from 12% to 13%. On AutomationBench-AA, Sol rose from 60% to 62% and Luna from 50% to 53%.

Results declined elsewhere. Sol lost about 100 Elo points and Luna about 75 on GDPval-AA v2.1. Luna also dropped about 45 Elo points on AA-Briefcase v1.1, while Sol was unchanged. Artificial Analysis attributed the declines to weaker presentation and responses that omitted elements required by the grading criteria, rather than a loss of underlying capability.

The evaluator estimated that Sol's maximum-setting cost per task fell by about half to $1.06, while Luna's dropped roughly 60% to 7 cents. Despite those savings, Sol gained only one point on Artificial Analysis' overall Intelligence Index, and Luna's score did not change.

Release adds to competition among U.S. AI companies

Anthropic announced Claude Opus 5.5 about 90 minutes before OpenAI introduced Sol and Luna, underscoring the rapid pace of competition among American AI developers. MacRumors reported that Anthropic's model outperformed GPT-6 Astra on some coding and knowledge-work tasks, though differences among the systems and test settings complicate direct cost-per-task comparisons.

The available results present a divided picture: GPT-6 Sol posted a higher independent coding score than its predecessor while both new models became cheaper to operate, but Luna declined on one coding index and both models lost ground on GDPval-AA. OpenAI's own DeepSWE results also left Sol behind the top reported score from Claude Fable 5.

Source: Original Korean article - Trendy News Korea

Comments