Articles

GPT-5.6 Sol, Terra, Luna: OpenAI's New Model Family

OpenAI is previewing GPT-5.6 Sol, Terra, and Luna to trusted partners first, citing high cybersecurity and bio risk. Benchmarks, pricing, and rollout.

Chisato Chisato · · 6 min read
Earth seen from orbit at night with glowing city lights

OpenAI has begun a limited preview of GPT-5.6, a three-model family that the company is releasing to a vetted set of partners and organizations before any general launch. The lineup splits into Sol, the flagship; Terra, a mid-tier model tuned for everyday work; and Luna, the fastest and cheapest of the three. What makes the rollout notable is not only the benchmark numbers but the way OpenAI is gating access — the models were shown to the US government ahead of launch, and general availability is being held back on safety grounds.

The framing echoes a pattern across the frontier: capability is advancing faster than the mechanisms meant to oversee it, a tension on full display this week as the UN’s first Global Dialogue on AI Governance convened in Geneva.

Three models, one frontier

OpenAI is positioning the GPT-5.6 series as a tiered family rather than a single flagship, a structure that lets developers trade cost against capability without leaving the ecosystem.

  • Sol is the top model — OpenAI’s most capable to date, aimed at the hardest agentic, coding, and reasoning work.
  • Terra is the balanced option, pitched at competitive performance with GPT-5.5 at roughly half the cost.
  • Luna is the speed-and-price tier, bringing what OpenAI calls strong capability at its lowest price point.

All three share a 1.5 million-token context window, a jump that narrows one of the largest advantages rivals like Google’s Gemini line had held on raw context length. The company says the family advances the frontier across software engineering, computer use, professional knowledge work, scientific research, and cybersecurity.

The headline benchmark: agentic coding

The number OpenAI led with is Terminal-Bench 2.1, a test of command-line workflows that require planning, iteration, and tool coordination — the kind of multi-step agentic work that has become the real battleground for coding models. On that benchmark, GPT-5.6 Sol scored 88.8%, rising to 91.9% in what OpenAI calls Ultra mode, which the company describes as a new state of the art for agentic coding.

Terminal-Bench is a deliberate choice. It rewards a model’s ability to carry a task across many steps — running a command, reading the output, adjusting, and continuing — rather than producing a single correct answer in one shot. That is precisely the skill profile that separates a chat assistant from an autonomous coding agent, and it is where a half-point improvement translates into a meaningfully higher rate of tasks completed without a human stepping in.

Notably, OpenAI foregrounded Terminal-Bench rather than SWE-bench Pro, the software-engineering benchmark where Anthropic’s models had led the previous generation. OpenAI has not published a Sol figure for SWE-bench Pro, a gap worth watching as independent evaluations arrive. Benchmark selection is itself a signal: labs tend to lead with the tests they win.

Pricing: cheaper capability, tiered

OpenAI set list prices per one million tokens across the three sizes:

  • Sol$5 input / $30 output
  • Terra$2.50 input / $15 output
  • Luna$1 input / $6 output

The pricing tells the strategic story as clearly as the benchmarks. Terra’s pitch is GPT-5.5-class quality at 2x lower cost, and Luna pushes a capable model to the bottom of OpenAI’s price ladder. That downward pressure on the cost of a given capability level has been the throughline of the last two years, and it reshapes what teams can afford to run at scale. For workloads that fan a model out across thousands or millions of calls, the difference between the tiers compounds fast — which is why techniques like prompt caching and careful model selection have become core to keeping AI features economical.

Why it is a preview, not a launch

The most consequential detail is not a benchmark. It is that GPT-5.6 is available only to selected trusted partners and organizations, through the OpenAI API and Codex, rather than to the general public. OpenAI says general availability will follow in the coming weeks.

The reason OpenAI gives is risk classification. Under its Preparedness Framework — the internal policy that governs how the company evaluates and releases models with dangerous-capability potential — all three GPT-5.6 models are treated as High capability in Cybersecurity and in Biological and Chemical risk. A “High” designation triggers additional safeguards and a more controlled rollout. OpenAI has said Sol is its most capable model yet for cybersecurity, including long-horizon tasks such as vulnerability research and exploitation — dual-use work that is valuable to defenders and dangerous in the wrong hands.

That is why the models were previewed to the US government ahead of a wider release. A model that materially raises the ceiling on automated vulnerability discovery, or that could assist with biological and chemical hazards, is exactly the kind of system that governance frameworks — corporate and national — are built to slow down at the point of release. The gated rollout is OpenAI applying its own tripwires.

How it fits the competitive picture

GPT-5.6 lands in a market where the top labs are shipping frontier models on a rhythm measured in weeks, not years. Anthropic recently pushed Claude Sonnet 5 across every plan tier, and Google’s Gemini line has continued to press on context length and reasoning. Each release nudges the others: a context-window jump here, a coding-benchmark record there, a price cut somewhere else.

Two structural themes stand out in this release. The first is agentic capability as the differentiator. Raw question-answering is increasingly commoditized; the value now sits in a model’s ability to plan and execute multi-step work reliably, the trait that reasoning models were built to strengthen. Terminal-Bench is a proxy for exactly that. The second is cost compression at every tier — the same capability that cost a fortune a year ago now ships at a fraction of the price, which is what pulls AI features from pilots into production.

What is still unknown

Several important things remain unverified until the models are broadly available and independently tested:

  • SWE-bench and other third-party numbers. OpenAI’s own benchmarks are a starting point, not a verdict. The comparisons that matter most will come from evaluations OpenAI does not control.
  • Real-world agentic reliability. A benchmark score of 88.8% is not the same as an 88.8% success rate on a company’s actual tasks, where messy context and long tool chains expose failure modes that clean tests do not.
  • The rollout timeline. “In the coming weeks” is a soft commitment, and a High-risk classification is precisely the kind of thing that can extend a preview.

What it means

GPT-5.6 is two stories at once. On the surface it is a routine frontier step — a new top model, a cheaper mid-tier, a bigger context window, a coding-benchmark record. Underneath it is a case study in how the safety and governance conversation is now baked into product releases.

Who benefits. Developers building agentic coding tools get a stronger flagship and, in Terra and Luna, cheaper options that make high-volume AI features more affordable. The tiered structure is a bet that most workloads do not need the top model, and that keeping developers inside one family — from cheapest to most capable — is stickier than winning any single benchmark.

The competitive read. By leading with Terminal-Bench rather than SWE-bench, OpenAI is telling the market where it believes it is strongest and steering attention toward agentic coding. Rivals will answer with their own numbers, and the 1.5-million-token context window puts direct pressure on the context-length advantage others had claimed. Expect the response within weeks, not months.

What to watch. The real signal is the gated rollout. A frontier lab holding back general availability because its own framework rates a model High for cybersecurity and biological risk is the clearest sign yet that capability and control are being weighed at the moment of release — not after. That is the same anxiety animating the UN’s Geneva dialogue and the tightening of rules like the EU AI Act. Whether the “coming weeks” promise holds — and what safeguards ship alongside general access — will say more about where the industry is heading than any benchmark on the slide.

Chisato Chisato · · 7 min read

OpenAI ChatGPT Work: The Super App Merging Codex

OpenAI merged ChatGPT and Codex into one desktop app and launched ChatGPT Work on GPT-5.6. What the super app does, pricing, and the fight with Anthropic.

#AI #OpenAI #LLM