GPT-6 Sol vs Luna vs Grok 4.7: Pricing & Benchmarks
OpenAI released GPT-6 Sol ($2/$10) and GPT-6 Luna ($0.10/$0.50 per 1M tokens) on Sep 22, 2026; xAI released Grok 4.7 ($2/$6, $4/$12 above 200K tokens) the day before. Updated Sep 2026: pricing tables, monthly cost estimates, and benchmarks, including Claude Opus 5.5 ($4/$20).
On September 22, 2026, OpenAI released GPT-6 Sol ($2 input / $10 output per 1M tokens) and GPT-6 Luna ($0.10 / $0.50). The day before, September 21, xAI released Grok 4.7 ($2 / $6 up to 200K tokens, $4 / $12 above that). Updated Sep 2026: for high-volume workloads GPT-6 Luna is by far the cheapest, general-purpose coding sits in a competitive middle tier between GPT-6 Sol and Grok 4.7, and only the hardest tasks justify routing to pricier models like Claude Opus 5.5 or GPT-6 Astra — that three-tier split is the practical takeaway as of September 2026.
What changed
- GPT-6 Sol/Luna get a permanent 50% price cut — Sol drops exactly 50% from GPT-5.6 Sol ($4/$20), and Luna drops 50% on input / 58.3% on output from GPT-5.6 Luna ($0.20/$1.20); OpenAI describes both as permanent, not promotional, pricing
- Grok 4.7 holds pricing flat while capability rises — xAI kept the same $2/$6 rate (at or below 200K tokens) as Grok 4.6 while lifting coding benchmarks like DeepSWE v1.1 and Terminal-Bench 4.0
- Three vendors shipped new models in the same week — Claude Opus 5.5 (Sept 22, $4/$20) launched the day after Grok 4.7 and roughly 90 minutes ahead of GPT-6 Sol/Luna, underscoring how fast the frontier-model price war is moving
Both new OpenAI models sit below GPT-6 Astra ($10/$50, announced Sept 3) in price, giving OpenAI a three-tier lineup (Astra / Sol / Luna), while xAI offers Grok 4.7 plus its faster Grok 4.7 Fast variant. Compared with the GPT-5.6 vs. Grok 4.5 pricing comparison from July 2026, per-token rates across the board have fallen sharply in just two months.
Pricing comparison
Rates per 1M tokens are as follows (yen figures at roughly 150 yen per dollar).
| Model | Input | Cached input | Output | Context | Notes |
|---|---|---|---|---|---|
| GPT-6 Sol | $2 | $0.20 | $10 | ~1.05M tokens | Released 2026-09-22, permanent pricing |
| GPT-6 Luna | $0.10 | $0.01 | $0.50 | ~1.05M tokens | Released 2026-09-22, permanent pricing |
| Grok 4.7 (≤200K tokens) | $2 | $0.50 | $6 | 500K tokens | Released 2026-09-21 |
| Grok 4.7 (>200K tokens) | $4 | $1 | $12 | 500K tokens | Applies to the whole request |
| Grok 4.7 Fast | 2x above | — | 2x above | 500K tokens | Cursor / Grok Build only |
| Claude Opus 5.5 (reference) | $4 | $0.20 | $20 | 1M tokens | Released 2026-09-22 |
| GPT-5.6 Sol (prior, reference) | $4 | — | $20 | — | Pre-cut pricing |
Two caveats stand out. First, Grok 4.7 has its own "200K-token cliff": once a prompt crosses that threshold, rates double across the entire request (structurally similar to GPT-6 Astra's 272K-token cliff). Second, cache-read discounts differ: GPT-6 Sol/Luna cut roughly 90% off the standard rate (down to 10%), while Grok 4.7 cuts 75% (down to 25%).
Monthly cost estimate
As a reference point, assume a monthly workload of 100M input tokens plus 20M output tokens. Estimated cost, based on published rates with no cached-input discount applied, breaks down as follows.

| Model | Est. monthly cost | Approx. yen | Fit |
|---|---|---|---|
| GPT-6 Luna | $20 | ≈¥3,000 | Routine tasks, large batches |
| Grok 4.7 (≤200K tokens) | $320 | ≈¥48,000 | Agentic coding |
| GPT-6 Sol | $400 | ≈¥60,000 | General-purpose coding, research |
| Claude Opus 5.5 | $800 | ≈¥120,000 | High-precision agentic tasks |
| GPT-5.6 Sol (prior, reference) | $800 | ≈¥120,000 | Pre-cut price level |
| GPT-6 Astra | $2,000 | ≈¥300,000 | Reserved for the hardest tasks |
Under this estimate, Luna costs roughly 1/20th of Sol and 1/16th of Grok 4.7. That gap only pays off if accuracy holds up — a task that needs re-running loses the cost advantage — so weigh it against the benchmark data below for your specific workload.
Benchmark comparison
The key benchmark results below come from OpenAI's and xAI's own announcement materials. Measurement conditions (reasoning effort, agent execution environment, etc.) differ across vendors and even across benchmarks, so treat the figures as directional.
| Benchmark | GPT-6 Sol | GPT-6 Luna |
|---|---|---|
| AutomationBench 1.0.6 | 33.2% (xhigh) | 20.7% (max) |
| DeepSWE v1.1 | 68.8% (max) | 66.6% (max) |
| FrontierCode 1.1 | 49.3% (max) | 42.4% (max) |
| OSWorld 2.0 | 64.4% (max) | 52.7% (max) |
| Agents' Last Exam v1 | 56.4% (max) | 50.9% (max) |
| Artificial Analysis Intelligence Index | 48 (#18/211) | 37 |
| Benchmark | Grok 4.7 | vs. Grok 4.6 |
|---|---|---|
| DeepSWE v1.1 (high effort) | 71.0% | +5.8pt |
| CursorBench 4.0 | 46.3% | +5.9pt |
| Senior SWE-Bench (pass@3) | 40.0% | +1.1pt (38.9% → 40.0%) |
| Terminal-Bench 4.0 (with Grok Build) | 33% | +15pt (18% → 33%) |
| Artificial Analysis Coding Agent Index (with Grok Build) | 56 | +9 (47 → 56) |
GPT-6 Sol beats Luna on DeepSWE v1.1, FrontierCode 1.1, and OSWorld 2.0, and the Artificial Analysis Intelligence Index gap (48 vs. 37) is clear-cut. Grok 4.7, by contrast, is best read through its gain over Grok 4.6 rather than a standalone score, with the biggest jumps on Terminal-Bench 4.0 (with Grok Build, +15pt) and DeepSWE v1.1 at high effort (+5.8pt). That said, OpenAI's AutomationBench and xAI's Artificial Analysis Coding Agent Index use different evaluation methodologies and scoring, so the two benchmark families should not be read as directly comparable.
When to use which
- GPT-6 Luna fits: high-volume routine classification, summarization, and simple Q&A, where per-task cost drives the bottom line directly
- GPT-6 Sol and Grok 4.7 fit: general-purpose coding assistance and agentic tasks that need more than Luna delivers but don't justify Astra- or Opus-5.5-level spend. Grok 4.7's 500K-token context is an advantage for bulk long-document code review
- Higher-priced models fit: high-difficulty math and scientific reasoning (FrontierMath, GPQA Diamond) and low-error-tolerance work such as security research, where Opus 5.5 or Astra are worth the premium. Tooling integration matters too — see GPT-6 Astra's GitHub Copilot GA for how that plays out in practice
- Practical cost management: Grok 4.7's cliff sits at 200K tokens and GPT-6 Astra's at 272K tokens — estimate which side of the relevant boundary your prompts fall on before finalizing prompt design
Getting started with the API
All three vendors let you switch models by changing a single string in your existing SDK code.
OpenAI: GPT-6 Sol / Luna
from openai import OpenAI
client = OpenAI()
response = client.responses.create(
model="gpt-6-sol", # use "gpt-6-luna" for lightweight tasks
input="Fix the bug in this code",
reasoning={"effort": "medium"},
)
print(response.output_text)xAI: Grok 4.7
from openai import OpenAI
# The Grok API exposes an OpenAI-compatible endpoint
client = OpenAI(api_key="XAI_API_KEY", base_url="https://api.x.ai/v1")
response = client.chat.completions.create(
model="grok-4.7",
messages=[
{"role": "user", "content": "Fix the failing tests in this repository"}
],
reasoning_effort="high", # default when omitted
)
print(response.choices[0].message.content)Anthropic: Claude Opus 5.5 (for reference)
import anthropic
client = anthropic.Anthropic()
response = client.messages.create(
model="claude-opus-5-5",
max_tokens=4096,
effort="medium",
messages=[
{"role": "user", "content": "Run the same task on Opus 5.5 to compare cost and accuracy"}
],
)
print(response.content)Caveats
All benchmark figures here come from OpenAI's, xAI's, and Anthropic's own announcement materials, not independent third-party verification. In particular, Grok 4.7's "with Grok Build" scores reflect the full agent execution environment, not the raw API model alone. Knowledge cutoffs are confirmed in official materials for GPT-6 Sol (April 20, 2026) and GPT-6 Luna (May 18, 2026), but sources vary between May and June 2026 for Grok 4.7, and we could not find an explicit date in xAI's own primary materials. Grok 4.7 Fast is not available on the public API — only through Cursor and Grok Build — and pricing and availability across all three vendors are subject to change, so check each vendor's official documentation before committing to a migration.
Frequently asked questions
How much do GPT-6 Sol and GPT-6 Luna cost?
Per OpenAI's announcement, rates per 1M tokens are $2 input / $0.20 cached input / $10 output for GPT-6 Sol, and $0.10 input / $0.01 cached input / $0.50 output for GPT-6 Luna. Both are described as permanent pricing (not promotional), roughly 50% below the prior GPT-5.6 Sol ($4/$20) and GPT-5.6 Luna ($0.20/$1.20).
What is the Grok 4.7 pricing structure?
Per xAI's announcement, rates per 1M tokens are $2 input / $0.50 cached input / $6 output for prompts up to 200K tokens, rising to $4 / $1 / $12 across the whole request once a prompt exceeds 200K tokens. These rates match the prior Grok 4.6 exactly, so capability improved with no price increase.
Which is the cheapest of GPT-6 Sol/Luna, Grok 4.7, and Claude Opus 5.5?
GPT-6 Luna is the cheapest on both per-token rate and estimated monthly cost. For a workload of 100M input plus 20M output tokens per month, Luna costs roughly $20, versus about $320 for Grok 4.7 (at or below 200K tokens), about $400 for GPT-6 Sol, and about $800 for Claude Opus 5.5.
Where can I use Grok 4.7 Fast?
Grok 4.7 Fast runs at twice the standard token rate for roughly twice the speed, but it is not available on the public xAI API — it can only be used through Cursor and Grok Build.
How should I choose between GPT-6 Sol/Luna and Claude Opus 5.5?
As a practical rule: use Luna for high-volume routine tasks where per-token cost matters most, use Sol or Grok 4.7 for general-purpose coding and agentic tasks, and reserve the pricier Opus 5.5 or GPT-6 Astra for specialized, high-stakes work where you need Fable-5.1-class capability.
Feel free to contact us
Contact Us