Skip to main content
株式会社オブライト
AI2026-09-288 min read

GPT-6 Sol vs Luna vs Grok 4.7: Pricing & Benchmarks

OpenAI released GPT-6 Sol ($2/$10) and GPT-6 Luna ($0.10/$0.50 per 1M tokens) on Sep 22, 2026; xAI released Grok 4.7 ($2/$6, $4/$12 above 200K tokens) the day before. Updated Sep 2026: pricing tables, monthly cost estimates, and benchmarks, including Claude Opus 5.5 ($4/$20).


On September 22, 2026, OpenAI released GPT-6 Sol ($2 input / $10 output per 1M tokens) and GPT-6 Luna ($0.10 / $0.50). The day before, September 21, xAI released Grok 4.7 ($2 / $6 up to 200K tokens, $4 / $12 above that). Updated Sep 2026: for high-volume workloads GPT-6 Luna is by far the cheapest, general-purpose coding sits in a competitive middle tier between GPT-6 Sol and Grok 4.7, and only the hardest tasks justify routing to pricier models like Claude Opus 5.5 or GPT-6 Astra — that three-tier split is the practical takeaway as of September 2026.

What changed

- GPT-6 Sol/Luna get a permanent 50% price cut — Sol drops exactly 50% from GPT-5.6 Sol ($4/$20), and Luna drops 50% on input / 58.3% on output from GPT-5.6 Luna ($0.20/$1.20); OpenAI describes both as permanent, not promotional, pricing
- Grok 4.7 holds pricing flat while capability rises — xAI kept the same $2/$6 rate (at or below 200K tokens) as Grok 4.6 while lifting coding benchmarks like DeepSWE v1.1 and Terminal-Bench 4.0
- Three vendors shipped new models in the same week — Claude Opus 5.5 (Sept 22, $4/$20) launched the day after Grok 4.7 and roughly 90 minutes ahead of GPT-6 Sol/Luna, underscoring how fast the frontier-model price war is moving

Both new OpenAI models sit below GPT-6 Astra ($10/$50, announced Sept 3) in price, giving OpenAI a three-tier lineup (Astra / Sol / Luna), while xAI offers Grok 4.7 plus its faster Grok 4.7 Fast variant. Compared with the GPT-5.6 vs. Grok 4.5 pricing comparison from July 2026, per-token rates across the board have fallen sharply in just two months.

Pricing comparison

Rates per 1M tokens are as follows (yen figures at roughly 150 yen per dollar).

ModelInputCached inputOutputContextNotes
GPT-6 Sol$2$0.20$10~1.05M tokensReleased 2026-09-22, permanent pricing
GPT-6 Luna$0.10$0.01$0.50~1.05M tokensReleased 2026-09-22, permanent pricing
Grok 4.7 (≤200K tokens)$2$0.50$6500K tokensReleased 2026-09-21
Grok 4.7 (>200K tokens)$4$1$12500K tokensApplies to the whole request
Grok 4.7 Fast2x above—2x above500K tokensCursor / Grok Build only
Claude Opus 5.5 (reference)$4$0.20$201M tokensReleased 2026-09-22
GPT-5.6 Sol (prior, reference)$4—$20—Pre-cut pricing

Two caveats stand out. First, Grok 4.7 has its own "200K-token cliff": once a prompt crosses that threshold, rates double across the entire request (structurally similar to GPT-6 Astra's 272K-token cliff). Second, cache-read discounts differ: GPT-6 Sol/Luna cut roughly 90% off the standard rate (down to 10%), while Grok 4.7 cuts 75% (down to 25%).

Monthly cost estimate

As a reference point, assume a monthly workload of 100M input tokens plus 20M output tokens. Estimated cost, based on published rates with no cached-input discount applied, breaks down as follows.

Horizontal bar chart comparing estimated monthly cost across GPT-6 Sol/Luna, Grok 4.7 and Claude Opus 5.5 for 100M input plus 20M output tokens per month. GPT-6 Luna: $20, Grok 4.7: $320, GPT-6 Sol: $400, Claude Opus 5.5 and GPT-5.6 Sol (prior) both $800, GPT-6 Astra: $2000.
ModelEst. monthly costApprox. yenFit
GPT-6 Luna$20≈¥3,000Routine tasks, large batches
Grok 4.7 (≤200K tokens)$320≈¥48,000Agentic coding
GPT-6 Sol$400≈¥60,000General-purpose coding, research
Claude Opus 5.5$800≈¥120,000High-precision agentic tasks
GPT-5.6 Sol (prior, reference)$800≈¥120,000Pre-cut price level
GPT-6 Astra$2,000≈¥300,000Reserved for the hardest tasks

Under this estimate, Luna costs roughly 1/20th of Sol and 1/16th of Grok 4.7. That gap only pays off if accuracy holds up — a task that needs re-running loses the cost advantage — so weigh it against the benchmark data below for your specific workload.

Benchmark comparison

The key benchmark results below come from OpenAI's and xAI's own announcement materials. Measurement conditions (reasoning effort, agent execution environment, etc.) differ across vendors and even across benchmarks, so treat the figures as directional.

BenchmarkGPT-6 SolGPT-6 Luna
AutomationBench 1.0.633.2% (xhigh)20.7% (max)
DeepSWE v1.168.8% (max)66.6% (max)
FrontierCode 1.149.3% (max)42.4% (max)
OSWorld 2.064.4% (max)52.7% (max)
Agents' Last Exam v156.4% (max)50.9% (max)
Artificial Analysis Intelligence Index48 (#18/211)37
BenchmarkGrok 4.7vs. Grok 4.6
DeepSWE v1.1 (high effort)71.0%+5.8pt
CursorBench 4.046.3%+5.9pt
Senior SWE-Bench (pass@3)40.0%+1.1pt (38.9% → 40.0%)
Terminal-Bench 4.0 (with Grok Build)33%+15pt (18% → 33%)
Artificial Analysis Coding Agent Index (with Grok Build)56+9 (47 → 56)

GPT-6 Sol beats Luna on DeepSWE v1.1, FrontierCode 1.1, and OSWorld 2.0, and the Artificial Analysis Intelligence Index gap (48 vs. 37) is clear-cut. Grok 4.7, by contrast, is best read through its gain over Grok 4.6 rather than a standalone score, with the biggest jumps on Terminal-Bench 4.0 (with Grok Build, +15pt) and DeepSWE v1.1 at high effort (+5.8pt). That said, OpenAI's AutomationBench and xAI's Artificial Analysis Coding Agent Index use different evaluation methodologies and scoring, so the two benchmark families should not be read as directly comparable.

When to use which

- GPT-6 Luna fits: high-volume routine classification, summarization, and simple Q&A, where per-task cost drives the bottom line directly
- GPT-6 Sol and Grok 4.7 fit: general-purpose coding assistance and agentic tasks that need more than Luna delivers but don't justify Astra- or Opus-5.5-level spend. Grok 4.7's 500K-token context is an advantage for bulk long-document code review
- Higher-priced models fit: high-difficulty math and scientific reasoning (FrontierMath, GPQA Diamond) and low-error-tolerance work such as security research, where Opus 5.5 or Astra are worth the premium. Tooling integration matters too — see GPT-6 Astra's GitHub Copilot GA for how that plays out in practice
- Practical cost management: Grok 4.7's cliff sits at 200K tokens and GPT-6 Astra's at 272K tokens — estimate which side of the relevant boundary your prompts fall on before finalizing prompt design

Getting started with the API

All three vendors let you switch models by changing a single string in your existing SDK code.

OpenAI: GPT-6 Sol / Luna

from openai import OpenAI

client = OpenAI()

response = client.responses.create(
    model="gpt-6-sol",  # use "gpt-6-luna" for lightweight tasks
    input="Fix the bug in this code",
    reasoning={"effort": "medium"},
)
print(response.output_text)

xAI: Grok 4.7

from openai import OpenAI

# The Grok API exposes an OpenAI-compatible endpoint
client = OpenAI(api_key="XAI_API_KEY", base_url="https://api.x.ai/v1")

response = client.chat.completions.create(
    model="grok-4.7",
    messages=[
        {"role": "user", "content": "Fix the failing tests in this repository"}
    ],
    reasoning_effort="high",  # default when omitted
)
print(response.choices[0].message.content)

Anthropic: Claude Opus 5.5 (for reference)

import anthropic

client = anthropic.Anthropic()

response = client.messages.create(
    model="claude-opus-5-5",
    max_tokens=4096,
    effort="medium",
    messages=[
        {"role": "user", "content": "Run the same task on Opus 5.5 to compare cost and accuracy"}
    ],
)
print(response.content)

Caveats

All benchmark figures here come from OpenAI's, xAI's, and Anthropic's own announcement materials, not independent third-party verification. In particular, Grok 4.7's "with Grok Build" scores reflect the full agent execution environment, not the raw API model alone. Knowledge cutoffs are confirmed in official materials for GPT-6 Sol (April 20, 2026) and GPT-6 Luna (May 18, 2026), but sources vary between May and June 2026 for Grok 4.7, and we could not find an explicit date in xAI's own primary materials. Grok 4.7 Fast is not available on the public API — only through Cursor and Grok Build — and pricing and availability across all three vendors are subject to change, so check each vendor's official documentation before committing to a migration.

Frequently asked questions

How much do GPT-6 Sol and GPT-6 Luna cost?

Per OpenAI's announcement, rates per 1M tokens are $2 input / $0.20 cached input / $10 output for GPT-6 Sol, and $0.10 input / $0.01 cached input / $0.50 output for GPT-6 Luna. Both are described as permanent pricing (not promotional), roughly 50% below the prior GPT-5.6 Sol ($4/$20) and GPT-5.6 Luna ($0.20/$1.20).

What is the Grok 4.7 pricing structure?

Per xAI's announcement, rates per 1M tokens are $2 input / $0.50 cached input / $6 output for prompts up to 200K tokens, rising to $4 / $1 / $12 across the whole request once a prompt exceeds 200K tokens. These rates match the prior Grok 4.6 exactly, so capability improved with no price increase.

Which is the cheapest of GPT-6 Sol/Luna, Grok 4.7, and Claude Opus 5.5?

GPT-6 Luna is the cheapest on both per-token rate and estimated monthly cost. For a workload of 100M input plus 20M output tokens per month, Luna costs roughly $20, versus about $320 for Grok 4.7 (at or below 200K tokens), about $400 for GPT-6 Sol, and about $800 for Claude Opus 5.5.

Where can I use Grok 4.7 Fast?

Grok 4.7 Fast runs at twice the standard token rate for roughly twice the speed, but it is not available on the public xAI API — it can only be used through Cursor and Grok Build.

How should I choose between GPT-6 Sol/Luna and Claude Opus 5.5?

As a practical rule: use Luna for high-volume routine tasks where per-token cost matters most, use Sol or Grok 4.7 for general-purpose coding and agentic tasks, and reserve the pricier Opus 5.5 or GPT-6 Astra for specialized, high-stakes work where you need Fable-5.1-class capability.

Feel free to contact us

Contact Us