Fireworks Ember-1: Kimi K3 With 40% Fewer Tokens
Fireworks AI's Ember-1 is Kimi K3 post-trained to cut output tokens by ~40%. Covers benchmarks, pricing, API usage, and how it differs from Kimi K3 itself.
Ember-1 is a reasoning model that Fireworks AI announced on September 23, 2026. It is Moonshot AI's Kimi K3 — the 2.78-trillion-parameter MoE model covered in our Kimi K3 specs and pricing guide — post-trained to keep the same accuracy while cutting output tokens substantially. It hit the Hacker News front page on September 28 with 462 points.
What problem it solves
Reasoning models improve accuracy by generating longer chains of thought, but that inflates token counts, API costs, and latency. Dialing down reasoning effort speeds things up but sacrifices accuracy — a tradeoff that has felt unavoidable. Ember-1 targets exactly this tradeoff. Rather than simply lowering a reasoning-effort setting at inference time, Fireworks retrained the model so it produces shorter, more targeted reasoning traces on its own. Fireworks Research ran 50+ training experiments and 200+ evaluations, applying new training algorithms across math, coding, instruction following, conversation, search, tool use, and software engineering tasks. No customer data was used in training.
How much token reduction, in practice
According to Fireworks, Ember-1 uses roughly 40% fewer tokens overall than Kimi K3 for comparable responses. Across seven benchmarks, reasoning tokens dropped 35–50%. In one production customer's A/B test, output tokens per task fell from 49,300 to 29,900 — about a 39% reduction — while the task score stayed essentially flat (0.753 vs 0.751). A separate measurement cited roughly 35% fewer tokens per task at comparable quality.

Benchmark comparison
| Benchmark | Ember-1 | Kimi K3 Max | Cost change |
|---|---|---|---|
| Terminal Bench 2.1 | 82.0% | 80.9% | -51.9% |
| SWE-bench Verified | 92.2% | 93.2% | -15.5% |
| SWE-Interact | 20.0% | 21.3% | -32.5% |
| DeepSWE 1.1 | 75.2% | 66.4% | -23.7% |
| τ²-Bench Airline | 66% | 64% | -5.9% |
Ember-1 trails Kimi K3 Max slightly on SWE-bench Verified and SWE-Interact, but leads on the other three benchmarks, and cost drops meaningfully across the board. These figures are vendor-reported by Fireworks; independent third-party verification is not yet available.
Pricing and availability
Ember-1 is API-only, offered as a Research Preview on Fireworks Serverless for an initial two-week window. Weights, training code, and the exact algorithms are not released, so unlike Kimi K3 there is no way to self-host it. Pricing matches Kimi K3's API rates — $3.00 per million input tokens, $0.30 per million cached input tokens, and $15.00 per million output tokens — so savings come purely from using fewer tokens, not a lower unit price. Applying that to the A/B example: 49,300 output tokens costs about $0.74, while 29,900 tokens costs about $0.45 — roughly a 40% cost reduction per task.
- Model ID: accounts/fireworks/models/ember-1
- Context length: ~1,040k tokens (about 1M)
- Input: text and image
- Supported: function/tool calling, implicit prompt caching, Zero Data Retention / no-training-on-data options
- Not supported: fine-tuning, embeddings, reranking
- Also available via: OpenRouter (fireworks/ember-1) and reportedly Vercel AI Gateway
How to use the API
Fireworks' endpoint is OpenAI-compatible, so you can call it with the standard OpenAI Python SDK by pointing base_url at Fireworks and specifying the model ID.
```python
from openai import OpenAI
import os
client = OpenAI(
base_url="https://api.fireworks.ai/inference/v1",
api_key=os.environ["FIREWORKS_API_KEY"],
)
response = client.chat.completions.create(
model="accounts/fireworks/models/ember-1",
messages=[
{"role": "user", "content": "Review this code and suggest improvements"}
],
)
print(response.choices[0].message.content)
```How it differs from Kimi K3
Ember-1 is not Kimi K3 itself — it's a derivative model post-trained on top of K3. It shares the base model's weights with the openly released Kimi K3, but Ember-1's own weights and training code are proprietary to Fireworks, so it can't be run locally the way our Kimi K3 self-hosting guide describes. If cost and speed for high-volume agentic workloads matter most, Ember-1 is the better fit; if you need to run on your own infrastructure or want full customization freedom, the open-weight Kimi K3 remains the choice. Fireworks also offers enterprises custom token-efficient training on their own data, of which Ember-1 is effectively the general-purpose version.
Caveats
- Research preview status: terms and availability may change after the two-week window
- Vendor-reported benchmarks: no independent third-party verification yet
- Slightly lower on some metrics: SWE-bench Verified trails Kimi K3 Max
- Not self-hostable: weights are not public, so it can't run on your own infrastructure
Can I run Ember-1 locally?
No. Weights and training code aren't public, so it's only available through the Fireworks AI API. If local deployment matters, the open-weight Kimi K3 is the alternative.
Is it cheaper than Kimi K3?
The per-token price is identical ($3.00 input / $15.00 output per million tokens), but because Ember-1 uses about 40% fewer tokens for the same task, the effective cost is lower.
How does it differ from Kimi K3?
Ember-1 is Kimi K3 post-trained to produce shorter reasoning traces at similar accuracy. The base Kimi K3 itself is unchanged; Ember-1 is Fireworks' own derivative on top of it.
What happens after the preview ends?
It's currently offered for an initial two-week window as a research preview, and Fireworks hasn't announced terms or availability beyond that period.
Feel free to contact us
Contact Us