DwarfStar ds4: antirez's Local LLM Engine for Mac & CUDA
DwarfStar (ds4) is antirez's MIT local LLM engine for DeepSeek V4, GLM 5.3, Qwen3.8 on Mac and DGX Spark. Oct 2026 guide: memory, setup, Claude Code, speed.
DwarfStar (ds4, GitHub: antirez/ds4) is an MIT-licensed local inference engine from Salvatore Sanfilippo (antirez), the creator of Redis. It is deliberately narrow: instead of running any model, it targets a handful of open-weight models, namely DeepSeek V4/V4.1 Flash, GLM 5.2/5.3, and Qwen3.8 Flash Next, on 96-128 GB Macs (Metal), NVIDIA DGX Spark, Strix Halo, and multi-GPU CUDA boxes. It is not a general GGUF runner; it only runs the GGUF files the project itself produces. The repository was created on 2026-05-06 and has reached about 23,000 stars, and on 2026-10-03 the project site dwarfstar.sh hit the Hacker News front page under the title "From the creator of Redis; run LLM locally with ds4".
This article draws on the official docs as of October 2026 (README, MODELS, PERFORMANCE, CLIENTS, SSD_STREAMING) and covers what ds4 can do, supported hardware and memory needs, setup, connecting Claude Code, Codex CLI and OpenCode, measured speeds, how it differs from other tools, and caveats. Up front: the project itself calls the software beta quality.
What DwarfStar (ds4) is: a deliberately narrow engine
The stated goal is to be the best way to run a few excellent large language models on consumer hardware, meaning hardware people can actually own. To get there, the project writes its own small native inference engine tuned for specific models. The first priority is DeepSeek V4 Flash (including the experimental vision model), joined by DeepSeek V4.1 Flash (Metal, plus text inference on CUDA), GLM 5.2 and 5.3, GLM 5.3 Flash, DeepSeek V4 PRO, and Qwen3.8 Flash Next (Metal and CUDA).
The code is self-contained, and model loading, prompt rendering, tool calls, KV state, the HTTP server, and the coding agent are built and tested together. It does not link against GGML, but it exists thanks to the kernels, quantization formats, and GGUF ecosystem developed by llama.cpp, and the README thanks Georgi Gerganov and the other contributors explicitly. Model support is also intentionally opportunistic: a model may be removed when a better replacement arrives. The sweet spot is 128 GB laptops and 256/512 GB workstations.

What DwarfStar can do
- Run very capable models on owned hardware: MacBooks, DGX Spark, Strix Halo and similar machines
- Asymmetric 2-bit quantization: the routed experts, which dominate the model, are compressed hard (IQ2_XXS gate/up, Q2_K down) while other components stay at higher precision such as Q8 and F16/F32
- SSD streaming: run models larger than RAM by caching hot experts and reading the rest from SSD
- Disk KV cache: save compatible KV snapshots so long prompts need not be recomputed
- Three ways to run: the interactive CLI ds4, the OpenAI/Anthropic-compatible ds4-server, and the native ds4-agent that talks to the engine directly
- Vision: PNG/JPEG input with a matching encoder passed via --vision
- MTP speculative decoding: --mtp uses the built-in draft block of GLM and Qwen
- Multi-Mac setups: tensor parallelism over RDMA with two 128 GB Macs, plus pipeline parallelism to pool RAM
- Multi-GPU CUDA: Ada Lovelace cards including the L40S are supported; an eight-L40S setup reached about 126 t/s aggregate generation with 16 sessions
The README's motivations: capable open-weight models now fit on high-end personal machines, DeepSeek V4 Flash/PRO and GLM 5.2 tolerate aggressive routed-expert quantization, and compressed KV caches plus fast local SSDs make long contexts practical.
Supported hardware and memory requirements per model
There are three backends. Metal is the primary target, on Macs with 96 GB or more; smaller machines use SSD streaming. On NVIDIA CUDA, the DGX Spark is the main goal, and multi-GPU systems unsupported by other backends also work (for example DeepSeek V4 Flash on Ada Lovelace cards). ROCm targets Strix Halo systems such as the Framework Desktop. Build with make (Metal), make cuda-spark (DGX Spark), make strix-halo, or make cuda-generic.
Main download targets and sizes, per the official MODELS.md (runtime memory for context and buffers is on top):
| Target | File size | Resident memory / notes | Suggested machine |
|---|---|---|---|
ds4f-q2 (DeepSeek V4 Flash 0731) | about 81 GiB | Recommended first model | 96/128 GB Mac, DGX Spark |
ds41f-q2 (V4.1 Flash) | 341 GiB (152 GiB main + 189 GiB Engram) | Engram is always read from disk; about 81 GiB main weights per rank with TP | SSD streaming on a 128 GB Mac or Spark, or TP across two |
ds41f-q4 (V4.1 Flash) | 483 GiB (294 GiB main) | Streaming on smaller Macs; resident needs a 512 GB Mac | 512 GB Mac |
qwen38-q2 (Qwen3.8 Flash Next) | 137.10 GiB (41.73 GiB main/MTP + 95.37 GiB n-grams) | N-grams stay on disk, not loaded into RAM; start at 8K context | 64 GB Mac and up |
qwen38-q4k | 165.11 GiB | 69.74 GiB resident weights, before runtime buffers | Larger Mac, single-GPU CUDA |
glm53-q2 (GLM 5.3 Flash) | about 90 GiB | Close to a 128 GB machine's memory budget | 128 GB Mac, DGX Spark, ROCm |
glm53-q4 | about 178 GiB | Larger Mac, two 128 GB Macs, or SSD streaming | Large-memory Mac |
glm53-full-q2 (full GLM 5.3) | about 197 GiB | A sufficiently large machine or streaming | Large Mac, streaming |
pro-q2-imatrix (V4 PRO) | n/a | 512 GB resident, or SSD streaming | 512 GB Mac |
To check quickly whether a given GPU or Mac can run a model, the VRAM Calculator (free, no sign-up) estimates the VRAM you need once you pick a model, quantization, and context length.
For the underlying models, see our guides on DeepSeek V4 requirements, VRAM and pricing, GLM 5.3 Flash requirements, and Qwen3.8 Flash Next requirements. ds4 is the runtime side of that story: how to actually run them on which machine.
Installation and quick start
Getting from source to a first run takes a few commands. Clone the repo and build for your platform (the example below is Apple Silicon). For a first run on a 96 or 128 GB machine, the official advice is to download DeepSeek V4 Flash Q2. Downloads go into gguf/, and rerunning the same command resumes an interrupted download.
git clone https://github.com/antirez/ds4.git
cd ds4
# Apple Silicon (Metal)
make
# First model for 96/128 GB machines
./download_model.sh ds4f-q2Once a model is downloaded, everyday use looks like this. The default model is ds4flash.gguf, a link updated by each main-model download; pass -m FILE to choose explicitly. The server listens on http://127.0.0.1:8000 by default.
./ds4
./ds4 -p "Explain Redis streams in one paragraph."
./ds4-agent
./ds4-server --ctx 32768The interactive CLI supports /help, /read FILE, /ctx N and /quit, and Ctrl+C interrupts generation. Thinking is on by default; --nothink gives direct answers. For V4.1, --think-level 25 sets reasoning effort from 1 to 100 (0 disables thinking). Sampling defaults are temperature 1, top-p 1, and min-p 0.05, and --temp 0 selects greedy output.
Running models larger than RAM: SSD streaming
Resident inference is faster whenever the model fits; SSD streaming trades speed for capacity. It keeps a bounded cache of routed experts in memory and reads missing ones from the GGUF, but it does not remove the memory needed for other weights, activations, scratch, and context. The official advice is to start with the automatic budget; the startup report shows the effective cache and memory requirements.
# Flash Q2 on a smaller Mac
./download_model.sh ds4f-q2
./ds4 --ssd-streaming --ctx 32768 --nothink
# GLM 5.3 Flash Q4 on a 128 GB Mac
./download_model.sh glm53-q4
./ds4 --ssd-streaming --ctx 4096
# Set the cache budget explicitly
./ds4 --ssd-streaming --ssd-streaming-cache-experts 32GBReference numbers on a 128 GB M5 Max (September 6, 2026; automatic cache, no speculative decoding): GLM 5.3 Flash Q4_K (177.77 GiB) gave 121 t/s initial prefill, 104 t/s continued prefill, and 11.9 / 14.9 t/s generation (three-run median); DeepSeek Flash Vision Exp MXFP4 (145.26 GiB) gave 300 t/s, 263 t/s, and 11.9 / 19.3 t/s (single run). Generation is more sensitive to cache misses than prefill, so a model that starts can still be too slow for interactive work; the docs recommend a short test generation first. PRO Q2 on a 128 GB Mac is described as usable for inspection but slow. These are workload references, not guarantees.
One more trap: an oversized expert cache can push out the non-routed weights that every token needs and slow decoding. More cache only helps while the rest of the working set still fits.
Using ds4 with Claude Code, Codex CLI, and OpenCode
Through ds4-server, ds4 can back coding agents such as Pi, OpenCode, Codex CLI, and Claude Code. Start the server first and keep the client's context limit at or below the server's; output tokens also consume context, and a client setting does not enlarge the server allocation. The dsv4-local strings below are placeholders, not authentication.
./ds4-server --ctx 100000 --kv-disk-dir /tmp/ds4-kv --kv-disk-space-mb 8192Claude Code uses the Anthropic-compatible endpoint. The official shell wrapper points the main agent, the model roles, and subagents at the local DeepSeek V4 Flash. Give the wrapper a name different from claude.
#!/bin/sh
unset ANTHROPIC_API_KEY
export ANTHROPIC_BASE_URL="http://127.0.0.1:8000"
export ANTHROPIC_AUTH_TOKEN="dsv4-local"
export ANTHROPIC_MODEL="deepseek-v4-flash"
export ANTHROPIC_DEFAULT_SONNET_MODEL="deepseek-v4-flash"
export ANTHROPIC_DEFAULT_HAIKU_MODEL="deepseek-v4-flash"
export ANTHROPIC_DEFAULT_OPUS_MODEL="deepseek-v4-flash"
export CLAUDE_CODE_SUBAGENT_MODEL="deepseek-v4-flash"
export CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1
export CLAUDE_STREAM_IDLE_TIMEOUT_MS=600000
exec claude "$@"Codex CLI uses the Responses API: add a provider to its configuration and launch it as shown. OpenCode takes an @ai-sdk/openai-compatible provider in ~/.config/opencode/opencode.json (baseURL http://127.0.0.1:8000/v1), with ds4/deepseek-v4-flash selected as the model. Pi works similarly through ~/.pi/agent/models.json.
[model_providers.ds4]
name = "DwarfStar"
base_url = "http://127.0.0.1:8000/v1"
wire_api = "responses"
stream_idle_timeout_ms = 1000000
# launch
codex --model deepseek-v4-flash -c model_provider=ds4Agent clients send large initial prompts, so the first prefill can take a while. Enabling the disk cache with --kv-disk-dir lets later sessions reuse compatible prefixes.
Measured performance: M5 Max vs DGX Spark
These are the recorded DeepSeek V4 Flash Q2 baselines from the official PERFORMANCE.md, using 2048-token continued-prefill intervals and 128 greedy generation tokens at each frontier. They are baselines, not fresh measurements of every commit.
| Machine | Context | Prefill | Generation |
|---|---|---|---|
| M5 Max 128 GB | 2,048 | 790.18 t/s | 39.35 t/s |
| M5 Max 128 GB | 16,384 | 572.53 t/s | 36.14 t/s |
| M5 Max 128 GB | 32,768 | 557.04 t/s | 34.36 t/s |
| M5 Max 128 GB | 65,536 | 398.50 t/s | 27.64 t/s |
| DGX Spark 128 GB | 2,048 | 825.76 t/s | 18.05 t/s |
| DGX Spark 128 GB | 16,384 | 872.44 t/s | 15.10 t/s |
| DGX Spark 128 GB | 32,768 | 855.94 t/s | 14.43 t/s |
| DGX Spark 128 GB | 65,536 | 822.98 t/s | 13.84 t/s |
The pattern: the M5 Max generates faster but its prefill drops as context grows, while the DGX Spark holds prefill in the 800s even at long contexts but generates at less than half the M5 Max's rate. Which one suits you depends on whether you feed long inputs or want snappy replies. For more on running DGX Spark locally, see confidential code analysis on DGX Spark. To measure your own machine, the bundled ds4-bench reports prefill and generation per context frontier.
MTP speculative decoding and multi-machine setups
Speculative decoding is opt-in. GLM and Qwen use --mtp with a built-in draft block, so no second model file is needed; V4 Flash DSpark needs a matching support GGUF. It can improve generation, but not every workload benefits, and on unpredictable prompts it can be slower than plain decoding. See the speculative decoding doc for the difference between default opportunistic sampling and --mtp-exact-sampling.
With two 128 GB Macs connected by RDMA, you can run 4-bit DeepSeek Flash or GLM 5.3 Flash with tensor parallelism. For V4.1 Flash Q2, each rank holds about 81 GiB of main weights and both machines need the full GGUF on disk. V4.1 Q4 does not fit resident TP on two 128 GB Macs, so it needs streaming or a 512 GB Mac. Pipeline parallelism can also pool the RAM of several systems.
How ds4 differs from llama.cpp, Ollama, LM Studio, and oMLX
The core difference is a general-purpose runner versus a narrow engine specialized for a few models. ds4 builds on llama.cpp's work, so it is less a competitor than a different tool for a specific job. The comparison below sticks to what the docs support.
| Item | ds4 (DwarfStar) | llama.cpp | Ollama / LM Studio | oMLX |
|---|---|---|---|---|
| Philosophy | Narrow engine for a few models | General GGUF runner | Easy-to-use wrappers around general runners | Inference server for Apple Silicon |
| Models | DeepSeek V4 family, GLM 5.x, Qwen3.8 Flash Next, etc. | Wide range of GGUF models | Wide range | mlx-lm family models |
| Model files | Project-made GGUFs only | Standard GGUFs | Each tool's model management | MLX format |
| Platforms | Metal, CUDA, ROCm | Broad | Broad | Apple Silicon only |
| Maturity | Beta (fast-changing) | Mature | Mature | n/a |
| Large-MoE features | Asymmetric 2-bit, SSD streaming, disk KV | General | General | Tiered KV cache |
If you swap models often for everyday work, general tools like those in our Ollama vs LM Studio comparison are easier to live with. If you want continuous batching and menu-bar management on Apple Silicon, see our oMLX guide. ds4 is for people who value running DeepSeek V4 or GLM 5.x-class MoE models realistically on a 128 GB-class machine they already own.
Caveats and limitations
- Beta quality: it changes very quickly; a QA run precedes each release, but instabilities and regressions are possible
- Project GGUFs only: other GGUFs may have unsupported tensor layouts, metadata, or quantization mixes; use the download_model.sh targets
- Built with strong AI assistance: the author states openly that it was developed with strong assistance from AI coding agents, with humans leading ideas, testing, and debugging, and says that if you dislike AI-developed code, it is not for you
- A fast local SSD is assumed: V4.1 Engram tables and Qwen3.8 n-grams are always read from disk, and SSD streaming makes SSD speed decisive
- Uneven feature support: V4.1 vision is Metal-only, agent sessions with images cannot be saved with /save, and GLM requires --power 100 and rejects --prefill-chunk and an external --mtp-model
- Privacy: saved conversations and traces may contain private information; point clients only at servers you trust
- No bit-exact guarantee: scalar, batched, and tensor-parallel execution are not numerically identical
The author also expects users to adapt the software with coding agents for their own setups. The official stance is that things not documented or implemented may be easy to achieve by asking an agent.
FAQ
What hardware can run ds4?
The main targets are Apple Silicon Macs with 96 GB or more (Metal), DGX Spark, Strix Halo, and multi-GPU CUDA systems. When memory falls short, SSD streaming extends capacity at the cost of speed. The smaller Qwen3.8 Flash Next Q2 build starts at 64 GB Macs.
Can I load any GGUF like in Ollama or LM Studio?
No. DwarfStar is not a general GGUF runner and only supports the GGUFs the project builds and distributes. Other GGUFs may use unsupported tensor layouts, metadata, or quantization mixes.
Can I use it as the model behind Claude Code?
Yes. ds4-server exposes an Anthropic-compatible endpoint, so a shell wrapper that sets ANTHROPIC_BASE_URL and related variables connects Claude Code. The docs also give setups for Codex CLI (Responses API), OpenCode, and Pi. Expect the first prefill to take a while.
What are the license terms?
The repository is MIT licensed, and because some GGML/llama.cpp code is retained or adapted, the GGML authors' copyright notice stays in the LICENSE file. The weights you run carry their own licenses, so check those separately.
How fast is it?
In the official records, DeepSeek V4 Flash Q2 on a 128 GB M5 Max reached 790 t/s prefill and 39 t/s generation at 2,048 tokens of context, and 826 t/s and 18 t/s on a DGX Spark under the same conditions. SSD streaming and model size change this a lot, and the docs say these are not speed guarantees.
Summary
DwarfStar (ds4) is an MIT-licensed, purpose-built inference engine from the creator of Redis, built around running a few excellent models on hardware you can own. By combining asymmetric 2-bit quantization, SSD streaming, disk KV caches, MTP, and multi-machine setups, it aims to make large models such as DeepSeek V4 and GLM 5.x realistic on 128 GB-class Macs and the DGX Spark, and ds4-server lets it back Claude Code and Codex CLI. It is also beta, runs only project-made GGUFs, and depends on SSD speed. A practical rule: pick established general tools for versatility and stability, and consider ds4 when you specifically want to run those large models locally.
Related free tools (no sign-up, instant results)
Feel free to contact us
Contact Us