Skip to main content
株式会社オブライト
Services
About
Company
Column
Glossary
Pricing
Free Tools
Contact
日本語
日本語
メニューを開く
Column
VRAM
Articles tagged "VRAM"
36 articles
AI
2026-09-24
Ternary Bonsai 2 27B Requirements — 5.9GB, 1.76-Bit Ternary Build of Qwen3.8-27B Runs on 16GB RAM or a 24GB GPU
Ternary Bonsai 2 27B (PrismML, Sep 2026): Qwen3.8-27B at 1.76-bit weights, 5.9GB, 16GB RAM or a 24GB GPU, 98.2% of FP16 score. Apache 2.0, updated Sep 2026.
ローカルLLM
VRAM
Requirements
AI
2026-09-22
MiMo-V2.6 Requirements: VRAM for Pro, Flash & 9B (2026)
MiMo-V2.6 Flash needs ~170-190GB at 4-bit, ~320-350GB at FP8; Pro ~550-600GB at 4-bit; the 9B distill ~6-7GB. Updated Sep 2026: specs, benchmarks, API prices.
Open Weight LLM
MoE
Requirements
AI
2026-09-22
Atria Dawn Preview Requirements: 744B MoE VRAM & Self-Host Guide
Atria Dawn Preview needs an estimated ~400GB at 4-bit, ~780GB at FP8, or ~1,500GB at BF16. A rundown of this 744B MoE model's specs, vendor-reported benchmarks, API options, and comparison with other open MoE models.
Open Weight LLM
MoE
Requirements
AI
2026-09-21
Qwen-Image-2.1 Requirements, VRAM & How to Use (7B Unified Generation + Editing, Sep 2026)
Qwen-Image-2.1, Alibaba's open model (Sep 20, 2026), unifies 7B gen and editing with transparent PNG, 10-image edits. VRAM unofficial; non-commercial license.
Qwen
Alibaba
画像生成
AI
2026-09-14
What Is Edge0-35B-A3B? How SSD Streaming Runs a 35B MoE Model in Under 3GB of RAM
Released Sep 2026, Edge0-35B-A3B-preview 4-bit-quantizes Qwen3.5-MoE 35B-A3B and streams experts off SSD on demand, running at under 3GiB of active memory and roughly 15-18 tok/s on MLX. Here's how it works, what hardware you need, how to run it, and how it compares to llama.cpp mmap offload and Colibri.
Requirements
VRAM
ローカルLLM
AI
2026-09-14
MiniCPM5-2B Requirements: RAM, VRAM and GPU for Phones, Raspberry Pi and Laptops (2026)
As of Sep 7, 2026, OpenBMB's MiniCPM5-2B (2.52B dense, 131K context, Apache-2.0) needs roughly 1.5GB to 5GB of RAM/VRAM depending on quantization. Here is a GGUF size table, how to run it with Ollama and llama.cpp, and a comparison with Qwen3.5-4B.
MiniCPM
Requirements
VRAM
AI
2026-09-14
Nex-N2.5 Requirements: VRAM, GPU and RAM for Mini, Pro and Max (2026)
As of Sep 2026, Nex-AGI's agentic Nex-N2.5 family (Mini 35B, Pro 397B, Max 1.6T), released Sep 8, needs anywhere from roughly 22GB to over a terabyte of VRAM depending on model and quantization. Here is a per-model, per-quantization sizing table plus GPU and Apple Silicon build guidance.
Requirements
VRAM
ローカルLLM
AI
2026-09-10
DeepSeek V4.1-Flash: 552B MoE, 890-byte/token KV Cache
DeepSeek V4.1-Flash launched Sep 10, 2026: 552B MoE, 8B/16B active params, 890-byte/token KV cache, quarter prior gen, MIT license. Pricing and VRAM specs.
DeepSeek V4
Open Weight LLM
MoE
AI
2026-09-07
K2 Horizon Requirements: VRAM, GPU and RAM for All 6 Models (0.9B–375B, 2026)
As of Sep 2026, K2 Horizon from MBZUAI's Institute of Foundation Models spans 0.9B to 375B parameters, needing anywhere from under 1GB to roughly 750GB of VRAM depending on quantization. Here is a per-model, per-quantization sizing table plus GPU and Apple Silicon build guidance.
K2 Horizon
Requirements
VRAM
AI
2026-09-06
AMD Threadripper Halo Station Explained: 96-Core CPU, 576GB HBM3E AI Workstation (Sept 2026)
AMD's Threadripper Halo Station: 96-core CPU, 4 MI350P GPUs, 576GB HBM3E, 16TB/s bandwidth, targets trillion-param models at 4-bit. Price, date unannounced.
AI
ローカルLLM
VRAM
AI
2026-09-01
Solar Open2 250B Requirements: VRAM & GPU by Quant (2026)
Solar-Open2-250B (15B active MoE): GGUF IQ4_XS ~127GB, Q2_K ~89GB, Q6_K ~191GB. Official spec is 4-8x H200, but a single 96GB GPU works too. As of Sep 2026.
Solar Open2
Upstage
Requirements
AI
2026-08-31
Local LLM Context Length and VRAM: KV Cache Formula Guide
Local LLM OOM errors usually come from the KV cache, not model weights, because it grows linearly with context length. Formula, sizing table and fixes.
VRAM
ローカルLLM
MoE
AI
2026-08-31
Qwen3.8-Flash-Next Requirements — VRAM 75GB to 354GB by Quantization [125B MoE, Runs on a Single RTX 4090, Aug 2026]
Qwen3.8-Flash-Next needs roughly 75GB (1-bit) to 354GB (BF16) of memory depending on quantization. This 125B-total/6B-active open-weight MoE has been run on a single 24GB RTX 4090 via MoE expert offloading, while the official vLLM/SGLang FP8 recipe needs ~250GB across multiple datacenter GPUs. VRAM and GPU tables inside. Updated August 2026.
Qwen 3.8
Requirements
VRAM
AI
2026-08-30
Hy4 Preview Requirements — VRAM 385GB to 1.5TB [Tencent's 770B MoE, 1M Context, Apache 2.0, Aug 2026]
Hy4 preview needs about 385GB VRAM at 4-bit and 1.54TB at BF16. Tencent open-weighted this 770B MoE with 1M context in Aug 2026. VRAM tables, GPUs, API cost.
Hy4
Tencent
Requirements
AI
2026-08-27
GLM-5.3-Flash Requirements — VRAM 190GB to 740GB by Quantization [320B MoE, the Model Behind "Ox Alpha", Aug 2026]
GLM-5.3-Flash needs about 190GB VRAM at 4-bit, 740GB at BF16. VRAM and GPU tables for this 320B/18B open MoE, plus local vs API costs. Updated August 2026.
GLM-5.3
Z.ai
Requirements
AI
2026-08-24
Ornith 1.5 Requirements: VRAM, GPU and Quantization by Model Size (9B / 35B-A3B / 397B, MIT, August 2026)
Ornith 1.5 is an MIT open-weight LLM needing 8GB to 800GB VRAM. Compare GPU picks and quantized VRAM needs for 9B, 35B-A3B and 397B. As of August 2026.
Ornith
DeepReinforce
Requirements
AI
2026-08-22
DeepSeek-V4-Flash-Vision-Exp: A First Look at the New Multimodal API
DeepSeek's V4-Flash-Vision-Exp launched Aug 21, 2026 at V4-Flash pricing. On Aug 31, DeepSeek open-sourced the ~168GB FP4/FP8 weights on Hugging Face under MIT. Update covers local hardware needs and when the API still wins.
DeepSeek V4
MoE
マルチモーダル
AI
2026-08-20
Unsloth Dynamic 3.0 GGUF Quantization Explained
Unsloth Dynamic 3.0 is a GGUF quant method with fresh imatrix data, up to 10% higher accuracy at same size vs 2.0. Sizes, VRAM, setup. Updated Aug 2026.
GGUF
量子化
ローカルLLM
AI
2026-08-17
LTX-2.5 Requirements: VRAM, GPU Picks and File Sizes for the Open-Weight Video+Audio Model (2026)
LTX-2.5 (22B DiT, Gemma 4 encoder) is an open video+audio model out Aug 11, 2026. VRAM: 16-80GB; files: 35.9-71.4GB by quant. See table and GPU picks.
LTX-2.5
Requirements
VRAM
AI
2026-08-15
Qwen3.8-27B System Requirements — VRAM 9–56GB, Apache-2.0 Licensed [2026]
Qwen3.8-27B weights landed on August 15, 2026. This at-a-glance requirements guide maps VRAM needs to the actual published file sizes: Q4_K_M is 17.1GB and won't fit a 16GB GPU, making IQ4_XS (15.7GB) the practical floor. Licensed Apache-2.0 for commercial use.
Qwen 3.8
Requirements
VRAM
AI
2026-08-13
Local LLM Inference Engines Compared — llama.cpp, Ollama, vLLM, LM Studio, MLX, TensorRT-LLM
Comparing local LLM engines: Ollama and llama.cpp solo, MLX on Apple Silicon, vLLM for concurrent serving, TensorRT-LLM for NVIDIA, LM Studio for GUI trials.
ローカルLLM
ローカルAI
Ollama
AI
2026-08-07
LFM2.5-2.6B Guide: Requirements, VRAM, and Benchmarks (2026)
Liquid AI's LFM2.5-2.6B is a 2.69B on-device agent model under 2.5GB with 128K context. This guide covers VRAM sizing, benchmarks, throughput, and licensing.
Liquid AI
LFM2.5
Requirements
AI
2026-08-06
Shieldstral 1.0 3B Explained: Mistral's Open-Weight Multimodal Moderation Model
Shieldstral 1.0 3B is Mistral AI's Apache 2.0 moderation model, released Aug 4, 2026. It needs ~16GB VRAM in BF16 and screens both text and images together.
Mistral
Requirements
VRAM
AI
2026-08-03
MiniMax H3 Requirements: VRAM, GPU Sizing & File Sizes (2026 Open-Weight Video+Audio Model)
MiniMax H3, an open video+audio model, released weights Aug 3, 2026. ComfyUI needs ~42.5GB files, ~24GB VRAM (12GB may work). Covers quantization, GPU sizing.
MiniMax
Requirements
VRAM
AI
2026-08-02
WASTE: Run Kimi K3's 2.78T Params on 29GB RAM
WASTE is a dependency-free C engine running Kimi K3 (2.78T params) on 29GB RAM by streaming MoE experts from NVMe. Covers real throughput, setup, and limits.
Kimi K3
Moonshot AI
MoE
AI
2026-08-01
K-EXAONE 2.0 750B-A37B: Self-Hosting a 750B MoE (Apache 2.0)
LG AI Research's K-EXAONE 2.0 750B-A37B needs ~1.5TB weights at BF16, ~750GB at FP8, and 375-420GB at 4-bit — hardware math and Apache 2.0 self-host limits.
K-EXAONE
LG AI Research
Open Weight LLM
AI
2026-07-27
Kimi K3 Open Weights Are Out — What It Actually Takes to Self-Host a 2.8T MoE (MXFP4, ~1.4TB, vLLM/SGLang)
Moonshot released Kimi K3 open weights on July 26, 2026. At MXFP4 the weights alone are ~1.4TB — self-hosting means a multi-node cluster, not one GPU.
Kimi K3
Moonshot AI
Open Weight LLM
AI
2026-07-24
GGUF Quantization: Which Level to Pick (Q4_K_M, Q5_K_M, Q8_0, IQ) for Local LLMs
Start with Q4_K_M; step up to Q5_K_M or Q6_K if you have VRAM headroom. This guide explains GGUF naming, the quality/speed/VRAM tradeoffs per level, IQ (imatrix) quants, and how to choose by task. Updated July 2026.
GGUF
量子化
ローカルLLM
AI
2026-07-21
LongCat-2.0 Requirements — VRAM, GPU, and API Pricing for a 1.6T Open MoE Model
LongCat-2.0 is Meituan's 1.6T-parameter, MIT-licensed MoE model. Running it locally needs roughly 3,800GB in BF16 or about 970GB even at INT4 — no personal PC can hold it. Requirements tables and API pricing ($0.75/$2.95 per 1M tokens), updated July 2026.
LongCat-2.0
Requirements
VRAM
AI
2026-07-21
DeepSeek V4 Requirements Reference — VRAM, RAM & GPU by Quantization, Plus API Pricing and the July 24 Legacy Retirement [Updated for the 0731 Build]
DeepSeek V4-Flash needs roughly 160GB at 4-bit and V4-Pro about 920GB. The July 31 V4-Flash-0731 build keeps the same footprint while beating V4-Pro on agent benchmarks. VRAM, RAM and GPU tables by quantization, API pricing, and the legacy model retirement. Updated August 2026.
DeepSeek V4
Requirements
VRAM
AI
2026-07-20
NVIDIA Nemotron 3 Requirements Reference — VRAM, GPU and RAM Quick-Lookup Tables for Nano, Super and Ultra (2026)
Nemotron 3 Nano runs in about 18GB at 4-bit, Super needs 8x H100-80GB, Ultra needs 4x B200 at NVFP4. VRAM, GPU and quantization tables for NVIDIA's open-weight MoE family. Updated July 2026.
Nemotron 3
NVIDIA
Requirements
AI
2026-07-18
GLM-5.2 Requirements Reference — VRAM, RAM & GPU Quick-Lookup Tables by Quantization [753B Open-Weight MoE, 2026]
GLM-5.2 needs roughly 430GB at 4-bit and about 1.5TB at BF16 in combined memory. Quick-lookup VRAM, RAM and quantization tables for running Z.ai's 753B / ~40B-active open-weight MoE (MIT-licensed) locally. Updated July 2026.
GLM-5.2
Z.ai
Requirements
AI
2026-07-18
Inkling (Thinking Machines) Requirements Reference — VRAM, RAM & GPU Quick-Lookup Tables by Quantization [975B Open-Weight MoE, 2026]
Inkling needs from ~280GB (1-bit) to ~600GB (4-bit) combined memory, and 1.9TB at BF16. Quick-lookup VRAM, RAM, disk and quantization tables for running this 975B / 41B-active open-weight MoE locally. Updated July 2026.
Inkling
Thinking Machines
Requirements
AI
2026-07-18
Kimi K3 (Moonshot AI) Explained — 2.8T MoE Specs, API Pricing & Local Requirements (vs K2, Weights Due July 27) [2026]
Kimi K3 is a 2.8-trillion-parameter open-weight MoE (weights due July 27, 2026), ranked #1 on the frontend-code arena. API pricing is $3 input / $15 output per million tokens. Local runs are estimated at 650GB–1TB, needing server-class hardware. How it differs from K2.
Kimi K3
Moonshot AI
Open Weight LLM
AI
2026-05-25
Gemma 4 System Requirements — 5–62GB VRAM, RTX 3060 to H100 by Variant (E2B/E4B/26B/31B) [2026 Guide]
Gemma 4 needs 5GB VRAM (E2B/E4B), 16GB (26B MoE), or 24-62GB (31B Dense) depending on quantization. Requirements by model: RTX 3060 to H100, Apple Silicon M1-M4, CPU-only operation, RAM sizing, and budget builds. Updated July 2026.
Gemma 4
ハードウェア
GPU
AI
2026-04-17
Gemma 4 Complete Requirements Reference — VRAM, RAM & GPU Quick-Lookup Tables [E2B/E4B/26B/31B All Variants]
Gemma 4 minimum: 5GB RAM (E2B Q4), recommended: 24GB VRAM (31B Dense Q4). Quick-lookup tables covering VRAM, RAM, and GPU requirements for all variants: E2B, E4B, 26B MoE, and 31B Dense.
Gemma 4
Requirements
VRAM