Skip to main content
株式会社オブライト
Services
About
Company
Column
Glossary
Pricing
Free Tools
Contact
日本語
日本語
メニューを開く
Column
MoE
Articles tagged "MoE"
33 articles
AI
2026-09-14
What Is colibri? The Pure-C Engine That Runs a 744B MoE Model on 25GB RAM
A guide to JustVugg/colibri, a pure-C engine that hit GitHub Trending by running the 744B GLM-5.2 MoE model on consumer hardware via on-demand expert streaming from disk. Covers how it works, supported models, build/usage, hardware requirements, and how it differs from llama.cpp, Ollama, Edge0, and turbo-fieldfare.
colibri
ローカルLLM
MoE
AI
2026-09-14
What Is Edge0-35B-A3B? How SSD Streaming Runs a 35B MoE Model in Under 3GB of RAM
Released Sep 2026, Edge0-35B-A3B-preview 4-bit-quantizes Qwen3.5-MoE 35B-A3B and streams experts off SSD on demand, running at under 3GiB of active memory and roughly 15-18 tok/s on MLX. Here's how it works, what hardware you need, how to run it, and how it compares to llama.cpp mmap offload and Colibri.
Requirements
VRAM
ローカルLLM
AI
2026-09-14
Nex-N2.5 Requirements: VRAM, GPU and RAM for Mini, Pro and Max (2026)
As of Sep 2026, Nex-AGI's agentic Nex-N2.5 family (Mini 35B, Pro 397B, Max 1.6T), released Sep 8, needs anywhere from roughly 22GB to over a terabyte of VRAM depending on model and quantization. Here is a per-model, per-quantization sizing table plus GPU and Apple Silicon build guidance.
Requirements
VRAM
ローカルLLM
AI
2026-09-10
DeepSeek V4.1-Flash: 552B MoE, 890-byte/token KV Cache
DeepSeek V4.1-Flash launched Sep 10, 2026: 552B MoE, 8B/16B active params, 890-byte/token KV cache, quarter prior gen, MIT license. Pricing and VRAM specs.
DeepSeek V4
Open Weight LLM
MoE
AI
2026-09-07
K2 Horizon Requirements: VRAM, GPU and RAM for All 6 Models (0.9B–375B, 2026)
As of Sep 2026, K2 Horizon from MBZUAI's Institute of Foundation Models spans 0.9B to 375B parameters, needing anywhere from under 1GB to roughly 750GB of VRAM depending on quantization. Here is a per-model, per-quantization sizing table plus GPU and Apple Silicon build guidance.
K2 Horizon
Requirements
VRAM
AI
2026-09-01
Solar Open2 250B Requirements: VRAM & GPU by Quant (2026)
Solar-Open2-250B (15B active MoE): GGUF IQ4_XS ~127GB, Q2_K ~89GB, Q6_K ~191GB. Official spec is 4-8x H200, but a single 96GB GPU works too. As of Sep 2026.
Solar Open2
Upstage
Requirements
AI
2026-08-31
Local LLM Context Length and VRAM: KV Cache Formula Guide
Local LLM OOM errors usually come from the KV cache, not model weights, because it grows linearly with context length. Formula, sizing table and fixes.
VRAM
ローカルLLM
MoE
AI
2026-08-31
Qwen3.8-Flash-Next Requirements — VRAM 75GB to 354GB by Quantization [125B MoE, Runs on a Single RTX 4090, Aug 2026]
Qwen3.8-Flash-Next needs roughly 75GB (1-bit) to 354GB (BF16) of memory depending on quantization. This 125B-total/6B-active open-weight MoE has been run on a single 24GB RTX 4090 via MoE expert offloading, while the official vLLM/SGLang FP8 recipe needs ~250GB across multiple datacenter GPUs. VRAM and GPU tables inside. Updated August 2026.
Qwen 3.8
Requirements
VRAM
AI
2026-08-30
Hy4 Preview Requirements — VRAM 385GB to 1.5TB [Tencent's 770B MoE, 1M Context, Apache 2.0, Aug 2026]
Hy4 preview needs about 385GB VRAM at 4-bit and 1.54TB at BF16. Tencent open-weighted this 770B MoE with 1M context in Aug 2026. VRAM tables, GPUs, API cost.
Hy4
Tencent
Requirements
AI
2026-08-27
GLM-5.3-Flash Requirements — VRAM 190GB to 740GB by Quantization [320B MoE, the Model Behind "Ox Alpha", Aug 2026]
GLM-5.3-Flash needs about 190GB VRAM at 4-bit, 740GB at BF16. VRAM and GPU tables for this 320B/18B open MoE, plus local vs API costs. Updated August 2026.
GLM-5.3
Z.ai
Requirements
AI
2026-08-22
DeepSeek-V4-Flash-Vision-Exp: A First Look at the New Multimodal API
DeepSeek's V4-Flash-Vision-Exp launched Aug 21, 2026 at V4-Flash pricing. On Aug 31, DeepSeek open-sourced the ~168GB FP4/FP8 weights on Hugging Face under MIT. Update covers local hardware needs and when the API still wins.
DeepSeek V4
MoE
マルチモーダル
AI
2026-08-15
GLM-5.3 Explained: Z.ai's New Model vs GLM-5.2
GLM-5.3, released Aug 14, 2026, reuses the 743B GLM-5.2 base and gains from post-training alone: a claimed 50% coding jump and top open Terminal-Bench scores.
GLM-5.3
Z.ai
Open Weight LLM
AI
2026-08-14
DeepSeek V4 Pro 0813 Hits GA: Pricing, Benchmarks, Harness v0.1
DeepSeek V4 Pro hit GA (V4-Pro-0813) Aug 12, 2026 with big benchmark gains. API prices rise Aug 16, output ~2.3x. Covers pricing, self-hosting, Harness v0.1.
DeepSeek V4
MoE
Open Weight LLM
AI
2026-08-04
Qwen3.8 Max: 2.4T MoE, 95B Active, Pricing & Open Weights
Qwen3.8 Max: Alibaba's 2.4T sparse MoE, ~95B active/token. SWE-bench 87.3%. API from $2.00/$6.00/M tokens (in/out). Open weights promised. Updated Aug. 2026.
Qwen 3.8
Alibaba
MoE
AI
2026-08-02
WASTE: Run Kimi K3's 2.78T Params on 29GB RAM
WASTE is a dependency-free C engine running Kimi K3 (2.78T params) on 29GB RAM by streaming MoE experts from NVMe. Covers real throughput, setup, and limits.
Kimi K3
Moonshot AI
MoE
AI
2026-08-01
K-EXAONE 2.0 750B-A37B: Self-Hosting a 750B MoE (Apache 2.0)
LG AI Research's K-EXAONE 2.0 750B-A37B needs ~1.5TB weights at BF16, ~750GB at FP8, and 375-420GB at 4-bit — hardware math and Apache 2.0 self-host limits.
K-EXAONE
LG AI Research
Open Weight LLM
AI
2026-07-27
Kimi K3 Open Weights Are Out — What It Actually Takes to Self-Host a 2.8T MoE (MXFP4, ~1.4TB, vLLM/SGLang)
Moonshot released Kimi K3 open weights on July 26, 2026. At MXFP4 the weights alone are ~1.4TB — self-hosting means a multi-node cluster, not one GPU.
Kimi K3
Moonshot AI
Open Weight LLM
AI
2026-07-21
LongCat-2.0 Requirements — VRAM, GPU, and API Pricing for a 1.6T Open MoE Model
LongCat-2.0 is Meituan's 1.6T-parameter, MIT-licensed MoE model. Running it locally needs roughly 3,800GB in BF16 or about 970GB even at INT4 — no personal PC can hold it. Requirements tables and API pricing ($0.75/$2.95 per 1M tokens), updated July 2026.
LongCat-2.0
Requirements
VRAM
AI
2026-07-21
DeepSeek V4 Requirements Reference — VRAM, RAM & GPU by Quantization, Plus API Pricing and the July 24 Legacy Retirement [Updated for the 0731 Build]
DeepSeek V4-Flash needs roughly 160GB at 4-bit and V4-Pro about 920GB. The July 31 V4-Flash-0731 build keeps the same footprint while beating V4-Pro on agent benchmarks. VRAM, RAM and GPU tables by quantization, API pricing, and the legacy model retirement. Updated August 2026.
DeepSeek V4
Requirements
VRAM
AI
2026-07-20
NVIDIA Nemotron 3 Requirements Reference — VRAM, GPU and RAM Quick-Lookup Tables for Nano, Super and Ultra (2026)
Nemotron 3 Nano runs in about 18GB at 4-bit, Super needs 8x H100-80GB, Ultra needs 4x B200 at NVFP4. VRAM, GPU and quantization tables for NVIDIA's open-weight MoE family. Updated July 2026.
Nemotron 3
NVIDIA
Requirements
AI
2026-07-18
GLM-5.2 Requirements Reference — VRAM, RAM & GPU Quick-Lookup Tables by Quantization [753B Open-Weight MoE, 2026]
GLM-5.2 needs roughly 430GB at 4-bit and about 1.5TB at BF16 in combined memory. Quick-lookup VRAM, RAM and quantization tables for running Z.ai's 753B / ~40B-active open-weight MoE (MIT-licensed) locally. Updated July 2026.
GLM-5.2
Z.ai
Requirements
AI
2026-07-18
Inkling (Thinking Machines) Requirements Reference — VRAM, RAM & GPU Quick-Lookup Tables by Quantization [975B Open-Weight MoE, 2026]
Inkling needs from ~280GB (1-bit) to ~600GB (4-bit) combined memory, and 1.9TB at BF16. Quick-lookup VRAM, RAM, disk and quantization tables for running this 975B / 41B-active open-weight MoE locally. Updated July 2026.
Inkling
Thinking Machines
Requirements
AI
2026-07-18
Kimi K3 (Moonshot AI) Explained — 2.8T MoE Specs, API Pricing & Local Requirements (vs K2, Weights Due July 27) [2026]
Kimi K3 is a 2.8-trillion-parameter open-weight MoE (weights due July 27, 2026), ranked #1 on the frontend-code arena. API pricing is $3 input / $15 output per million tokens. Local runs are estimated at 650GB–1TB, needing server-class hardware. How it differs from K2.
Kimi K3
Moonshot AI
Open Weight LLM
AI
2026-07-08
Gemma 4 Technical Report Deep Dive — Google DeepMind's Open-Weight, Natively Multimodal 2.3B–31B LLMs with an Encoder-Free 12B Unified Design and Built-In Reasoning Mode [arXiv:2607.02770](https://arxiv.org/abs/2607.02770), Published 2026-07-02, 300+ Authors, Both Dense and MoE Variants
**Google DeepMind's Gemma Team released the Gemma 4 Technical Report as [arXiv:2607.02770](https://arxiv.org/abs/2607.02770) on 2026-07-02**. The paper introduces **2.3B / 12B / 31B parameter models**, **both Dense and MoE variants**, **natively multimodal (text / image / audio)**, a **12B encoder-free unified design** (raw audio and image patches processed directly without separate encoders), a **built-in reasoning (thinking) mode**, **improved vision / audio encoders**, **architectural refinements for inference speed, memory efficiency, and long context**, and **competitive performance against larger open models on STEM, multimodal, and long-context benchmarks**. Over 300 authors contributed. Open weights allow commercial use, distributed via Hugging Face and Ollama. Sits alongside [Qwen 3.6-35B-A3B](../columns/qwen36-35b-a3b-uncensored-abliterated-2026-07) and the [Local LLM June 2026 update](../columns/local-llm-landscape-2026-june-update) as a new chapter at the open-weights frontier. **A milestone in Google's open-weights strategy**; the encoder-free unified design departs from Qwen / Llama / DeepSeek multimodality (separate vision encoder + projection). The **reasoning mode** mirrors the extended-thinking modes of Anthropic and OpenAI closed models — the open ecosystem catching up. Caveats: commercial-license fine print, potential systemic-risk classification (EU AI Act's 10^25 FLOPs threshold), and heavy Google Cloud Vertex AI integration bias.
Gemma 4
Google DeepMind
Open Weight
AI
2026-07-05
Qwen3.6-35B-A3B Uncensored / Abliterated Deep Dive — 35B MoE / 3B Active / 262K Context / 3:1 Hybrid Linear+Softmax Attention / Native Text + Image + Video, 0/465 Refusal Rate; the Technique and Ethics of Community Uncensored Variants HauhauCS Aggressive, huihui-ai abliterated, wangzhang abliterated, prithivMLmods and Other Variants Distributed via Hugging Face and Ollama
**Qwen3.6-35B-A3B-Uncensored / Abliterated** is a family of **community-produced derivatives of Alibaba's Qwen 3.6-35B-A3B** (a 35B MoE with 3B active parameters, 262K context, and hybrid attention) with **refusal behaviors surgically removed** ([HackerNoon overview](https://hackernoon.com/qwen36-35b-a3b-uncensored-a-35b-moe-model-with-262k-context) / [HauhauCS Aggressive](https://huggingface.co/HauhauCS/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive) / [huihui-ai abliterated](https://huggingface.co/huihui-ai/Huihui-Qwen3.6-35B-A3B-abliterated) / [wangzhang abliterated](https://huggingface.co/wangzhang/Qwen3.6-35B-A3B-abliterated) / [prithivMLmods Aggressive](https://huggingface.co/prithivMLmods/Qwen3.6-35B-A3B-Uncensored-Aggressive)). **Base model specs**: **35B total / 3B active parameters** (MoE, sparse experts), **40 layers**, **hybrid attention** in a **3:1 ratio** (linear + full softmax), **native 262K-token context**, and **native text / image / video** multimodal input. Alibaba positions it as a flagship of its open-weights strategy. **The abliteration technique**: **the refusal direction is removed with LoRA-based steering** on attention and MLP projections. It layers on **Expert-Granular Abliteration (EGA)** (abliterating per-expert `down_proj` slices per layer) and **MoE router suppression** (deactivating safety experts at the router stage) — techniques adapted for the MoE architecture. HauhauCS reports **0 refusals across 465 test prompts**. The philosophy: preserve 100% of the base Qwen 3.6-35B's capability, remove only refusal. **Available variants**: - **HauhauCS-Aggressive** ([HF](https://huggingface.co/HauhauCS/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive) / [Ollama](https://ollama.com/fredrezones55/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive)): the most aggressive refusal removal - **huihui-ai Huihui-Qwen3.6-abliterated** ([HF](https://huggingface.co/huihui-ai/Huihui-Qwen3.6-35B-A3B-abliterated) / [Ollama](https://ollama.com/huihui_ai/Qwen3.6-abliterated)): from the established huihui-ai team - **wangzhang abliterated** ([HF](https://huggingface.co/wangzhang/Qwen3.6-35B-A3B-abliterated)) - **prithivMLmods Uncensored-Aggressive** ([HF](https://huggingface.co/prithivMLmods/Qwen3.6-35B-A3B-Uncensored-Aggressive)) Each ships with **quantization variants** (GGUF Q4 / Q5 / Q8 / FP16), spanning consumer GPUs (RTX 5090 32GB) up to H100 servers. **Ethical and legal considerations**: abliterated models **can produce content that Qwen would normally refuse** (illegal drugs, offensive-security code, dangerous-material synthesis, etc.). Legitimate research, jailbreak-resistance testing, roleplay, and adult-content use cases exist, but **enterprise and commercial adoption carries meaningful legal risk**. Compliance with the EU AI Act (effective August 2026) and Japanese PPC guidance is also in question. **The responsibility sits entirely with the user**; Alibaba's Qwen team is not involved. **Positioning**: alongside our [Local LLM June 2026 Update](../columns/local-llm-landscape-2026-june-update), [Kimi K2.7-Code](../columns/kimi-k2-7-code-moonshot-ai-2026-06), and [Ornith-1.0](../columns/ornith-1-0-deepreinforce-agentic-coding-2026-06), this is a case study showing that **"safety-stripping techniques" scale into the MoE era** in the open-weights ecosystem.
Qwen3.6
Uncensored
Abliterated
AI
2026-06-15
Kimi K2.7-Code Deep Dive — Moonshot AI's June 12, 2026 Coding-Specialized 1T MoE Open-Weights Model, Modified MIT License, $0.95/$4.00 per 1M, 256K Context — But Japanese Enterprises Face Two Critical Caveats (Cross-Border Data and Unverified Benchmarks)
A primary-source deep dive on **Kimi K2.7-Code**, released June 12, 2026 by Moonshot AI (Beijing). Grounded in the [Hugging Face model card](https://huggingface.co/moonshotai/Kimi-K2.7-Code), [MarkTechPost](https://www.marktechpost.com/2026/06/12/moonshot-ai-releases-kimi-k2-7-code-a-coding-model-reporting-21-8-on-kimi-code-bench-v2-over-k2-6/), and [VentureBeat's skepticism piece](https://venturebeat.com/technology/kimi-k2-7-code-cuts-thinking-tokens-30-practitioners-say-benchmarks-dont-check-out). Covers the 1T-total / 32B-active MoE architecture (384 experts, 8 routed + 1 shared), 256K context, MoonViT ~400M vision encoder, native INT4, forced-on thinking mode. License is **Modified MIT** (attribution required only above 100M MAU or $20M MRR), API pricing is $0.95 input / $0.19 cache-hit / $4.00 output per 1M tokens — roughly **1/18 of Claude Opus 4.8's output price**. OpenAI + Anthropic-compatible endpoints drop straight into Claude Code / Cursor / Aider / Cline / cmux. Moonshot self-reports **+21.8% vs K2.6 on its own Kimi Code Bench v2 and -30% reasoning tokens**, but **all public benchmarks are Moonshot's own proprietary suites; independent SWE-bench Verified / Pro / FrontierCode scores are not yet available as of June 15, 2026** (VentureBeat). For Japanese enterprises the column flags two critical caveats: **(1) both `api.moonshot.cn` and the Singapore-subsidiary-run `api.moonshot.ai` remain exposed to PRC National Intelligence Law Article 7 compelled disclosure (set against Japan's PPC DeepSeek alert of February 3, 2025 and the Digital Agency notice of February 6, 2025), and (2) the only reliable mitigation is Hugging Face self-hosting (~4-8 H100, ~595GB INT4) following the Mizuho / Lion Qwen-on-domestic-infrastructure precedent**.
Moonshot AI
Kimi K2.7-Code
Open Weight LLM
AI
2026-04-24
DeepSeek V4 Preview Released — 1.6T MoE / 1M-Token Context Open-Weight Model [April 2026]
Overview of DeepSeek V4 Preview, released on April 24, 2026: two open-weight Mixture-of-Experts variants (V4-Pro at 1.6T total / 49B active and V4-Flash at 284B / 13B), 1-million-token context, weights on Hugging Face, and rollout via API and chat — based on official information.
DeepSeek V4
オープンソースLLM
MoE
AI
2026-04-10
GLM-5.1 Complete Guide — #1 SWE-bench Pro Open-Source LLM [April 2026]
GLM-5.1 by Z.ai (released April 7, 2026) is the first open-source LLM to top SWE-bench Pro at 58.4%, surpassing GPT-5.4 (57.7%) and Claude Opus 4.6 (57.3%). This guide covers its 744B/40B-active MoE architecture, MIT license, 8-hour autonomous task capability, and setup via Ollama.
GLM-5.1
Z.ai
SWE-bench
AI
2026-04-10
Kimi K2.5 Complete Guide — 1 Trillion Parameter MIT-Licensed Open-Source LLM [2026]
Kimi K2.5, released by Moonshot AI on January 27, 2026, is a 1 trillion parameter (32B active) MoE model under the MIT License. It scores 76.8% on SWE-bench, 99.0% on HumanEval, and 87.6% on GPQA Diamond. This guide covers its architecture, hardware requirements, Ollama setup, and practical use cases.
Kimi K2.5
Moonshot AI
1兆パラメータ
AI
2026-04-10
Mistral Small 4 Complete Guide — Unified Reasoning, Multimodal & Code in 119B MoE [2026]
Mistral Small 4, released March 2026, unifies reasoning, multimodal vision, and agentic coding in a 119B MoE model under Apache 2.0. Supports 11 languages including Japanese. Full specs, setup guide, and model comparisons.
Mistral Small 4
MoE
マルチモーダル
AI
2026-04-10
MiniMax M2.5 Complete Guide — Lightning Attention Achieves 80.2% SWE-bench [2026]
MiniMax M2.5 achieves 80.2% on SWE-bench Verified using proprietary Lightning Attention in a 230B MoE model. Full breakdown of architecture, benchmarks, license terms, and setup instructions.
MiniMax M2.5
SWE-bench
Lightning Attention
AI
2026-03-17
Complete Guide to Rakuten AI 3.0 Architecture: Next-Gen Japanese LLM with MoE
A comprehensive analysis of Rakuten AI 3.0's Mixture of Experts architecture with 700B parameters. Explore the 8-expert configuration, 40B active parameter efficiency, and technical background behind achieving 8.88 on Japanese MT-Bench.
Rakuten AI 3.0
MoE
Mixture of Experts
AI
2026-03-17
NemoClaw's NIM Inference Microservices and Nemotron Models — Deployment Strategies from Edge to Cloud
A technical deep dive into NemoClaw's NIM inference microservices and Nemotron model family. We examine containerized API endpoints, elastic scaling, Nemotron 3 Super performance (120B parameters, MoE with 12B active), deployment comparisons across AWS, Azure, GCP, and on-premises, lightweight edge device operations, and partner integration use cases with Salesforce, CrowdStrike, and more.
NemoClaw
NIM
Nemotron