Skip to main content
株式会社オブライト

Articles tagged "Open Weight LLM"

30 articles

AI2026-09-14
MiniCPM5-2B Requirements: RAM, VRAM and GPU for Phones, Raspberry Pi and Laptops (2026)
As of Sep 7, 2026, OpenBMB's MiniCPM5-2B (2.52B dense, 131K context, Apache-2.0) needs roughly 1.5GB to 5GB of RAM/VRAM depending on quantization. Here is a GGUF size table, how to run it with Ollama and llama.cpp, and a comparison with Qwen3.5-4B.
MiniCPMRequirementsVRAM
AI2026-09-14
Nex-N2.5 Requirements: VRAM, GPU and RAM for Mini, Pro and Max (2026)
As of Sep 2026, Nex-AGI's agentic Nex-N2.5 family (Mini 35B, Pro 397B, Max 1.6T), released Sep 8, needs anywhere from roughly 22GB to over a terabyte of VRAM depending on model and quantization. Here is a per-model, per-quantization sizing table plus GPU and Apple Silicon build guidance.
RequirementsVRAMローカルLLM
AI2026-09-10
DeepSeek V4.1-Flash: 552B MoE, 890-byte/token KV Cache
DeepSeek V4.1-Flash launched Sep 10, 2026: 552B MoE, 8B/16B active params, 890-byte/token KV cache, quarter prior gen, MIT license. Pricing and VRAM specs.
DeepSeek V4Open Weight LLMMoE
AI2026-09-07
K2 Horizon Requirements: VRAM, GPU and RAM for All 6 Models (0.9B–375B, 2026)
As of Sep 2026, K2 Horizon from MBZUAI's Institute of Foundation Models spans 0.9B to 375B parameters, needing anywhere from under 1GB to roughly 750GB of VRAM depending on quantization. Here is a per-model, per-quantization sizing table plus GPU and Apple Silicon build guidance.
K2 HorizonRequirementsVRAM
AI2026-09-01
Solar Open2 250B Requirements: VRAM & GPU by Quant (2026)
Solar-Open2-250B (15B active MoE): GGUF IQ4_XS ~127GB, Q2_K ~89GB, Q6_K ~191GB. Official spec is 4-8x H200, but a single 96GB GPU works too. As of Sep 2026.
Solar Open2UpstageRequirements
AI2026-08-31
Qwen3.8-Flash-Next Requirements — VRAM 75GB to 354GB by Quantization [125B MoE, Runs on a Single RTX 4090, Aug 2026]
Qwen3.8-Flash-Next needs roughly 75GB (1-bit) to 354GB (BF16) of memory depending on quantization. This 125B-total/6B-active open-weight MoE has been run on a single 24GB RTX 4090 via MoE expert offloading, while the official vLLM/SGLang FP8 recipe needs ~250GB across multiple datacenter GPUs. VRAM and GPU tables inside. Updated August 2026.
Qwen 3.8RequirementsVRAM
AI2026-08-30
Hy4 Preview Requirements — VRAM 385GB to 1.5TB [Tencent's 770B MoE, 1M Context, Apache 2.0, Aug 2026]
Hy4 preview needs about 385GB VRAM at 4-bit and 1.54TB at BF16. Tencent open-weighted this 770B MoE with 1M context in Aug 2026. VRAM tables, GPUs, API cost.
Hy4TencentRequirements
AI2026-08-27
GLM-5.3-Flash Requirements — VRAM 190GB to 740GB by Quantization [320B MoE, the Model Behind "Ox Alpha", Aug 2026]
GLM-5.3-Flash needs about 190GB VRAM at 4-bit, 740GB at BF16. VRAM and GPU tables for this 320B/18B open MoE, plus local vs API costs. Updated August 2026.
GLM-5.3Z.aiRequirements
AI2026-08-24
Ornith 1.5 Requirements: VRAM, GPU and Quantization by Model Size (9B / 35B-A3B / 397B, MIT, August 2026)
Ornith 1.5 is an MIT open-weight LLM needing 8GB to 800GB VRAM. Compare GPU picks and quantized VRAM needs for 9B, 35B-A3B and 397B. As of August 2026.
OrnithDeepReinforceRequirements
AI2026-08-22
DeepSeek-V4-Flash-Vision-Exp: A First Look at the New Multimodal API
DeepSeek's V4-Flash-Vision-Exp launched Aug 21, 2026 at V4-Flash pricing. On Aug 31, DeepSeek open-sourced the ~168GB FP4/FP8 weights on Hugging Face under MIT. Update covers local hardware needs and when the API still wins.
DeepSeek V4MoEマルチモーダル
AI2026-08-20
Unsloth Dynamic 3.0 GGUF Quantization Explained
Unsloth Dynamic 3.0 is a GGUF quant method with fresh imatrix data, up to 10% higher accuracy at same size vs 2.0. Sizes, VRAM, setup. Updated Aug 2026.
GGUF量子化ローカルLLM
AI2026-08-15
GLM-5.3 Explained: Z.ai's New Model vs GLM-5.2
GLM-5.3, released Aug 14, 2026, reuses the 743B GLM-5.2 base and gains from post-training alone: a claimed 50% coding jump and top open Terminal-Bench scores.
GLM-5.3Z.aiOpen Weight LLM
AI2026-08-15
Qwen3.8-27B System Requirements — VRAM 9–56GB, Apache-2.0 Licensed [2026]
Qwen3.8-27B weights landed on August 15, 2026. This at-a-glance requirements guide maps VRAM needs to the actual published file sizes: Q4_K_M is 17.1GB and won't fit a 16GB GPU, making IQ4_XS (15.7GB) the practical floor. Licensed Apache-2.0 for commercial use.
Qwen 3.8RequirementsVRAM
AI2026-08-14
DeepSeek V4 Pro 0813 Hits GA: Pricing, Benchmarks, Harness v0.1
DeepSeek V4 Pro hit GA (V4-Pro-0813) Aug 12, 2026 with big benchmark gains. API prices rise Aug 16, output ~2.3x. Covers pricing, self-hosting, Harness v0.1.
DeepSeek V4MoEOpen Weight LLM
AI2026-08-09
DOE Genesis Open Models Initiative Explained: Genesis-Science-1
DOE's Genesis Open Models Initiative, announced Aug 7, 2026, builds open-weight AI for science. First model: Genesis-Science-1, co-developed with Arcee AI.
Open Weight LLMオープンソースLLMAI業界動向
AI2026-08-07
LFM2.5-2.6B Guide: Requirements, VRAM, and Benchmarks (2026)
Liquid AI's LFM2.5-2.6B is a 2.69B on-device agent model under 2.5GB with 128K context. This guide covers VRAM sizing, benchmarks, throughput, and licensing.
Liquid AILFM2.5Requirements
AI2026-08-06
Shieldstral 1.0 3B Explained: Mistral's Open-Weight Multimodal Moderation Model
Shieldstral 1.0 3B is Mistral AI's Apache 2.0 moderation model, released Aug 4, 2026. It needs ~16GB VRAM in BF16 and screens both text and images together.
MistralRequirementsVRAM
AI2026-08-04
Qwen3.8 Max: 2.4T MoE, 95B Active, Pricing & Open Weights
Qwen3.8 Max: Alibaba's 2.4T sparse MoE, ~95B active/token. SWE-bench 87.3%. API from $2.00/$6.00/M tokens (in/out). Open weights promised. Updated Aug. 2026.
Qwen 3.8AlibabaMoE
AI2026-08-03
MiniMax H3 Requirements: VRAM, GPU Sizing & File Sizes (2026 Open-Weight Video+Audio Model)
MiniMax H3, an open video+audio model, released weights Aug 3, 2026. ComfyUI needs ~42.5GB files, ~24GB VRAM (12GB may work). Covers quantization, GPU sizing.
MiniMaxRequirementsVRAM
AI2026-08-02
WASTE: Run Kimi K3's 2.78T Params on 29GB RAM
WASTE is a dependency-free C engine running Kimi K3 (2.78T params) on 29GB RAM by streaming MoE experts from NVMe. Covers real throughput, setup, and limits.
Kimi K3Moonshot AIMoE
AI2026-08-01
K-EXAONE 2.0 750B-A37B: Self-Hosting a 750B MoE (Apache 2.0)
LG AI Research's K-EXAONE 2.0 750B-A37B needs ~1.5TB weights at BF16, ~750GB at FP8, and 375-420GB at 4-bit — hardware math and Apache 2.0 self-host limits.
K-EXAONELG AI ResearchOpen Weight LLM
AI2026-07-27
Kimi K3 Open Weights Are Out — What It Actually Takes to Self-Host a 2.8T MoE (MXFP4, ~1.4TB, vLLM/SGLang)
Moonshot released Kimi K3 open weights on July 26, 2026. At MXFP4 the weights alone are ~1.4TB — self-hosting means a multi-node cluster, not one GPU.
Kimi K3Moonshot AIOpen Weight LLM
AI2026-07-24
GGUF Quantization: Which Level to Pick (Q4_K_M, Q5_K_M, Q8_0, IQ) for Local LLMs
Start with Q4_K_M; step up to Q5_K_M or Q6_K if you have VRAM headroom. This guide explains GGUF naming, the quality/speed/VRAM tradeoffs per level, IQ (imatrix) quants, and how to choose by task. Updated July 2026.
GGUF量子化ローカルLLM
AI2026-07-21
LongCat-2.0 Requirements — VRAM, GPU, and API Pricing for a 1.6T Open MoE Model
LongCat-2.0 is Meituan's 1.6T-parameter, MIT-licensed MoE model. Running it locally needs roughly 3,800GB in BF16 or about 970GB even at INT4 — no personal PC can hold it. Requirements tables and API pricing ($0.75/$2.95 per 1M tokens), updated July 2026.
LongCat-2.0RequirementsVRAM
AI2026-07-21
DeepSeek V4 Requirements Reference — VRAM, RAM & GPU by Quantization, Plus API Pricing and the July 24 Legacy Retirement [Updated for the 0731 Build]
DeepSeek V4-Flash needs roughly 160GB at 4-bit and V4-Pro about 920GB. The July 31 V4-Flash-0731 build keeps the same footprint while beating V4-Pro on agent benchmarks. VRAM, RAM and GPU tables by quantization, API pricing, and the legacy model retirement. Updated August 2026.
DeepSeek V4RequirementsVRAM
AI2026-07-20
NVIDIA Nemotron 3 Requirements Reference — VRAM, GPU and RAM Quick-Lookup Tables for Nano, Super and Ultra (2026)
Nemotron 3 Nano runs in about 18GB at 4-bit, Super needs 8x H100-80GB, Ultra needs 4x B200 at NVFP4. VRAM, GPU and quantization tables for NVIDIA's open-weight MoE family. Updated July 2026.
Nemotron 3NVIDIARequirements
AI2026-07-18
GLM-5.2 Requirements Reference — VRAM, RAM & GPU Quick-Lookup Tables by Quantization [753B Open-Weight MoE, 2026]
GLM-5.2 needs roughly 430GB at 4-bit and about 1.5TB at BF16 in combined memory. Quick-lookup VRAM, RAM and quantization tables for running Z.ai's 753B / ~40B-active open-weight MoE (MIT-licensed) locally. Updated July 2026.
GLM-5.2Z.aiRequirements
AI2026-07-18
Inkling (Thinking Machines) Requirements Reference — VRAM, RAM & GPU Quick-Lookup Tables by Quantization [975B Open-Weight MoE, 2026]
Inkling needs from ~280GB (1-bit) to ~600GB (4-bit) combined memory, and 1.9TB at BF16. Quick-lookup VRAM, RAM, disk and quantization tables for running this 975B / 41B-active open-weight MoE locally. Updated July 2026.
InklingThinking MachinesRequirements
AI2026-07-18
Kimi K3 (Moonshot AI) Explained — 2.8T MoE Specs, API Pricing & Local Requirements (vs K2, Weights Due July 27) [2026]
Kimi K3 is a 2.8-trillion-parameter open-weight MoE (weights due July 27, 2026), ranked #1 on the frontend-code arena. API pricing is $3 input / $15 output per million tokens. Local runs are estimated at 650GB–1TB, needing server-class hardware. How it differs from K2.
Kimi K3Moonshot AIOpen Weight LLM
AI2026-06-15
Kimi K2.7-Code Deep Dive — Moonshot AI's June 12, 2026 Coding-Specialized 1T MoE Open-Weights Model, Modified MIT License, $0.95/$4.00 per 1M, 256K Context — But Japanese Enterprises Face Two Critical Caveats (Cross-Border Data and Unverified Benchmarks)
A primary-source deep dive on **Kimi K2.7-Code**, released June 12, 2026 by Moonshot AI (Beijing). Grounded in the [Hugging Face model card](https://huggingface.co/moonshotai/Kimi-K2.7-Code), [MarkTechPost](https://www.marktechpost.com/2026/06/12/moonshot-ai-releases-kimi-k2-7-code-a-coding-model-reporting-21-8-on-kimi-code-bench-v2-over-k2-6/), and [VentureBeat's skepticism piece](https://venturebeat.com/technology/kimi-k2-7-code-cuts-thinking-tokens-30-practitioners-say-benchmarks-dont-check-out). Covers the 1T-total / 32B-active MoE architecture (384 experts, 8 routed + 1 shared), 256K context, MoonViT ~400M vision encoder, native INT4, forced-on thinking mode. License is **Modified MIT** (attribution required only above 100M MAU or $20M MRR), API pricing is $0.95 input / $0.19 cache-hit / $4.00 output per 1M tokens — roughly **1/18 of Claude Opus 4.8's output price**. OpenAI + Anthropic-compatible endpoints drop straight into Claude Code / Cursor / Aider / Cline / cmux. Moonshot self-reports **+21.8% vs K2.6 on its own Kimi Code Bench v2 and -30% reasoning tokens**, but **all public benchmarks are Moonshot's own proprietary suites; independent SWE-bench Verified / Pro / FrontierCode scores are not yet available as of June 15, 2026** (VentureBeat). For Japanese enterprises the column flags two critical caveats: **(1) both `api.moonshot.cn` and the Singapore-subsidiary-run `api.moonshot.ai` remain exposed to PRC National Intelligence Law Article 7 compelled disclosure (set against Japan's PPC DeepSeek alert of February 3, 2025 and the Digital Agency notice of February 6, 2025), and (2) the only reliable mitigation is Hugging Face self-hosting (~4-8 H100, ~595GB INT4) following the Mizuho / Lion Qwen-on-domestic-infrastructure precedent**.
Moonshot AIKimi K2.7-CodeOpen Weight LLM