Skip to main content
株式会社オブライト
Services
About
Company
Column
Glossary
Pricing
Free Tools
Contact
日本語
日本語
メニューを開く
Column
ローカルLLM
Articles tagged "ローカルLLM"
30 articles
AI
2026-08-17
Needle 2 Explained: Cactus Compute's 14MB On-Device Agent LLM
Needle 2 is Cactus Compute's 45M-parameter on-device agent LLM, shipped as a single 14MB binary for tool calling, device use, and structured extraction. It runs from Raspberry Pi 5 down to budget phones. Here's how it works and how to install it (Updated August 2026).
Needle 2
Cactus Compute
On-device AI
AI
2026-08-15
Qwen3.8-27B System Requirements — VRAM 9–56GB, Apache-2.0 Licensed [2026]
Qwen3.8-27B weights landed on August 15, 2026. This at-a-glance requirements guide maps VRAM needs to the actual published file sizes: Q4_K_M is 17.1GB and won't fit a 16GB GPU, making IQ4_XS (15.7GB) the practical floor. Licensed Apache-2.0 for commercial use.
Qwen 3.8
Requirements
VRAM
AI
2026-08-13
Local LLM Inference Engines Compared — llama.cpp, Ollama, vLLM, LM Studio, MLX, TensorRT-LLM
Comparing local LLM engines: Ollama and llama.cpp solo, MLX on Apple Silicon, vLLM for concurrent serving, TensorRT-LLM for NVIDIA, LM Studio for GUI trials.
ローカルLLM
ローカルAI
Ollama
AI
2026-08-12
Nemotron 3.5 Lightning & NeMo Switchyard: A New Agent Stack
NVIDIA unveiled Nemotron 3.5 Lightning, a 30B-A3B MoE model delivering 4x faster output, plus NeMo Switchyard, cutting agent costs to about one-third.
NVIDIA
Nemotron
ローカルLLM
AI
2026-08-11
Muse Glimmer 30B Hardware Requirements — VRAM, GPU & Mac Guide for Meta's Local Agentic Model (2026)
Muse Glimmer 30B fits in ~20GB at 4-bit quantization on a 24GB GPU or Mac. VRAM needs by quant level, plus RTX 5090 and M4 Max speeds. Updated August 2026.
ローカルLLM
オープンソースLLM
AIエージェント
AI
2026-08-07
LFM2.5-2.6B Guide: Requirements, VRAM, and Benchmarks (2026)
Liquid AI's LFM2.5-2.6B is a 2.69B on-device agent model under 2.5GB with 128K context. This guide covers VRAM sizing, benchmarks, throughput, and licensing.
Liquid AI
LFM2.5
Requirements
AI
2026-08-06
Shieldstral 1.0 3B Explained: Mistral's Open-Weight Multimodal Moderation Model
Shieldstral 1.0 3B is Mistral AI's Apache 2.0 moderation model, released Aug 4, 2026. It needs ~16GB VRAM in BF16 and screens both text and images together.
Mistral
Requirements
VRAM
AI
2026-08-03
MiniMax H3 Requirements: VRAM, GPU Sizing & File Sizes (2026 Open-Weight Video+Audio Model)
MiniMax H3, an open video+audio model, released weights Aug 3, 2026. ComfyUI needs ~42.5GB files, ~24GB VRAM (12GB may work). Covers quantization, GPU sizing.
MiniMax
Requirements
VRAM
AI
2026-08-02
WASTE: Run Kimi K3's 2.78T Params on 29GB RAM
WASTE is a dependency-free C engine running Kimi K3 (2.78T params) on 29GB RAM by streaming MoE experts from NVMe. Covers real throughput, setup, and limits.
Kimi K3
Moonshot AI
MoE
AI
2026-08-01
K-EXAONE 2.0 750B-A37B: Self-Hosting a 750B MoE (Apache 2.0)
LG AI Research's K-EXAONE 2.0 750B-A37B needs ~1.5TB weights at BF16, ~750GB at FP8, and 375-420GB at 4-bit — hardware math and Apache 2.0 self-host limits.
K-EXAONE
LG AI Research
Open Weight LLM
AI
2026-07-30
Qwen Scribe Explained: Fully Local Transcription and System-Wide Dictation on Apple Silicon
Qwen Scribe runs Qwen3-ASR on Apple Silicon via MLX for fully offline transcription and system-wide dictation — 3.4GB unified memory for the 1.7B model, 1.2GB for 0.6B. Covers requirements, setup, push-to-talk usage, SRT export, the three macOS permissions, and how it compares to Whisper-based tools.
ローカルLLM
音声入力
音声認識
AI
2026-07-27
Kimi K3 Open Weights Are Out — What It Actually Takes to Self-Host a 2.8T MoE (MXFP4, ~1.4TB, vLLM/SGLang)
Moonshot released Kimi K3 open weights on July 26, 2026. At MXFP4 the weights alone are ~1.4TB — self-hosting means a multi-node cluster, not one GPU.
Kimi K3
Moonshot AI
Open Weight LLM
AI
2026-07-24
GGUF Quantization: Which Level to Pick (Q4_K_M, Q5_K_M, Q8_0, IQ) for Local LLMs
Start with Q4_K_M; step up to Q5_K_M or Q6_K if you have VRAM headroom. This guide explains GGUF naming, the quality/speed/VRAM tradeoffs per level, IQ (imatrix) quants, and how to choose by task. Updated July 2026.
GGUF
量子化
ローカルLLM
AI
2026-07-23
Ollama vs LM Studio — How to Choose a Local LLM Runner (2026)
A 2026 comparison of Ollama and LM Studio, the two leading local LLM runners: CLI vs GUI, install and usage, API, OS support, and model formats, with guidance on which to pick.
Ollama
LM Studio
ローカルLLM
AI
2026-07-21
LongCat-2.0 Requirements — VRAM, GPU, and API Pricing for a 1.6T Open MoE Model
LongCat-2.0 is Meituan's 1.6T-parameter, MIT-licensed MoE model. Running it locally needs roughly 3,800GB in BF16 or about 970GB even at INT4 — no personal PC can hold it. Requirements tables and API pricing ($0.75/$2.95 per 1M tokens), updated July 2026.
LongCat-2.0
Requirements
VRAM
AI
2026-07-21
DeepSeek V4 Requirements Reference — VRAM, RAM & GPU by Quantization, Plus API Pricing and the July 24 Legacy Retirement [Updated for the 0731 Build]
DeepSeek V4-Flash needs roughly 160GB at 4-bit and V4-Pro about 920GB. The July 31 V4-Flash-0731 build keeps the same footprint while beating V4-Pro on agent benchmarks. VRAM, RAM and GPU tables by quantization, API pricing, and the legacy model retirement. Updated August 2026.
DeepSeek V4
Requirements
VRAM
AI
2026-07-20
NVIDIA Nemotron 3 Requirements Reference — VRAM, GPU and RAM Quick-Lookup Tables for Nano, Super and Ultra (2026)
Nemotron 3 Nano runs in about 18GB at 4-bit, Super needs 8x H100-80GB, Ultra needs 4x B200 at NVFP4. VRAM, GPU and quantization tables for NVIDIA's open-weight MoE family. Updated July 2026.
Nemotron 3
NVIDIA
Requirements
AI
2026-07-18
GLM-5.2 Requirements Reference — VRAM, RAM & GPU Quick-Lookup Tables by Quantization [753B Open-Weight MoE, 2026]
GLM-5.2 needs roughly 430GB at 4-bit and about 1.5TB at BF16 in combined memory. Quick-lookup VRAM, RAM and quantization tables for running Z.ai's 753B / ~40B-active open-weight MoE (MIT-licensed) locally. Updated July 2026.
GLM-5.2
Z.ai
Requirements
AI
2026-07-18
Inkling (Thinking Machines) Requirements Reference — VRAM, RAM & GPU Quick-Lookup Tables by Quantization [975B Open-Weight MoE, 2026]
Inkling needs from ~280GB (1-bit) to ~600GB (4-bit) combined memory, and 1.9TB at BF16. Quick-lookup VRAM, RAM, disk and quantization tables for running this 975B / 41B-active open-weight MoE locally. Updated July 2026.
Inkling
Thinking Machines
Requirements
AI
2026-07-18
Kimi K3 (Moonshot AI) Explained — 2.8T MoE Specs, API Pricing & Local Requirements (vs K2, Weights Due July 27) [2026]
Kimi K3 is a 2.8-trillion-parameter open-weight MoE (weights due July 27, 2026), ranked #1 on the frontend-code arena. API pricing is $3 input / $15 output per million tokens. Local runs are estimated at 650GB–1TB, needing server-class hardware. How it differs from K2.
Kimi K3
Moonshot AI
Open Weight LLM
AI
2026-05-05
NVIDIA DGX Spark in 2026 — A Two-Stage Workflow for Code Migrations Where "Confidential Analysis Stays Local, Cloud LLMs Only Touch Sanitized Code"
An overview of NVIDIA DGX Spark (GB10 Grace Blackwell Superchip, 128GB unified memory, up to 1 PFLOP at FP4, $4,699) and a concrete two-stage workflow for confidential code-migration projects: analyze and sanitize locally, then hand a clean, PII-free representation to cloud frontier LLMs for the actual migration. Practical answers to the "executives won't approve cloud AI even with opt-out" problem.
NVIDIA
DGX Spark
ローカルLLM
AI
2026-04-17
NousResearch Hermes Complete Guide — Hermes 4.3 36B, Function Calling & Hermes Agent [2026]
Complete guide to NousResearch Hermes 4.3 36B (512K context) and the Hermes Agent framework. Covers Function Calling implementation, Ollama setup, hardware requirements, and benchmarks including RefusalBench dominance — updated for 2026.
Hermes
NousResearch
Function Calling
AI
2026-04-10
Local LLM Landscape April 2026 — Top 10 Open-Source Models Comprehensive Comparison [Ollama Guide]
Comprehensive comparison of the top 10 local LLMs as of April 2026. Covers SWE-bench scores, Japanese language performance, VRAM requirements, Ollama commands, and licensing for Gemma 4, Llama 4, Qwen 3.5, GLM-5.1, Kimi K2.5, MiniMax M2.5, and more.
ローカルLLM
オープンソースAI
2026年
AI
2026-04-10
Qwen 3.5 27B Dense & 35B-A3B MoE Complete Guide — DFlash Acceleration Breaks 24GB GPU Limits [2026]
Compare Qwen 3.5 27B Dense vs 35B-A3B MoE, check 24GB GPU requirements, learn DFlash 2–3x acceleration, and follow step-by-step Ollama setup instructions.
Qwen 3.5
27B Dense
35B-A3B MoE
AI
2026-04-07
Gemma 4 E4B Complete Guide — 4.5B Parameter Multimodal Model for Edge Deployment [2026]
Gemma 4 E4B is Google's 4.5B parameter edge AI model released in April 2026. This guide covers local deployment on Apple Silicon and Raspberry Pi, multimodal features, quantization settings, and benchmark comparisons.
Gemma 4
Gemma 4 E4B
エッジAI
AI
2026-04-04
Claude Alternative Local LLM Comparison 2026 — Qwen 3.5, Mistral Small 4, DeepSeek R1 & Gemma 4 Reviewed
Following Anthropic Claude restrictions, comprehensive comparison of local LLMs including Qwen 3.5-9B, Mistral Small 4, DeepSeek R1, Gemma 4, and Llama 4. Detailed analysis of Japanese performance, hardware requirements, and use-case recommendations.
ローカルLLM
Qwen 3.5
Mistral Small 4
AI
2026-04-04
AI API Cost Optimization in the Pay-Per-Use Era — Smart Strategies for Claude, GPT, Gemini & Local LLMs [2026]
Comprehensive guide to AI API cost optimization in the pay-per-use era. Covers Claude, GPT, Gemini pricing comparisons, 5 reduction techniques including prompt caching, batch APIs, local LLM hybrid operations, monthly cost simulations, and ROI calculation methods.
AI API
コスト最適化
従量課金
AI
2026-04-04
Hybrid AI Strategy Guide — Achieving 50% Cost Reduction with Cloud API + Local LLM [2026]
A practical guide to reducing AI operational costs by over 50% with a hybrid AI strategy combining cloud APIs and local LLMs. Learn optimal architecture design and implementation steps using local models like Qwen 3.5 and DeepSeek R1 with Claude, GPT, and Gemini.
ハイブリッドAI
ローカルLLM
コスト削減
AI
2026-04-03
Gemma 4 Complete Guide — Features, System Requirements & Ollama Setup [2026]
Complete guide to Google Gemma 4 (released April 2, 2026): 4 model variants (E2B/E4B/26B MoE/31B Dense), Apache 2.0 license, system requirements, multimodal capabilities, AIME 89% benchmark, 140+ languages, and step-by-step Ollama installation and setup instructions.
Gemma 4
Ollama
Google
AI
2026-04-03
Gemma 4 vs Llama 4 vs Qwen 3.5 Comparison — 2026 Local LLM Selection Guide
Comprehensive comparison of Gemma 4, Llama 4, and Qwen 3.5 local LLMs. Detailed analysis of benchmark performance, licensing, Japanese support, hardware requirements, and use case selection criteria.
Gemma 4
Llama 4
Qwen 3.5