Skip to main content
株式会社オブライト
Services
About
Company
Column
Glossary
Pricing
Free Tools
Contact
日本語
日本語
メニューを開く
Column
音声AI
Articles tagged "音声AI"
6 articles
AI
2026-10-01
Irodori-TTS-v4-Large Explained: A 3.29B Japanese TTS — Voice Cloning, Voice Design, Emoji Control, Differences from v4.1-Small, and License [Released September 2026]
Irodori-TTS-v4-Large is a 3.29B Japanese-only TTS with voice cloning, Voice Design and emoji control. Setup, benchmarks vs v4.1-Small, Gemma license notes.
音声AI
ローカルAI
Hugging Face
AI
2026-09-14
What Is VoiceStudio? The Fully-Local, Open-Source ElevenLabs Alternative
A guide to debpalash/VoiceStudio (formerly OmniVoice-Studio), which surged on GitHub Trending in September 2026. Covers how it runs voice cloning, voice design, video dubbing, dictation, transcription, and audiobook creation entirely on local hardware, its 16 TTS / 11 ASR engine lineup, installation, local API/MCP support, and how it compares to ElevenLabs and other tools.
音声AI
ローカルAI
オープンソース
AI
2026-09-13
GPT-Live-1 Guide: $0.05/min Full-Duplex Voice API Explained
GPT-Live-1 is OpenAI's full-duplex voice model, released Sep 10, 2026, billed $0.05/min on a dedicated Live endpoint, with 87% tool accuracy and 0.8s latency.
OpenAI
Realtime API
音声AI
AI
2026-05-08
OpenAI GPT-Realtime-2 and the Three New Voice Models — A Practitioner's 2026 Look at Reasoning Voice Agents, Live Translation, and Streaming Whisper
On May 7, 2026, OpenAI released a trio of new voice models: GPT-Realtime-2 (the first voice model with GPT-5-class reasoning), GPT-Realtime-Translate (live translation across 70+ input / 13 output languages), and GPT-Realtime-Whisper (streaming speech-to-text). This article summarizes capabilities, benchmark deltas vs 1.5, pricing, when to pick which, and the upgrade decision from 1.5 — based on official information.
OpenAI
gpt-realtime-2
Realtime API
AI
2026-04-28
OpenAI gpt-realtime-1.5 and the Official realtime-voice-component — A Practitioner's Look at the New Voice-Agent Stack [2026]
OpenAI released the gpt-realtime-1.5 audio model on February 26, 2026, and openai/realtime-voice-component on GitHub provides an official React reference for voice UIs. This article summarizes the documented gains (+5% audio reasoning, +10.23% transcription, +7% instruction following), pricing, the component's positioning as a reference implementation, and practical considerations for business adoption.
OpenAI
gpt-realtime
Realtime API
AI
2026-04-06
NVIDIA PersonaPlex 7B Complete Guide — Real-Time Full-Duplex Voice AI Architecture & Use Cases [2026]
NVIDIA PersonaPlex 7B, released in January 2026, is an open-source voice AI that integrates the traditional ASR→LLM→TTS pipeline into a single end-to-end model, achieving true full-duplex voice interaction. This guide covers architecture, performance benchmarks, setup procedures, and practical use cases.
PersonaPlex
NVIDIA
音声AI