Skip to main content
株式会社オブライト
AI2026-10-0115 min read

Irodori-TTS-v4-Large: 3.29B Japanese TTS Guide & License

Irodori-TTS-v4-Large is a 3.29B Japanese-only TTS with voice cloning, Voice Design and emoji control. Setup, benchmarks vs v4.1-Small, Gemma license notes.


Irodori-TTS-v4-Large is a Japanese-only text-to-speech model of about 3.29B parameters, released on Hugging Face by Aratako. The repository was created on September 27, 2026. It unifies three conditions (input text, reference audio, and a caption) in one model, enabling zero-shot voice cloning, Voice Design (creating a voice from a text description), and voice cloning with style control. Emojis in the input text can also steer delivery and non-verbal sounds.

This article is based on the official model card, the quantized-variant model card, and the GitHub README as of October 1, 2026. Where the official sources are silent, such as required VRAM, we write 'not stated officially.' All benchmark figures are automatic evaluations by the developer; no large-scale human evaluation was conducted, so read the numbers with that in mind.

What Is Irodori-TTS-v4-Large? A Japanese TTS That Unifies Three Conditions in One Model

According to the model card, the model scales the v4 architecture from about 766M to about 3.29B parameters (roughly 4.3x). The method is a Rectified Flow Diffusion Transformer (RF-DiT), and the acknowledgments state that the design draws on Echo-TTS. Audio is represented with a codec using a 32-dimensional continuous latent and output at 48kHz. The basic specifications are as follows.

ItemDetails
DeveloperAratako (the model card's citation uses the name Chihiro Arata)
ReleaseRepository created 2026-09-27 (the quantized repo and demo Space were created the same day)
ParametersAbout 3.29B (24-layer, 2,048-dimension Diffusion Transformer)
MethodRectified Flow Diffusion Transformer (RF-DiT), Flow Matching family
LanguageJapanese only
Output48kHz. Represented with the Aratako/Semantic-DACVAE-Japanese-32dim codec (32-dimensional continuous latent)
Text encoderT5Gemma 2 (a fine-tuned text encoder from google/t5gemma-2-1b-1b), shared by the input text and the caption
Reference audio limit120 seconds in total
WatermarkInvisible watermark by Sony's SilentCipher, applied automatically (when the dependency and model files are available)
LicenseModel weights: Gemma Terms of Use. Code: MIT
Distribution sizeFull-precision model.safetensors is 13,152,893,348 bytes (about 13.15GB)

- Zero-shot voice cloning: pass a reference audio clip and the model reads the text in a voice close to that speaker
- Voice Design: create a voice from a text description (caption) alone, with no reference audio
- Voice cloning with style control: layer a caption on top of reference audio to specify emotion, delivery, and environment

In addition, emojis in the input text can control non-verbal sounds such as laughter, coughs, and sighs, as well as delivery (see below).

How It Works (Architecture): Five Components

As described in the model card and README, v4-Large consists of five components. The input text and the caption are read by the same encoder, speaker characteristics are extracted from the reference audio, and the Diffusion Transformer generates the audio latent conditioned on all of these.

How Irodori-TTS-v4-Large works: text with emojis and a caption go to a shared T5Gemma 2 encoder, reference audio (up to 120 seconds) to a reference latent encoder, a duration predictor estimates length, a 24-layer 2,048-dim Diffusion Transformer generates latents, and the DACVAE decoder outputs 48kHz speech with a SilentCipher watermark

- Shared text/caption encoder: a fine-tuned T5Gemma 2 that reads both the input text and the caption
- Condition projectors: trained separately for text and for captions
- Reference latent encoder: conditions on speaker identity. Reference audio is up to 120 seconds in total
- Diffusion Transformer: a joint-attention DiT using Low-Rank AdaLN, half-RoPE, and SwiGLU MLPs
- Duration Predictor: automatically estimates the length of the output audio, built from SwiGLU MLP blocks

v4-Small used ModernBERT-ja (sbintuitions/modernbert-ja-310m) as its text encoder. v4-Large replaces it with one derived from the T5Gemma 2 text encoder, and this swap is also the reason for the license difference discussed below. The Duration Predictor is trained separately after the main model, with the other parameters frozen. The model card says this is the same approach as v4.1-Small.

What Changed in v4-Large: Differences from v4.1-Small

The model card's 'What's New' lists four changes: scaling up, the T5Gemma 2 encoder, a separately trained Duration Predictor, and a higher Voice Design score. The table below sets confirmed values against v4.1-Small, the smaller model in the same Irodori-TTS series (a minor update of v4-Small).

Itemv4-Largev4.1-Small
ParametersAbout 3.29BAbout 766M (the value in the v4-Small model card)
Text encoderT5Gemma 2 (fine-tuned)ModernBERT-ja (fine-tuned)
Weights licenseGemma Terms of UseMIT
Full-precision fileAbout 13.15GBAbout 3.06GB
INT8 (weight-only)3,662 MiB872 MiB
INT4 (weight-only)2,818 MiB813 MiB
Reading accuracy (Joyo Parakeet edition)92.80%93.42%
Few-step distilled versionNot stated officiallyYes (v4.1-Small-MF, 4 steps)
Reference audio limit120 seconds in total120 seconds in total

Note that the Voice Design and voice cloning benchmarks compare v4-Large with v4-Small (not v4.1). The series history, by Hugging Face repository creation date, runs: 500M (2026-02-25), 500M-v2 (03-23), 500M-v2-VoiceDesign (03-30), 500M-v3 (05-12), 600M-v3-VoiceDesign (05-31), v4-Small (08-01), v4.1-Small (08-11), v4.1-Small-MF (09-12), and v4-Large (09-27).

Benchmarks: Small Reads Better, Large Is More Expressive

According to the model card, the evaluation conditions were FP32 inference, 40 RF steps, and text CFG 3.0, reported as the mean and population standard deviation over 5 seeds (0 to 4). Joyo, JSUT, Coco-Nut, and JVS are not in the training data. Everything is an automatic evaluation by the developer, and the card states that no large-scale human MOS evaluation was run. First, Japanese reading (no reference audio, no caption).

Reference length vs speaker similarity (JVS, CAM++ cosine): v4-Large scores 0.6711 with one clip, 0.7593 at ~30 s, 0.7700 at ~60 s and 0.7788 at 120 s; v4-Small scores 0.6610, 0.7521, 0.7646 and 0.7753. Most of the gain arrives by about 30 seconds
BenchmarkMetricv4.1-Smallv4-Large
Joyo kanji reading Parakeet Edition (13,536 sentences, 4,512 kanji-reading pairs)Reading accuracy↑93.42 ± 0.04%92.80 ± 0.08%
SameTarget Kana-CER↓6.88 ± 0.08%7.96 ± 0.36%
SameSentence Kana-CER↓1.25 ± 0.01%1.51 ± 0.06%
SameText CER↓4.68 ± 0.06%4.87 ± 0.08%
JSUT BASIC5000Sentence Kana-CER↓3.43 ± 0.01%3.67 ± 0.03%
JSUT BASIC5000Standard CER↓7.22 ± 0.12%7.30 ± 0.01%

The model card states that reading error rates are slightly higher for v4-Large than for v4.1-Small on both benchmarks. On reading accuracy alone, Small is ahead. Next is Voice Design (caption following). The 2,890 public voice descriptions in the Coco-Nut test set were combined with short, medium, and long texts, giving 8,670 clips per model (no reference audio), which Gemini 3.6 Flash scored from 1 to 5. This is a single run with seed 0.

ModelMean score↑Score 1–2↓Score 4–5↑
600M-v3-VoiceDesign4.209617.09%73.30%
v4-Small4.233916.78%74.12%
v4-Large4.299114.63%76.24%

The gain over v4-Small is +0.0652 (95% confidence interval [+0.0353, +0.0955]). The largest improvement was on long texts, from 3.9851 to 4.1343. The model card itself notes that this is a lightweight internal comparison using a single automatic judge, with no human verification, and not a research-grade benchmark. Voice cloning was evaluated on all 100 JVS speakers, 5 texts, 5 seeds, speaker CFG 5.0, broken down by reference audio length.

Reference audiov4-Small CAM++ cosine↑v4-Small top-1↑v4-Large CAM++ cosine↑v4-Large top-1↑
1 clip0.661084.60%0.671187.00%
About 30 s0.752198.56%0.759399.48%
About 60 s0.764699.56%0.770099.60%
120 s0.775399.76%0.778899.64%

How to read this: cosine similarity is higher for v4-Large in all four conditions, but the top-1 at 120 seconds is slightly lower for v4-Large, and the gap is small. Most of the improvement is already obtained at about 30 seconds of reference audio, and similarity drops substantially with a single short clip. Where possible, use clean reference audio of about 30 seconds or more. To repeat, these are automatic evaluations with no human evaluation, and the cloning comparison is against v4-Small.

Hardware and Quantized Variants: No Official VRAM Requirement

Neither the model card nor the quantized-variant model card states a required VRAM; this is not stated officially. All we have are checkpoint file sizes: the full-precision model.safetensors is about 13.15GB. Five post-training quantized variants made with torchao are published, listed below. File size is only a rough guide, since runtime usage adds reference audio, CFG branches, and decoding, so required VRAM cannot be derived from file size.

VariantQuantizationSizeGPU supportNotes
int8-weight-onlyW8A163,662 MiBNVIDIA CUDAGeneral-purpose recommended INT8
int8-dynamicW8A83,662 MiBNVIDIA CUDADynamic INT8 activations as well
int4-weight-onlyW4A16 (group size 128)2,818 MiBNVIDIA Ampere or later (compute capability 8.0+)Smallest
float8-weight-onlyFP8 weights, BF16 activations3,665 MiBNVIDIA Ada / Hopper / Blackwell (8.9+)
float8-dynamicFP8 weights, dynamic FP8 activations3,659 MiBSame (8.9+)

Execution of the quantized variants was verified on NVIDIA CUDA only. CPU, ROCm, and Intel XPU depend on PyTorch / torchao kernels and are untested in this release. Specify --model-precision bf16 at inference. To simply try the model first, the demo Space, which runs on ZeroGPU, is the easiest way to start in a browser.

How to Use It: Install, Voice Cloning, Voice Design

The code is published on GitHub, and the README uses uv. Clone the repository, then sync with the extra that matches your backend.

```bash
git clone https://github.com/Aratako/Irodori-TTS.git
cd Irodori-TTS
uv sync --extra cu128
```

The backend extras are mutually exclusive: cu128 (NVIDIA CUDA 12.8, Linux / Windows), rocm (AMD, Linux / WSL), xpu (Intel, Linux / Windows), and cpu (CPU only, or CPU / MPS on macOS). After syncing, run commands with uv run --no-sync. Below is the int8 inference command from the quantized-variant model card (no reference audio).

```bash
uv run --no-sync python infer.py \
  --hf-checkpoint Aratako/Irodori-TTS-v4-Large-Quantized/int8-weight-only \
  --model-precision bf16 \
  --text "こんにちは、私はAIです。これは音声合成のテストです。" \
  --no-ref \
  --output-wav outputs/sample.wav
```

To use the full-precision model, replace --hf-checkpoint with Aratako/Irodori-TTS-v4-Large. The README examples are written with the v4.1-Small repository ID, so you need to substitute the ID when using v4-Large. The main options listed in the README are below (examples are written for v4.1-Small).

OptionPurpose
--ref-wav path/to/reference.wavVoice cloning from reference audio
--ref-wavs ref_01.wav ref_02.wav ref_03.wavMultiple clips, concatenated in order (truncated at the 120-second limit)
--caption "落ち着いた、近い距離感の女性話者" with --no-refVoice Design from a caption alone
--ref-wav with --captionVoice cloning with style control
--secondsIf omitted, the Duration Predictor estimates the length automatically. Adjust with --duration-scale
Defaults--num-steps 40 (RF), --cfg-scale-text 3.0, --cfg-scale-caption 3.0, --cfg-scale-speaker 5.0

For browser-based operation there are Gradio UIs: gradio_app.py for voice cloning (port 7860) and gradio_app_voicedesign.py for Voice Design (7861). As of the README, both UIs default to v4.1-Small. As of October 1, 2026, the latest GitHub commit (2026-09-12) is titled 'Add MeanFlow distillation and v4-Large support,' and the configs include train_v4_large.yaml and train_v4_large_duration.yaml. On the other hand, v4-Large support for the OpenAI-compatible server (Irodori-TTS-Server), LoRA fine-tuning, and the MeanFlow distilled version is not stated officially; the README targets v4-Small / v4.1-Small for these.

Controlling Delivery with Emojis

Placing emojis in the input text controls delivery, emotion, and non-verbal sounds such as laughter, coughs, and sighs. EMOJI_ANNOTATIONS.md lists all 45; below is an excerpt of 12.

EmojiMeaning
👂Whisper, sound near the ear
😮‍💨Breath, sigh
⏸️Pause, silence
🤭Laughter (giggle, suppressed laugh)
📢Echo, reverb
📞Sound as if over a phone or speaker
⏩Fast speech
🐢Slowly
😊Cheerful, happy
😠Anger, displeasure
😭Sobbing, crying, sadness
📖Narration, monologue

The README's example sentence is 「あははっ🤭、それ本当に言ってるの?…😮‍💨まぁ、君らしいけどね。」 with the caption 「余裕のある大人の男性。親しい相手に対して、くだけた雰囲気で呆れながらも楽しそうに話している。」 (a composed adult man talking casually to someone close, amused and a little exasperated). Repeating the same emoji strengthens the effect. The official docs state, however, that control is not perfect, and that the effect depends on context and can be inconsistent.

License and Ethical Restrictions: Small Is MIT, Large Is Under the Gemma Terms

The biggest caveat for v4-Large is the license. According to the model card, because the shared encoder derives from google/t5gemma-2-1b-1b, the v4-Large weights fall under the Gemma Terms of Use, and use and redistribution must follow those terms and the Gemma Prohibited Use Policy. The repository bundles GEMMA_TERMS_OF_USE.md, GEMMA_PROHIBITED_USE_POLICY.md, and NOTICE. Even though the code is MIT, the weights are under a separate agreement, so the model cannot simply be called 'open source.' For the Gemma family itself, see our Gemma 4 introduction.

TargetLicense
v4-Large model weightsGemma Terms of Use
v4-Small / v4.1-Small / v4.1-Small-MF model weightsMIT
Code (GitHub)MIT

- Google claims no rights in Outputs
- On redistribution, you must include usage restrictions in your agreement, provide a copy of the terms, mark modifications, and include the NOTICE
- Google reserves the right to restrict, by remote or other means, use it reasonably believes violates the terms
- There is no clause banning commercial use, but it is conditional as above (the terms were last updated 2026-04-01)

- Do not clone or impersonate the voices of voice actors, celebrities, public figures, or others without their explicit consent
- Do not generate content for deepfakes or disinformation
- Even without reference audio, output may happen to resemble a real person (a probabilistic result; training was not intended to reproduce any specific individual)
- The developer takes no responsibility for misuse, and legal compliance is the user's responsibility

None of this is legal advice. If you plan commercial use or redistribution, read the original Gemma Terms of Use and Prohibited Use Policy before making a final decision. Voice cloning presupposes the consent of the person whose voice is used as the reference.

Which One to Choose

Use case / prioritySuited modelReason
Reading accuracy, light footprint, MIT licensev4.1-SmallHigher reading accuracy (93.42%), 872 MiB at INT8, MIT weights
Generation speedv4.1-Small-MFA 4-step distilled version exists (a v4-Large distilled version is not stated officially)
Voice Design following, voice reproductionv4-LargeMean Voice Design score +0.0652 over v4-Small, and cloning similarity higher in all 4 conditions

v4-Large trails Small on reading accuracy, so bigger is not always better. It is the model to pick when voice specification and reproduction matter. In positioning, it is close to local voice tools such as VoiceStudio in that it can run locally. For a cloud TTS API example, there is the xAI Grok voice API; check each provider's official documentation for pricing, language support, and latency. For the reverse direction, fully local speech input (STT), see TypeWhisper.

Cautions for Business Use

- Have a person listen for misreadings. The official docs state that rare proper nouns, technical terms, and context-dependent readings can be misread
- Get the speaker's consent and use clean reference audio of about 30 seconds or more. With a single short clip, speaker similarity drops substantially
- Avoid contradictions between the caption and the reference audio. Contradictions can cause instability or unnatural output, or one side may take priority. Use the caption for emotion, style, and environment, and match voice quality to the reference audio
- Do not strip the watermark. Disclose that the audio is synthetic when you publish it
- Check the license. The v4-Large weights are under the Gemma Terms of Use, which carry obligations on redistribution
- Japanese only. It cannot be used for narration in other languages

For business uses such as phone answering, misreadings and disclosure of synthetic audio are especially important. For a cost overview, see AI phone answering cost guide.

FAQ

Can I use it commercially?

The v4-Large weights fall under the Gemma Terms of Use. There is no clause banning commercial use, but there are conditions such as obligations on redistribution. Check the original terms before deciding; this is not legal advice. The code is MIT, and the weights of the Small models such as v4.1-Small are also MIT.

How much VRAM does it need?

Not stated officially. The only reference is file size: about 13.15GB at full precision and 2,818 to 3,665 MiB for the quantized variants. Runtime usage adds reference audio, CFG branches, and decoding, so required VRAM cannot be derived from file size. Try the demo Space first.

Does it run on a Mac?

The README says the cpu extra targets CPU only, or CPU / MPS on macOS, so an installation path exists. However, execution of the quantized variants was verified only on NVIDIA CUDA, and CPU and others are untested. Speed on a Mac is not stated officially.

Which is better, v4-Large or v4.1-Small?

Choose v4.1-Small for reading accuracy, light footprint, and the MIT license; choose v4-Large for Voice Design following and voice reproduction. Reading accuracy is 93.42% for v4.1-Small and 92.80% for v4-Large.

How long should the reference audio be for voice cloning?

Up to 120 seconds in total. In the official evaluation, most of the improvement comes at about 30 seconds, and similarity drops substantially with a single short clip. Use clean audio of about 30 seconds or more where possible.

Can it speak English?

No. The official docs state that only Japanese text is supported.

What is the watermark?

It is an inaudible watermark from Sony's SilentCipher, applied automatically to generated audio (when the dependency and model files are available). For how detection works, see the official SilentCipher repository (sony/silentcipher).

Sources

Feel free to contact us

Contact Us