VoxCPM: tokenizer-free multilingual speech synthesis
OpenBMB/VoxCPM · 37,974★ · 4,314 forks
Everything worth knowing about OpenBMB/VoxCPM: an open-source, open-weight text-to-speech system that generates continuous audio directly, with no discrete tokenization, whose version 2 supports 30 languages, description-based voice design, and 48 kHz cloning.
What VoxCPM is
VoxCPM is a tokenizer-free text-to-speech (TTS) system: instead of converting audio into a vocabulary of discrete tokens and then decoding them, it generates continuous audio representations directly through an end-to-end autoregressive diffusion architecture. The product is an open-weight model (Apache-2.0) with a Python API, CLI, and web demo.
The current version is VoxCPM2, a 2B-parameter model trained on over 2 million hours of multilingual speech, supporting 30 languages and 9 Chinese dialects, Voice Design (creating a new voice from a natural-language description with no reference audio), controllable cloning (cloning a voice from a short clip and adjusting style, pace, and emotion), and ultimate-fidelity cloning via audio continuation, with native 48 kHz output. It’s built on the same organization’s MiniCPM-4 backbone.
Origin
The repository was created on September 16, 2025 by the OpenBMB organization, the MiniCPM model team at ModelBest (a Chinese company) in collaboration with THUHCSI (a Tsinghua University research group).
The timeline the README documents: September 2025, publication of the VoxCPM technical report and the launch of VoxCPM-0.5B, which reached #1 trending on HuggingFace (per the README); December 5, 2025, publication of VoxCPM1.5 weights with SFT and LoRA fine-tuning, #1 trending on GitHub (per the README); April 6, 2026, the launch of VoxCPM2 — 2B, 30 languages, Voice Design and controllable cloning, 48 kHz output — with a technical report on arXiv.
The line of work extends the MiniCPM family (10,222 stars), sharing its philosophy of small, capable models that run locally. The README explicitly credits DiTAR for the autoregressive diffusion backbone, CosyVoice (FunAudioLLM) for the LocDiT Flow Matching implementation, and DAC (descript-audio-codec) for the audio VAE backbone.

Philosophy and principles
The technical report and README reveal four coherent design decisions:
- No discrete tokenization: the model operates entirely in the AudioVAE V2 latent space and “paints” the audio with a diffusion step (LocDiT) at each position, instead of choosing among quantized tokens. The argument is that a discrete quantizer introduces a quality bottleneck; without it, the VAE can reconstruct directly at 48 kHz.
- One model, several modes: the technical report describes a “unified sequence organization” where every generation mode (text-to-speech, voice design, reference cloning, continuation cloning) is expressed as different arrangements of the same input blocks.
- Studio quality without an external upsampler: AudioVAE V2 has an asymmetric design — encoding at 16 kHz and reconstructing at 48 kHz — building implicit super-resolution into the decoder itself.
- Full openness and commercial use: weights and code under Apache-2.0, “free for commercial use” per the README. The README includes a “Risks and Limitations” section that explicitly prohibits use for impersonation, fraud, or disinformation and recommends labeling generated content.


How it works
VoxCPM2’s architecture follows four stages operating in the AudioVAE V2 latent space: LocEnc (local encoder, fusing text and optional reference audio); TSLM (the autoregressive language model, a 2B-parameter MiniCPM-4 backbone, deciding which audio “patch” comes next, at 6.25 Hz); RALM (residual autoregressive language model, a second pass refining each patch with prosodic detail); LocDiT (a diffusion estimator that paints the audio latent, decoded by AudioVAE V2 to 48 kHz).
The four documented generation modes: Text-to-Speech (model.generate(text=...)); Voice Design (the description goes in parentheses at the start of the text, no reference audio required); Controllable Voice Cloning (pass reference_wav_path, a 5–30s clip, with optional style instructions); Ultimate Cloning (pass prompt_wav_path and prompt_text, the clip’s exact transcript).

There’s also a streaming API (model.generate_streaming) and a CLI (voxcpm design, voxcpm clone, voxcpm batch) with word- or character-level timestamp support. The README reports RTF ~0.30 on an NVIDIA RTX 4090 with standard PyTorch, ~0.13 with Nano-vLLM, and ~1.76 (Q8_0) on Apple M4 Pro / Metal via llama.cpp-omni. VRAM: ~8 GB for VoxCPM2.

The ecosystem
The README maintains an official ecosystem table (community projects, not maintained by OpenBMB). Inference and deployment engines: vllm-project/vllm-omni (the official vLLM project’s omni-modal extension with native VoxCPM2 support, an OpenAI-compatible endpoint; 6,263 stars — the most relevant semi-official deployment, published by the vLLM organization, not OpenBMB); a710128/nanovllm-voxcpm (292 stars, RTF ~0.13 on a 4090); tc-mb/llama.cpp-omni (253 stars, native GGUF support on CPU/Metal/CUDA/Vulkan); bluryar/VoxCPM.cpp (89 stars); 0xShug0/audio.cpp (1,957 stars, a unified C++ framework for several audio models).

Community interfaces and apps: ComfyUI nodes (wildminder/ComfyUI-VoxCPM, 509 stars; Saganaki22/ComfyUI-VoxCPM2, 195); liuzhao1225/YouDub-webui (5,345 stars, a video-localization tool with voice-clone dubbing — the largest adoption case as an engine inside another application); maomao-2001/Whispera (100 stars, real-time conversation with streaming TTS).

Official / semi-official status
- vLLM-Omni publishes native VoxCPM2 support with deployment examples in its own repository: this is adoption by the official vLLM project, which in practice means VoxCPM2 can be served with the same infrastructure as the highest-traffic language models, with an OpenAI-compatible API.
- ICLR 2026: the README links VoxCPM-0.5B’s review on OpenReview under “ICLR 2026.” Acceptance couldn’t be verified directly in this run; it’s cited as the status the project itself declares.
- The README states VoxCPM-0.5B was #1 trending on HuggingFace and VoxCPM1.5 was #1 trending on GitHub. These are the README’s own claims.
- There’s no official marketplace in the agent sense: the de facto status is more that of a reference for local multilingual TTS, reinforced by presence on PyPI, HuggingFace, ModelScope, and the engine ecosystem.

Quick-start guide
Installation and first run
Documented prerequisites: Python ≥ 3.10 and < 3.13, PyTorch ≥ 2.5.0, CUDA ≥ 12.0 for an NVIDIA GPU. VoxCPM2 uses ~8 GB of VRAM.
pip install voxcpm
First run (Python API):
from voxcpm import VoxCPM
import soundfile as sf
model = VoxCPM.from_pretrained(
"openbmb/VoxCPM2",
load_denoiser=False,
)
wav = model.generate(
text="VoxCPM2 is the current recommended release for realistic multilingual speech synthesis.",
cfg_value=2.0,
inference_timesteps=10,
seed=42,
)
sf.write("demo.wav", wav, model.tts_model.sample_rate)
The first call downloads the weights from HuggingFace. Alternative: python app.py --port 8808 opens a web demo at http://localhost:8808, with --device auto|cpu|mps|cuda|cuda:N.
Common workflows
- Design a voice with no reference audio:
voxcpm design --text "..." --control "Young female voice, warm and gentle" --seed 42 --output out.wav. - Clone a voice from a short clip:
voxcpm clone --text "..." --reference-audio voice.wav --output out.wav. - Ultimate-fidelity cloning: pass
prompt_wav_pathwith the clip,prompt_textwith its exact transcript, and optionally the same clip inreference_wav_path. - Batches:
voxcpm batch --input examples/input.txt --output-dir outs. With timestamps:pip install "voxcpm[timestamps]"and--timestamps --timestamp-level word. - Serve in production:
vllm serve openbmb/VoxCPM2 --omni --port 8000exposesPOST /v1/audio/speech, compatible with OpenAI clients. - Fine-tuning (5–10 minutes of audio is enough):
python scripts/train_voxcpm_finetune.py --config_path conf/voxcpm_v2/voxcpm_finetune_lora.yaml.
Essential configuration
cfg_value: diffusion guidance scale; the README uses 2.0.inference_timesteps: number of diffusion steps; the README uses 10.seed: random seed for reproducibility; without it, Voice Design and controllable cloning can vary between runs.load_denoiser: enables the denoiser for reference audio; recommendedFalsewhen not needed.optimize:optimize=Falsedisablestorch.compile; recommended if hitting Triton/torch.compile errors.
Common pitfalls and fixes
- Triton errors on Windows: install Triton from the community
triton-windowsproject and pair PyTorch↔Triton versions; if it persists,optimize=False. libtorchcodecwon’t load: torchaudio ≥ 2.9 requires FFmpeg installed; workaround:torchaudio.set_audio_backend("soundfile").torch.compileerrors on first run: disable it withoptimize=False.- MPS on Apple Silicon: if it errors, force
device="cpu". - Python 3.14+: installation can fail; use 3.10–3.12.
- Voice cutoff at the end of words (open issue with 51 comments): a known issue documented in the repository.
Integrations and migration
vLLM-Omni: the path to integrating into existing serving infrastructure. ComfyUI: several community nodes. TTS WebUI: an available extension. llama.cpp-omni: for on-device deployment without Python. Migrating from cloud APIs: Soniqo’s blog (May 2026) frames migration from ElevenLabs — same 48 kHz output, text-based voice design, but local and with no per-call cost.
Repo numbers
Measured: August 23, 2026, GitHub API.
| Metric | Value |
|---|---|
| Stars | 35,986 |
| Forks | 4,114 |
| Subscribers | 155 |
| Commits | 162 |
| Open issues | 114 |
| Primary language | Python |
| License | Apache-2.0 |
| Created | September 16, 2025 |
| Latest release | 2.0.3, May 11, 2026 |
Top contributors: Labmem-Zhouyx (29), a710128 (16), liuxin99 (13), VoxInstruct (11), ZMXJJ (5). PyPI (voxcpm 2.0.3): 2,794 downloads in the last day, 101,622 in the last month. HuggingFace: openbmb/VoxCPM2 logs 327,211 downloads and 1,542 likes.
How the community received it
Hacker News: the main submission, “VoxCPM: Tokenizer-Free TTS for Context-Aware Speech Generation and Voice Cloning” (December 5, 2025), got 3 points and 0 comments. No broad launch thread exists on HN.
GitHub (issues): the most active issue is “Voice cutoff at the end of words — Polish language” (open, 51 comments); the torchcodec issue (closed, 26 comments). The “When is VoxCPM2 getting released?” issue (12 comments) shows the pre-launch anticipation.
YouTube: “VoxCPM 2 Review - 2026 | Clone Realistic AI Voice in 30+ Languages” (57,828 views, Dan - Smart Tutorials); “VoxCPM 2 Review” (38,336 views, Bijan Bowen); a Spanish-language video, “¿El fin de ElevenLabs? OmniVoice vs VoxCPM2” (15,580 views). Reviews consistently frame it as local-vs-ElevenLabs.
Third-party blog: Soniqo (May 17, 2026) published a technical analysis of the pipeline and a comparison table against ElevenLabs, concluding the real difference “is whether the audio leaves the device or not.”
VoxCPM versus other proposals
| Project | Verifiable overlap | Verifiable difference |
|---|---|---|
| ElevenLabs | Voice cloning, text-based voice design, 48 kHz output, multilingual coverage. | Proprietary and hosted: audio is uploaded to its servers, cost per character, no local weights. |
| CosyVoice / CosyVoice3 | Open-source TTS with cloning; CosyVoice contributes the LocDiT Flow Matching implementation VoxCPM2 uses. | CosyVoice is tokenized-architecture and its v3 is closed; VoxCPM2 is tokenizer-free and covers 30 languages. |
| Qwen3-TTS (1.7B) | Open-source TTS compared in the README’s own tables. | Beats VoxCPM2 on English WER in self-reported tables. |
| FishAudio S2 (4B) | Directly compared in the README’s benchmarks. | Better overall WER; VoxCPM2 beats Fish on SIM across nearly all Minimax-MLS languages. |
The structural, verifiable differentiator versus most alternatives: native 48 kHz output with no discrete tokenizer + voice design + controllable cloning in a single 2B backbone under Apache-2.0.
Use cases and who this repository can help
- Developers who need local multilingual TTS: a single Apache-2.0 model covers 30 languages with the standard Python API, no per-character cost and no network dependency.
- Product teams with privacy or offline requirements: on-device cloning means audio never leaves the machine.
- High-traffic serving platforms: with vLLM-Omni, VoxCPM2 is served with PagedAttention and continuous batching behind an OpenAI-compatible endpoint.
- Content creators and animation studios: ComfyUI nodes plug the synthesis into node-based pipelines; YouDub-webui uses it as a dubbing engine.
- Researchers and teams fine-tuning voice models: documented SFT/LoRA fine-tuning lets you adapt the model to a specific speaker, language, or domain.
- Users on consumer and edge hardware: the llama.cpp-omni binary and GGUF weights extend use to laptops without an NVIDIA GPU.
Resources
- Repository: https://github.com/OpenBMB/VoxCPM
- Documentation: https://voxcpm.readthedocs.io/en/latest/
- Models: https://huggingface.co/openbmb/VoxCPM2
- Playground: https://huggingface.co/spaces/OpenBMB/VoxCPM-Demo
- Technical report (VoxCPM2): https://arxiv.org/abs/2606.06928
- Community: Discord https://discord.gg/KZUx7tVNwz
- Official site: https://voxcpm.com
Note: this article combines the README, ReadTheDocs FAQ, releases, and GitHub API of OpenBMB/VoxCPM, PyPI and HuggingFace counters, the Hacker News API, a YouTube search, and Soniqo’s blog, consulted on August 23, 2026. Benchmark figures are self-reported by OpenBMB. Figures change over time.
Comments