August 28, 2026 · By YasKad
OpenBMB/VoxCPM

VoxCPM: tokenizer-free multilingual speech synthesis

OpenBMB/VoxCPM · 37,974★ · 4,314 forks

Everything worth knowing about OpenBMB/VoxCPM: an open-source, open-weight text-to-speech system that generates continuous audio directly, with no discrete tokenization, whose version 2 supports 30 languages, description-based voice design, and 48 kHz cloning.


What VoxCPM is

VoxCPM is a tokenizer-free text-to-speech (TTS) system: instead of converting audio into a vocabulary of discrete tokens and then decoding them, it generates continuous audio representations directly through an end-to-end autoregressive diffusion architecture. The product is an open-weight model (Apache-2.0) with a Python API, CLI, and web demo.

The current version is VoxCPM2, a 2B-parameter model trained on over 2 million hours of multilingual speech, supporting 30 languages and 9 Chinese dialects, Voice Design (creating a new voice from a natural-language description with no reference audio), controllable cloning (cloning a voice from a short clip and adjusting style, pace, and emotion), and ultimate-fidelity cloning via audio continuation, with native 48 kHz output. It’s built on the same organization’s MiniCPM-4 backbone.

Origin

The repository was created on September 16, 2025 by the OpenBMB organization, the MiniCPM model team at ModelBest (a Chinese company) in collaboration with THUHCSI (a Tsinghua University research group).

The timeline the README documents: September 2025, publication of the VoxCPM technical report and the launch of VoxCPM-0.5B, which reached #1 trending on HuggingFace (per the README); December 5, 2025, publication of VoxCPM1.5 weights with SFT and LoRA fine-tuning, #1 trending on GitHub (per the README); April 6, 2026, the launch of VoxCPM2 — 2B, 30 languages, Voice Design and controllable cloning, 48 kHz output — with a technical report on arXiv.

The line of work extends the MiniCPM family (10,222 stars), sharing its philosophy of small, capable models that run locally. The README explicitly credits DiTAR for the autoregressive diffusion backbone, CosyVoice (FunAudioLLM) for the LocDiT Flow Matching implementation, and DAC (descript-audio-codec) for the audio VAE backbone.

A detailed dark-mode cyberpunk isometric technical illustration of a four-stage tokenizer-free speech synthesis architecture. Four glowing modular stages are arranged in a pipeline: a local encoder fusing text and optional reference audio into one continuous vector stream, a 2B autoregressive language backbone deciding audio patches at 6.25 Hz, a residual refinement model adding prosodic detail, and a diffusion or flow-matching painter rendering each latent audio patch. The latent space is shown as a smooth continuous field, not discrete tokens, with a translucent AudioVAE V2 asymmetric decoder converting 16 kHz latent audio into a crisp 48 kHz waveform. Neon cyan, violet, and amber accents, dark grid background, ultra-detailed, 8K resolution, no readable text, abstract module icons only.

Philosophy and principles

The technical report and README reveal four coherent design decisions:

  1. No discrete tokenization: the model operates entirely in the AudioVAE V2 latent space and “paints” the audio with a diffusion step (LocDiT) at each position, instead of choosing among quantized tokens. The argument is that a discrete quantizer introduces a quality bottleneck; without it, the VAE can reconstruct directly at 48 kHz.
  2. One model, several modes: the technical report describes a “unified sequence organization” where every generation mode (text-to-speech, voice design, reference cloning, continuation cloning) is expressed as different arrangements of the same input blocks.
  3. Studio quality without an external upsampler: AudioVAE V2 has an asymmetric design — encoding at 16 kHz and reconstructing at 48 kHz — building implicit super-resolution into the decoder itself.
  4. Full openness and commercial use: weights and code under Apache-2.0, “free for commercial use” per the README. The README includes a “Risks and Limitations” section that explicitly prohibits use for impersonation, fraud, or disinformation and recommends labeling generated content.

A dark-mode cyberpunk image about responsible speech synthesis. A central audio waveform is protected by a translucent shield, with warning icons blocking misuse scenarios such as impersonation, fraud, and disinformation. A small holographic label icon indicates generated-content labeling, while a balance scale and lock represent ethical guardrails. The background shows a dark data center with restrained red warning accents and calm cyan safety lines, conveying risk limits and responsible use. Ultra-detailed, 8K resolution, no readable text.

A conceptual dark-mode cyberpunk image about voice design by natural language description. A glowing speech bubble and microphone icon dissolve into a new synthesized voice portrait made of soft neon waveform particles. On the left, abstract descriptive words appear as floating luminous glyphs, with no reference audio clip present, emphasizing text-only voice creation. The emerging voice has a gentle, warm timbre represented by a smooth magenta and cyan waveform, while 30 language nodes hover in the background. Dark tech background, holographic UI panels, neon accents, ultra-detailed, 8K resolution, no readable text.

How it works

VoxCPM2’s architecture follows four stages operating in the AudioVAE V2 latent space: LocEnc (local encoder, fusing text and optional reference audio); TSLM (the autoregressive language model, a 2B-parameter MiniCPM-4 backbone, deciding which audio “patch” comes next, at 6.25 Hz); RALM (residual autoregressive language model, a second pass refining each patch with prosodic detail); LocDiT (a diffusion estimator that paints the audio latent, decoded by AudioVAE V2 to 48 kHz).

The four documented generation modes: Text-to-Speech (model.generate(text=...)); Voice Design (the description goes in parentheses at the start of the text, no reference audio required); Controllable Voice Cloning (pass reference_wav_path, a 5–30s clip, with optional style instructions); Ultimate Cloning (pass prompt_wav_path and prompt_text, the clip’s exact transcript).

A dark-mode cyberpunk visualization of controllable voice cloning and ultimate audio continuation. A short reference audio clip is rendered as a compact neon capsule on a timeline; from it, a continuous waveform extends forward, preserving timbre, rhythm, emotion, and style. Three control dials float nearby, labeled only by abstract icons for style, speed, and emotion, with smooth glowing sliders. A second capsule shows prompt audio plus exact transcript as abstract aligned text particles, enabling seamless continuation. 48 kHz waveform details, dark circuit background, cyan and orange neon, ultra-detailed, 8K resolution, no readable text.

There’s also a streaming API (model.generate_streaming) and a CLI (voxcpm design, voxcpm clone, voxcpm batch) with word- or character-level timestamp support. The README reports RTF ~0.30 on an NVIDIA RTX 4090 with standard PyTorch, ~0.13 with Nano-vLLM, and ~1.76 (Q8_0) on Apple M4 Pro / Metal via llama.cpp-omni. VRAM: ~8 GB for VoxCPM2.

A dark-mode cyberpunk globe representing 30-language multilingual speech synthesis. The sphere is built from translucent continents connected by luminous audio beams, with 30 major language nodes glowing around it and 9 smaller Chinese dialect nodes orbiting the Asia region. Continuous waveforms travel between nodes instead of discrete token packets, suggesting tokenizer-free generation. Neon blue, green, and magenta accents, dark space-like background with subtle grid, floating abstract glyphs, ultra-detailed, 8K resolution, no readable text.

The ecosystem

The README maintains an official ecosystem table (community projects, not maintained by OpenBMB). Inference and deployment engines: vllm-project/vllm-omni (the official vLLM project’s omni-modal extension with native VoxCPM2 support, an OpenAI-compatible endpoint; 6,263 stars — the most relevant semi-official deployment, published by the vLLM organization, not OpenBMB); a710128/nanovllm-voxcpm (292 stars, RTF ~0.13 on a 4090); tc-mb/llama.cpp-omni (253 stars, native GGUF support on CPU/Metal/CUDA/Vulkan); bluryar/VoxCPM.cpp (89 stars); 0xShug0/audio.cpp (1,957 stars, a unified C++ framework for several audio models).

A dark-mode cyberpunk ecosystem diagram for deployment and inference engines surrounding a central VoxCPM-like neural core. Around the core, abstract module icons represent vLLM-omni, llama.cpp-omni, GGUF weights, ONNX export, Rust reimplementation, Apple Neural Engine, ComfyUI nodes, and an OpenAI-compatible API endpoint, connected by neon data pipes. The central core emits a continuous 48 kHz audio waveform into a server rack, laptop, phone, and edge device, showing local and cloud deployment. Dark tech background, holographic panels, cyan and violet neon, ultra-detailed, 8K resolution, no readable text.

Community interfaces and apps: ComfyUI nodes (wildminder/ComfyUI-VoxCPM, 509 stars; Saganaki22/ComfyUI-VoxCPM2, 195); liuzhao1225/YouDub-webui (5,345 stars, a video-localization tool with voice-clone dubbing — the largest adoption case as an engine inside another application); maomao-2001/Whispera (100 stars, real-time conversation with streaming TTS).

A dark-mode cyberpunk image of open-source open-weight speech AI. A transparent neural model core floats above a local workstation, with an unlocked holographic license seal and an Apache-style abstract emblem, emphasizing commercial-friendly open weights. Around it, small devices include a GPU card, laptop, phone, and desktop PC, with a subtle 8 GB VRAM gauge and a crisp 48 kHz studio waveform. The scene feels accessible and local-first, with neon cyan and amber accents, dark grid, ultra-detailed, 8K resolution, no readable text.

Official / semi-official status

  • vLLM-Omni publishes native VoxCPM2 support with deployment examples in its own repository: this is adoption by the official vLLM project, which in practice means VoxCPM2 can be served with the same infrastructure as the highest-traffic language models, with an OpenAI-compatible API.
  • ICLR 2026: the README links VoxCPM-0.5B’s review on OpenReview under “ICLR 2026.” Acceptance couldn’t be verified directly in this run; it’s cited as the status the project itself declares.
  • The README states VoxCPM-0.5B was #1 trending on HuggingFace and VoxCPM1.5 was #1 trending on GitHub. These are the README’s own claims.
  • There’s no official marketplace in the agent sense: the de facto status is more that of a reference for local multilingual TTS, reinforced by presence on PyPI, HuggingFace, ModelScope, and the engine ecosystem.

A dark-mode cyberpunk timeline of a speech model family's evolution. A glowing horizontal path moves from left to right through three major milestones: a compact 0.5B model release, a 1.5B intermediate version with fine-tuning and LoRA, and a 2B flagship release with 30 languages, voice design, controllable cloning, and 48 kHz output. Abstract trending badges appear as neon sparkles, and a MiniCPM family emblem is suggested by layered rings. Background includes a research lab, an OpenBMB-style abstract logo, and a Tsinghua-inspired academic motif, all rendered as futuristic holograms. Ultra-detailed, 8K resolution, no readable text.

Quick-start guide

Installation and first run

Documented prerequisites: Python ≥ 3.10 and < 3.13, PyTorch ≥ 2.5.0, CUDA ≥ 12.0 for an NVIDIA GPU. VoxCPM2 uses ~8 GB of VRAM.

pip install voxcpm

First run (Python API):

from voxcpm import VoxCPM
import soundfile as sf

model = VoxCPM.from_pretrained(
  "openbmb/VoxCPM2",
  load_denoiser=False,
)
wav = model.generate(
    text="VoxCPM2 is the current recommended release for realistic multilingual speech synthesis.",
    cfg_value=2.0,
    inference_timesteps=10,
    seed=42,
)
sf.write("demo.wav", wav, model.tts_model.sample_rate)

The first call downloads the weights from HuggingFace. Alternative: python app.py --port 8808 opens a web demo at http://localhost:8808, with --device auto|cpu|mps|cuda|cuda:N.

Common workflows

  • Design a voice with no reference audio: voxcpm design --text "..." --control "Young female voice, warm and gentle" --seed 42 --output out.wav.
  • Clone a voice from a short clip: voxcpm clone --text "..." --reference-audio voice.wav --output out.wav.
  • Ultimate-fidelity cloning: pass prompt_wav_path with the clip, prompt_text with its exact transcript, and optionally the same clip in reference_wav_path.
  • Batches: voxcpm batch --input examples/input.txt --output-dir outs. With timestamps: pip install "voxcpm[timestamps]" and --timestamps --timestamp-level word.
  • Serve in production: vllm serve openbmb/VoxCPM2 --omni --port 8000 exposes POST /v1/audio/speech, compatible with OpenAI clients.
  • Fine-tuning (5–10 minutes of audio is enough): python scripts/train_voxcpm_finetune.py --config_path conf/voxcpm_v2/voxcpm_finetune_lora.yaml.

Essential configuration

  • cfg_value: diffusion guidance scale; the README uses 2.0.
  • inference_timesteps: number of diffusion steps; the README uses 10.
  • seed: random seed for reproducibility; without it, Voice Design and controllable cloning can vary between runs.
  • load_denoiser: enables the denoiser for reference audio; recommended False when not needed.
  • optimize: optimize=False disables torch.compile; recommended if hitting Triton/torch.compile errors.

Common pitfalls and fixes

  • Triton errors on Windows: install Triton from the community triton-windows project and pair PyTorch↔Triton versions; if it persists, optimize=False.
  • libtorchcodec won’t load: torchaudio ≥ 2.9 requires FFmpeg installed; workaround: torchaudio.set_audio_backend("soundfile").
  • torch.compile errors on first run: disable it with optimize=False.
  • MPS on Apple Silicon: if it errors, force device="cpu".
  • Python 3.14+: installation can fail; use 3.10–3.12.
  • Voice cutoff at the end of words (open issue with 51 comments): a known issue documented in the repository.

Integrations and migration

vLLM-Omni: the path to integrating into existing serving infrastructure. ComfyUI: several community nodes. TTS WebUI: an available extension. llama.cpp-omni: for on-device deployment without Python. Migrating from cloud APIs: Soniqo’s blog (May 2026) frames migration from ElevenLabs — same 48 kHz output, text-based voice design, but local and with no per-call cost.

Repo numbers

Measured: August 23, 2026, GitHub API.

MetricValue
Stars35,986
Forks4,114
Subscribers155
Commits162
Open issues114
Primary languagePython
LicenseApache-2.0
CreatedSeptember 16, 2025
Latest release2.0.3, May 11, 2026

Top contributors: Labmem-Zhouyx (29), a710128 (16), liuxin99 (13), VoxInstruct (11), ZMXJJ (5). PyPI (voxcpm 2.0.3): 2,794 downloads in the last day, 101,622 in the last month. HuggingFace: openbmb/VoxCPM2 logs 327,211 downloads and 1,542 likes.

How the community received it

Hacker News: the main submission, “VoxCPM: Tokenizer-Free TTS for Context-Aware Speech Generation and Voice Cloning” (December 5, 2025), got 3 points and 0 comments. No broad launch thread exists on HN.

GitHub (issues): the most active issue is “Voice cutoff at the end of words — Polish language” (open, 51 comments); the torchcodec issue (closed, 26 comments). The “When is VoxCPM2 getting released?” issue (12 comments) shows the pre-launch anticipation.

YouTube: “VoxCPM 2 Review - 2026 | Clone Realistic AI Voice in 30+ Languages” (57,828 views, Dan - Smart Tutorials); “VoxCPM 2 Review” (38,336 views, Bijan Bowen); a Spanish-language video, “¿El fin de ElevenLabs? OmniVoice vs VoxCPM2” (15,580 views). Reviews consistently frame it as local-vs-ElevenLabs.

Third-party blog: Soniqo (May 17, 2026) published a technical analysis of the pipeline and a comparison table against ElevenLabs, concluding the real difference “is whether the audio leaves the device or not.”

VoxCPM versus other proposals

ProjectVerifiable overlapVerifiable difference
ElevenLabsVoice cloning, text-based voice design, 48 kHz output, multilingual coverage.Proprietary and hosted: audio is uploaded to its servers, cost per character, no local weights.
CosyVoice / CosyVoice3Open-source TTS with cloning; CosyVoice contributes the LocDiT Flow Matching implementation VoxCPM2 uses.CosyVoice is tokenized-architecture and its v3 is closed; VoxCPM2 is tokenizer-free and covers 30 languages.
Qwen3-TTS (1.7B)Open-source TTS compared in the README’s own tables.Beats VoxCPM2 on English WER in self-reported tables.
FishAudio S2 (4B)Directly compared in the README’s benchmarks.Better overall WER; VoxCPM2 beats Fish on SIM across nearly all Minimax-MLS languages.

The structural, verifiable differentiator versus most alternatives: native 48 kHz output with no discrete tokenizer + voice design + controllable cloning in a single 2B backbone under Apache-2.0.

Use cases and who this repository can help

  • Developers who need local multilingual TTS: a single Apache-2.0 model covers 30 languages with the standard Python API, no per-character cost and no network dependency.
  • Product teams with privacy or offline requirements: on-device cloning means audio never leaves the machine.
  • High-traffic serving platforms: with vLLM-Omni, VoxCPM2 is served with PagedAttention and continuous batching behind an OpenAI-compatible endpoint.
  • Content creators and animation studios: ComfyUI nodes plug the synthesis into node-based pipelines; YouDub-webui uses it as a dubbing engine.
  • Researchers and teams fine-tuning voice models: documented SFT/LoRA fine-tuning lets you adapt the model to a specific speaker, language, or domain.
  • Users on consumer and edge hardware: the llama.cpp-omni binary and GGUF weights extend use to laptops without an NVIDIA GPU.

Resources


Note: this article combines the README, ReadTheDocs FAQ, releases, and GitHub API of OpenBMB/VoxCPM, PyPI and HuggingFace counters, the Hacker News API, a YouTube search, and Soniqo’s blog, consulted on August 23, 2026. Benchmark figures are self-reported by OpenBMB. Figures change over time.

Comments