VibeVoice: Microsoft's open-source voice family, pulled for misuse and rebuilt around speech recognition
microsoft/VibeVoice · 54,494★ · 6,137 forks
An open-source family of voice models (TTS and ASR) from Microsoft, MIT-licensed, whose distinguishing trait is handling long audio: it synthesizes conversations of up to 90 minutes with up to 4 speakers and transcribes up to 60 minutes of audio in a single pass, with speaker identification and timestamps. The project has accumulated 53,121 stars, and its history includes an unusual episode: Microsoft temporarily pulled the original TTS model’s code over misuse, then rebuilt the project’s family around speech recognition (VibeVoice-ASR).
Origin
The repository was created on August 25, 2025. It launched alongside VibeVoice-TTS (7B “Large” and 1.5B variants), a long-form conversational text-to-speech model documented in the technical report arXiv:2508.19205, whose abstract presents it as a way to synthesize “the true conversational vibe,” outperforming both open and proprietary dialogue models. The work was accepted as an Oral at ICLR 2026.
The story’s central episode happened 9 days later: on September 4, 2025, Microsoft removed the 7B model from Hugging Face and deleted the repository; on September 5 it republished it without the code, with this notice reproduced in full in the current README: “VibeVoice is an open-source research framework intended to advance collaboration in the speech synthesis community. After release, we discovered instances where the tool was used in ways inconsistent with the stated intent. Since responsible use of AI is one of Microsoft’s guiding principles, we have removed the VibeVoice-TTS code from this repository.”
The community’s reaction was documented: backup forks and a PyPI package appeared, vibevoice 0.0.1 (uploaded September 26, 2025), whose changelog-style description records day by day the deletion, the codeless restoration, and the appearance of unofficial implementations. On the Hacker News thread about the removal, user mustaphah summarized it this way: the repo passed 8,000 stars in a few days, and upon review the code had disappeared and the stars dropped below 200 because it was deleted and a new one created.
From then on, the family was rebuilt around ASR and more controlled TTS variants: 2025-12-03, VibeVoice-Realtime-0.5B opened (real-time TTS with streaming text input); 2026-01-21, VibeVoice-ASR opened, the unified speech recognition model (60 minutes in a single pass, 50+ languages), with a technical report at arXiv:2601.18184; 2026-03-06, VibeVoice-ASR entered the Hugging Face Transformers library; 2026-03-12, it was integrated into Azure AI Foundry Labs; 2026-07-23, VibeASR.cpp launched, a CPU inference engine with heterogeneous quantization.

Philosophy and principles
The principles verifiable in the README, documentation, and CONTRIBUTING.md:
- Research ahead of product: the README explicitly states that “VibeVoice is not recommended for use in commercial or real-world applications without further testing and development.”
- Long-form voice as its own category: against the common practice of cutting audio into short fragments, the project’s thesis is to natively handle long audio within a 64K-token window.
- Efficiency via continuous 7.5 Hz tokenizers: the continuous voice tokenizer operates at an ultra-low frame rate of 7.5 Hz. The technical report claims it improves data compression 80-fold compared to EnCodec.
- Code minimalism: “our core principles are code minimalism, high readability, and functional purity.” Over-engineering, purely cosmetic PRs, and non-English contributions are explicitly rejected.
- Responsible use as a condition: deepfake risk is addressed in every model document.
- MIT license for code and weights, though the community debated its practical meaning once the repo went “codeless.”

How it works
VibeVoice isn’t a single tool but a family of models with a shared architecture: a next-token diffusion framework —the LLM understands textual context and dialogue flow, and a diffusion head generates fine acoustic detail— built on a Qwen2.5 base (1.5B in TTS-1.5B, 0.5B in Realtime-0.5B, 7B in ASR).

The repository’s active models: VibeVoice-ASR: structured “who/when/what” transcription — it performs ASR, diarization, and timestamping simultaneously, with custom-hotword support and 50+ languages with no need to specify the language; it handles up to 60 minutes of continuous audio in a single pass. VibeVoice-TTS: conversational generation of up to 90 minutes with up to 4 distinct speakers; the installation section literally reads “Disabled due to widespread misuse.” VibeVoice-Realtime-0.5B: real-time TTS (~200–300 ms latency to the first audible word), streaming text input, ~10 minutes of robust generation. VibeVoice-ASR-BitNet (in microsoft/VibeASR.cpp): a CPU variant with heterogeneous quantization, compressing the model from 4.62 GB to 1.58 GB.



The ecosystem
Official Microsoft repos: microsoft/VibeASR.cpp — the official CPU inference engine (166 stars, 21 forks); Hugging Face Transformers (microsoft/VibeVoice-ASR-HF); a Hugging Face collection bundling the weights; and integration into Azure AI Foundry Labs.
Searching GitHub for “VibeVoice” returns 371 repositories. The most notable by stars: Enemyx-net/VibeVoice-ComfyUI (1,547, a ComfyUI integration for the TTS), vibevoice-community/VibeVoice (1,527, a community fork of the original model), wildminder/ComfyUI-VibeVoice (596), zhao-kun/VibeVoiceFusion (490), voicepowered-ai/VibeVoice-finetuning (373), mpaepper/vibevoice (160, local STT with faster-whisper), localai-org/vibevoice.cpp (121, a C++ port over ggml), marhensa/vibevoice-realtime-openai-api (88), danielclough/vibevoice-rs (66, Rust).
The vibevoice 0.0.1 PyPI package is a backup of the original TTS code from before the deletion. Simon Willison documented a one-liner to run VibeVoice-ASR on a Mac with mlx-audio and the 4-bit MLX conversion — evidence of a community conversion ecosystem (MLX, GGUF, quantizations).

Official / semi-official status
- An official project of the
microsoftorganization on GitHub, with its own site, technical reports on arXiv, and acceptance as an Oral at ICLR 2026. - Integration into Azure AI Foundry Labs (March 12, 2026): the only verifiable “vendor backing,” and it comes from the vendor itself.
- Distribution in Hugging Face Transformers (March 6, 2026) for the ASR.
- A permanent official restriction on the 7B TTS: the original model is “Disabled” in the README’s table, its installation disabled “due to widespread misuse.”
- No verifiable “de facto standard” status; the README states the project is for research and discourages commercial use.

Repo numbers
From the GitHub API, as of 00:35 UTC on August 23, 2026:
| Metric | Value |
|---|---|
| Stars | 53,121 |
| Forks | 5,990 |
| Subscribers (watchers) | 259 |
| Open issues + PRs | 183 |
| Commits on the main branch | 140 |
| Releases | none |
| Language | Python |
| License | MIT |
Top contributors: YaoyaoChang (50), MSLDCherryPick (18), pengzhiliang (16), jsoref (12), Damon-Salvetore (10). Total Hugging Face downloads: VibeVoice-ASR 697,553, VibeVoice-Realtime-0.5B 647,155, VibeVoice-1.5B 119,947, VibeVoice-ASR-BitNet 19,273.
Quick-start guide
Installation and first run
Prerequisites: an NVIDIA GPU (recommended) and the NVIDIA Deep Learning Container for CUDA.
For the ASR:
git clone https://github.com/microsoft/VibeVoice.git
cd VibeVoice
pip install -e .
For real-time TTS:
git clone https://github.com/microsoft/VibeVoice.git
cd VibeVoice/
pip install -e .[streamingtts]
For VibeASR.cpp (CPU, no GPU):
git clone --recursive https://github.com/microsoft/VibeASR.cpp.git
cd VibeASR.cpp
pip install -r requirements.txt
python setup_env.py
The model downloads to the Hugging Face cache; the ASR-7B takes up 17.3 GB in FP16.
Common workflows
- Transcribe a long meeting:
python demo/vibevoice_asr_gradio_demo.py --model_path microsoft/VibeVoice-ASR --share(requiresffmpeg). - Transcribe with domain hotwords:
python demo/vibevoice_asr_inference_from_file.py --model_path microsoft/VibeVoice-ASR --audio_files [path]. - Try real-time TTS without your own GPU: the Colab notebook, or
python demo/vibevoice_realtime_demo.py --model_path microsoft/VibeVoice-Realtime-0.5B. - Serve the ASR as an OpenAI-compatible API:
docker run -d --gpus all ... vllm/vllm-openai:v0.14.1 -c "python3 /app/vllm_plugin/scripts/start_server.py --dp 4".
Essential configuration
--model_path: the Hugging Face model name; downloads automatically to the HF cache.--hotwords: a list of terms that improve accuracy on domain jargon.VIBEVOICE_FFMPEG_MAX_CONCURRENCY(vLLM server, default 64).--tp/--dp(vLLM server): tensor and data parallelism.
Common pitfalls and fixes
- Random music/sounds in the TTS: these are spontaneous and uncontrollable (“an easter egg”); the Large model is more stable.
- Instability with Chinese text: use English punctuation, the Large variant, and split the text into several turns with the same speaker tag.
- Insufficient
--max-tokenson the ASR: the default (8192) covers ~25 minutes; a full hour requires raising it (e.g. to 32768). - “CUDA out of memory”: lower
--gpu-memory-utilization,--max-num-seqs, or--max-model-len. - The original 7B TTS is not officially available: the only routes are community forks/backups, with no official support.
Integrations and migration
vLLM: an official plugin exposing OpenAI-compatible endpoints. Hugging Face Transformers: the ASR is consumed as microsoft/VibeVoice-ASR-HF. ComfyUI: community nodes. MLX (Apple Silicon): via mlx-audio and a 4-bit conversion. Migrating from Whisper: VibeASR.cpp is directly benchmarked against Whisper.cpp (1.6–2.3× faster at comparable size).
Contributing
CONTRIBUTING.md documents strict criteria: an academic, research-oriented focus; preferred contributions are bug fixes and new features; over-engineering, purely cosmetic PRs, and non-English contributions are rejected; manual line-by-line review by maintainers, with explicit rejection of large blocks of LLM-generated code without rigorous human cleanup.
How the community received it
Hacker News — TTS launch (448 points, 170 comments): baal80spam says it “sounds VERY impressive”; the most repeated criticism is that male voices sound robotic (simiones, IshKebab, strangescript). TheAceOfHearts flagged a hardware problem: “generated a 66-second clip in 832 seconds on CPU with float32.” echelon: “close to emotional SOTA… wonder if ElevenLabs will keep their huge ARR”; odie5533 replied that Chatterbox feels more realistic to them.
Hacker News — code removal: a thread where mustaphah documented the before/after: over 8,000 stars in a few days, then “code gone, stars below 200.”
Hacker News — ASR revival (386 points, 181 comments): CubsFan1060 linked Simon Willison’s writeup, who ran the ASR on a MacBook Pro M5 Max using mlx-audio: 8 min 45 s to transcribe an hour of podcast. walthamstow: “seems pretty heavy for STT; Parakeet and Whisper are much smaller and perform well for quick dictation.”
VibeVoice versus other proposals
| Project | Relationship to VibeVoice |
|---|---|
| Whisper / Whisper.cpp (OpenAI) | The classic STT benchmark; VibeASR.cpp benchmarks against it directly (1.6–2.3× faster at comparable size). |
| Parakeet (NVIDIA) | A smaller, faster STT alternative for dictation. |
| SenseVoice / FunASR (Alibaba) | Included in the same WER tables. |
| ElevenLabs | A commercial expressive-TTS reference cited by HN commenters. |
| Kokoro TTS | A local alternative cited as preferred, with phoneme (IPA) reading. |
| Chatterbox | Cited as “more realistic, without the robotic sound.” |
Use cases and who this repository can help
- Teams transcribing long meetings/podcasts: the 60-minute single-pass ASR with diarization, timestamps, and hotwords covers the full workflow without external chunking chains; the BitNet variant opens up GPU-free deployments.
- Product teams building live voice: VibeVoice-Realtime-0.5B is designed so assistants can start speaking from the first token.
- Conversational audio production: the 90-minute, 4-speaker TTS-1.5B covers dialogue prototyping; community ComfyUI integrations plug it into visual/audio pipelines.
- Researchers and ML speech teams: citable technical reports, documented LoRA fine-tuning, and a reference architecture.
- Apple Silicon / edge developers: the 4-bit MLX conversion and the VibeASR.cpp runtime enable local transcription without a GPU.
Resources
- Repository: https://github.com/microsoft/VibeVoice
- Documentation: https://microsoft.github.io/VibeVoice/
- Official models: https://huggingface.co/microsoft/VibeVoice-ASR
- ASR playground: https://aka.ms/vibevoice-asr
- Technical papers: TTS https://arxiv.org/abs/2508.19205 · ASR https://arxiv.org/abs/2601.18184
- Official sibling repo: https://github.com/microsoft/VibeASR.cpp
- Reviews: Simon Willison (April 27, 2026) https://simonwillison.net/2026/Apr/27/vibevoice/
Note: this article combines VibeVoice’s README and official documentation, the GitHub API, technical papers on arXiv, Hacker News discussions, and Simon Willison’s writeup, consulted on August 22, 2026. Figures change over time. No recoverable Reddit evidence was obtained during this run.
Comments