August 27, 2026 · By YasKad
microsoft/VibeVoice

VibeVoice: Microsoft's open-source voice family, pulled for misuse and rebuilt around speech recognition

microsoft/VibeVoice · 54,494★ · 6,137 forks

An open-source family of voice models (TTS and ASR) from Microsoft, MIT-licensed, whose distinguishing trait is handling long audio: it synthesizes conversations of up to 90 minutes with up to 4 speakers and transcribes up to 60 minutes of audio in a single pass, with speaker identification and timestamps. The project has accumulated 53,121 stars, and its history includes an unusual episode: Microsoft temporarily pulled the original TTS model’s code over misuse, then rebuilt the project’s family around speech recognition (VibeVoice-ASR).


Origin

The repository was created on August 25, 2025. It launched alongside VibeVoice-TTS (7B “Large” and 1.5B variants), a long-form conversational text-to-speech model documented in the technical report arXiv:2508.19205, whose abstract presents it as a way to synthesize “the true conversational vibe,” outperforming both open and proprietary dialogue models. The work was accepted as an Oral at ICLR 2026.

The story’s central episode happened 9 days later: on September 4, 2025, Microsoft removed the 7B model from Hugging Face and deleted the repository; on September 5 it republished it without the code, with this notice reproduced in full in the current README: “VibeVoice is an open-source research framework intended to advance collaboration in the speech synthesis community. After release, we discovered instances where the tool was used in ways inconsistent with the stated intent. Since responsible use of AI is one of Microsoft’s guiding principles, we have removed the VibeVoice-TTS code from this repository.”

The community’s reaction was documented: backup forks and a PyPI package appeared, vibevoice 0.0.1 (uploaded September 26, 2025), whose changelog-style description records day by day the deletion, the codeless restoration, and the appearance of unofficial implementations. On the Hacker News thread about the removal, user mustaphah summarized it this way: the repo passed 8,000 stars in a few days, and upon review the code had disappeared and the stars dropped below 200 because it was deleted and a new one created.

From then on, the family was rebuilt around ASR and more controlled TTS variants: 2025-12-03, VibeVoice-Realtime-0.5B opened (real-time TTS with streaming text input); 2026-01-21, VibeVoice-ASR opened, the unified speech recognition model (60 minutes in a single pass, 50+ languages), with a technical report at arXiv:2601.18184; 2026-03-06, VibeVoice-ASR entered the Hugging Face Transformers library; 2026-03-12, it was integrated into Azure AI Foundry Labs; 2026-07-23, VibeASR.cpp launched, a CPU inference engine with heterogeneous quantization.

A dramatic cyberpunk visualization of the origin and temporary removal of the Microsoft VibeVoice repository, a glowing GitHub-style repository panel suspended in a dark digital void, a red deletion beam sweeping through the main TTS code section, backup forks emerging as branching neon constellations, a small PyPI package capsule floating nearby as a community backup, abstract Hacker News comment threads represented as translucent ribbons of light, a timeline showing the initial release, the removal, and the later reconstruction around ASR, dark mode, cyberpunk tech aesthetic, neon accents, ultra-detailed, 8K resolution

Philosophy and principles

The principles verifiable in the README, documentation, and CONTRIBUTING.md:

  • Research ahead of product: the README explicitly states that “VibeVoice is not recommended for use in commercial or real-world applications without further testing and development.”
  • Long-form voice as its own category: against the common practice of cutting audio into short fragments, the project’s thesis is to natively handle long audio within a 64K-token window.
  • Efficiency via continuous 7.5 Hz tokenizers: the continuous voice tokenizer operates at an ultra-low frame rate of 7.5 Hz. The technical report claims it improves data compression 80-fold compared to EnCodec.
  • Code minimalism: “our core principles are code minimalism, high readability, and functional purity.” Over-engineering, purely cosmetic PRs, and non-English contributions are explicitly rejected.
  • Responsible use as a condition: deepfake risk is addressed in every model document.
  • MIT license for code and weights, though the community debated its practical meaning once the repo went “codeless.”

A conceptual image of VibeVoice's research-first philosophy and long-audio principle, a minimalist dark interface displaying a continuous 64K token window as a long glowing track, tiny 7.5 Hz continuous voice tokens moving in a precise low-rate stream, a shield icon representing responsible AI use, a subtle warning glyph indicating research-only use, clean functional code panels with abstract syntax, no clutter, pure functional design, elegant neon lines, dark mode, cyberpunk tech aesthetic, cyan and violet neon accents, ultra-detailed, 8K resolution

How it works

VibeVoice isn’t a single tool but a family of models with a shared architecture: a next-token diffusion framework —the LLM understands textual context and dialogue flow, and a diffusion head generates fine acoustic detail— built on a Qwen2.5 base (1.5B in TTS-1.5B, 0.5B in Realtime-0.5B, 7B in ASR).

A detailed architectural visualization of VibeVoice's next-token diffusion framework, a large translucent Qwen2.5 language model core on the left understanding dialogue context and text flow, a diffusion head on the right generating fine acoustic detail, token arrows flowing from semantic text understanding into high-fidelity waveform synthesis, a layered neural map showing 1.5B, 0.5B, and 7B model variants as nested glowing cores, a 64K token context window rendered as a curved luminous ring, dark mode, cyberpunk tech aesthetic, neon accents, ultra-detailed, 8K resolution

The repository’s active models: VibeVoice-ASR: structured “who/when/what” transcription — it performs ASR, diarization, and timestamping simultaneously, with custom-hotword support and 50+ languages with no need to specify the language; it handles up to 60 minutes of continuous audio in a single pass. VibeVoice-TTS: conversational generation of up to 90 minutes with up to 4 distinct speakers; the installation section literally reads “Disabled due to widespread misuse.” VibeVoice-Realtime-0.5B: real-time TTS (~200–300 ms latency to the first audible word), streaming text input, ~10 minutes of robust generation. VibeVoice-ASR-BitNet (in microsoft/VibeASR.cpp): a CPU variant with heterogeneous quantization, compressing the model from 4.62 GB to 1.58 GB.

A high-tech scene representing VibeVoice-ASR, a 60-minute continuous audio file being processed in a single pass, structured transcription panels showing who spoke, when it happened, and what was said, speaker diarization visualized as color-coded voice lanes, timestamp markers glowing along the waveform, hotword tokens highlighted as small neon chips, multilingual support represented by abstract language glyphs for more than 50 languages, code-switching shown as blended waveform colors, dark mode, cyberpunk tech aesthetic, cyan green and magenta neon accents, ultra-detailed, 8K resolution

A cinematic visualization of VibeVoice TTS long-form generation, a 90-minute conversational waveform unfolding across a dark cyberpunk stage, four distinct speaker avatars speaking in natural turns, each with a unique neon color signature, dialogue continuity shown as smooth connecting arcs between speakers, a locked 7B model section marked with a red restricted-use seal, a smaller 1.5B model and a 0.5B realtime model shown as compact glowing engines, streaming text input flowing from one side into instant speech output, dark mode, cyberpunk tech aesthetic, neon accents, ultra-detailed, 8K resolution

A technical visualization of VibeASR.cpp CPU inference, a compact laptop or embedded CPU board running speech recognition without a GPU, a large 4.62 GB model being compressed into a 1.58 GB BitNet-optimized engine, heterogeneous quantization shown as I8_S and I2_S token streams, multiple CPU threads visualized as parallel neon pipelines, a real-time factor indicator below 1 displayed as a smooth glowing gauge, waveform input and text output moving in synchronized streams, dark mode, cyberpunk tech aesthetic, neon accents, ultra-detailed, 8K resolution

The ecosystem

Official Microsoft repos: microsoft/VibeASR.cpp — the official CPU inference engine (166 stars, 21 forks); Hugging Face Transformers (microsoft/VibeVoice-ASR-HF); a Hugging Face collection bundling the weights; and integration into Azure AI Foundry Labs.

Searching GitHub for “VibeVoice” returns 371 repositories. The most notable by stars: Enemyx-net/VibeVoice-ComfyUI (1,547, a ComfyUI integration for the TTS), vibevoice-community/VibeVoice (1,527, a community fork of the original model), wildminder/ComfyUI-VibeVoice (596), zhao-kun/VibeVoiceFusion (490), voicepowered-ai/VibeVoice-finetuning (373), mpaepper/vibevoice (160, local STT with faster-whisper), localai-org/vibevoice.cpp (121, a C++ port over ggml), marhensa/vibevoice-realtime-openai-api (88), danielclough/vibevoice-rs (66, Rust).

The vibevoice 0.0.1 PyPI package is a backup of the original TTS code from before the deletion. Simon Willison documented a one-liner to run VibeVoice-ASR on a Mac with mlx-audio and the 4-bit MLX conversion — evidence of a community conversion ecosystem (MLX, GGUF, quantizations).

An ecosystem map of the VibeVoice project, a central glowing VibeVoice core connected to official and community nodes, Hugging Face Transformers shown as a standard library portal, Azure AI Foundry Labs shown as a cloud service gateway, VibeASR.cpp shown as a CPU inference chip, community forks represented as branching neon repositories, MLX and GGUF conversion nodes rendered as compact model cartridges, a web studio interface and ComfyUI node panel floating nearby, dark mode, cyberpunk tech aesthetic, neon accents, ultra-detailed, 8K resolution

Official / semi-official status

  • An official project of the microsoft organization on GitHub, with its own site, technical reports on arXiv, and acceptance as an Oral at ICLR 2026.
  • Integration into Azure AI Foundry Labs (March 12, 2026): the only verifiable “vendor backing,” and it comes from the vendor itself.
  • Distribution in Hugging Face Transformers (March 6, 2026) for the ASR.
  • A permanent official restriction on the 7B TTS: the original model is “Disabled” in the README’s table, its installation disabled “due to widespread misuse.”
  • No verifiable “de facto standard” status; the README states the project is for research and discourages commercial use.

A responsible-use and safety-themed image for VibeVoice, a translucent shield protecting a speech waveform from misuse, deepfake risk represented as distorted glitched voice clones being filtered out, embedded voice sample capsules shown as secure neon containers, a MIT license seal and research-only badge rendered as abstract emblems, community feedback streams flowing from public discussion into improved safeguards, a balanced composition of openness and accountability, dark mode, cyberpunk tech aesthetic, cyan violet and amber neon accents, ultra-detailed, 8K resolution

Repo numbers

From the GitHub API, as of 00:35 UTC on August 23, 2026:

MetricValue
Stars53,121
Forks5,990
Subscribers (watchers)259
Open issues + PRs183
Commits on the main branch140
Releasesnone
LanguagePython
LicenseMIT

Top contributors: YaoyaoChang (50), MSLDCherryPick (18), pengzhiliang (16), jsoref (12), Damon-Salvetore (10). Total Hugging Face downloads: VibeVoice-ASR 697,553, VibeVoice-Realtime-0.5B 647,155, VibeVoice-1.5B 119,947, VibeVoice-ASR-BitNet 19,273.

Quick-start guide

Installation and first run

Prerequisites: an NVIDIA GPU (recommended) and the NVIDIA Deep Learning Container for CUDA.

For the ASR:

git clone https://github.com/microsoft/VibeVoice.git
cd VibeVoice
pip install -e .

For real-time TTS:

git clone https://github.com/microsoft/VibeVoice.git
cd VibeVoice/
pip install -e .[streamingtts]

For VibeASR.cpp (CPU, no GPU):

git clone --recursive https://github.com/microsoft/VibeASR.cpp.git
cd VibeASR.cpp
pip install -r requirements.txt
python setup_env.py

The model downloads to the Hugging Face cache; the ASR-7B takes up 17.3 GB in FP16.

Common workflows

  1. Transcribe a long meeting: python demo/vibevoice_asr_gradio_demo.py --model_path microsoft/VibeVoice-ASR --share (requires ffmpeg).
  2. Transcribe with domain hotwords: python demo/vibevoice_asr_inference_from_file.py --model_path microsoft/VibeVoice-ASR --audio_files [path].
  3. Try real-time TTS without your own GPU: the Colab notebook, or python demo/vibevoice_realtime_demo.py --model_path microsoft/VibeVoice-Realtime-0.5B.
  4. Serve the ASR as an OpenAI-compatible API: docker run -d --gpus all ... vllm/vllm-openai:v0.14.1 -c "python3 /app/vllm_plugin/scripts/start_server.py --dp 4".

Essential configuration

  • --model_path: the Hugging Face model name; downloads automatically to the HF cache.
  • --hotwords: a list of terms that improve accuracy on domain jargon.
  • VIBEVOICE_FFMPEG_MAX_CONCURRENCY (vLLM server, default 64).
  • --tp / --dp (vLLM server): tensor and data parallelism.

Common pitfalls and fixes

  • Random music/sounds in the TTS: these are spontaneous and uncontrollable (“an easter egg”); the Large model is more stable.
  • Instability with Chinese text: use English punctuation, the Large variant, and split the text into several turns with the same speaker tag.
  • Insufficient --max-tokens on the ASR: the default (8192) covers ~25 minutes; a full hour requires raising it (e.g. to 32768).
  • “CUDA out of memory”: lower --gpu-memory-utilization, --max-num-seqs, or --max-model-len.
  • The original 7B TTS is not officially available: the only routes are community forks/backups, with no official support.

Integrations and migration

vLLM: an official plugin exposing OpenAI-compatible endpoints. Hugging Face Transformers: the ASR is consumed as microsoft/VibeVoice-ASR-HF. ComfyUI: community nodes. MLX (Apple Silicon): via mlx-audio and a 4-bit conversion. Migrating from Whisper: VibeASR.cpp is directly benchmarked against Whisper.cpp (1.6–2.3× faster at comparable size).

Contributing

CONTRIBUTING.md documents strict criteria: an academic, research-oriented focus; preferred contributions are bug fixes and new features; over-engineering, purely cosmetic PRs, and non-English contributions are rejected; manual line-by-line review by maintainers, with explicit rejection of large blocks of LLM-generated code without rigorous human cleanup.

How the community received it

Hacker News — TTS launch (448 points, 170 comments): baal80spam says it “sounds VERY impressive”; the most repeated criticism is that male voices sound robotic (simiones, IshKebab, strangescript). TheAceOfHearts flagged a hardware problem: “generated a 66-second clip in 832 seconds on CPU with float32.” echelon: “close to emotional SOTA… wonder if ElevenLabs will keep their huge ARR”; odie5533 replied that Chatterbox feels more realistic to them.

Hacker News — code removal: a thread where mustaphah documented the before/after: over 8,000 stars in a few days, then “code gone, stars below 200.”

Hacker News — ASR revival (386 points, 181 comments): CubsFan1060 linked Simon Willison’s writeup, who ran the ASR on a MacBook Pro M5 Max using mlx-audio: 8 min 45 s to transcribe an hour of podcast. walthamstow: “seems pretty heavy for STT; Parakeet and Whisper are much smaller and perform well for quick dictation.”

VibeVoice versus other proposals

ProjectRelationship to VibeVoice
Whisper / Whisper.cpp (OpenAI)The classic STT benchmark; VibeASR.cpp benchmarks against it directly (1.6–2.3× faster at comparable size).
Parakeet (NVIDIA)A smaller, faster STT alternative for dictation.
SenseVoice / FunASR (Alibaba)Included in the same WER tables.
ElevenLabsA commercial expressive-TTS reference cited by HN commenters.
Kokoro TTSA local alternative cited as preferred, with phoneme (IPA) reading.
ChatterboxCited as “more realistic, without the robotic sound.”

Use cases and who this repository can help

  • Teams transcribing long meetings/podcasts: the 60-minute single-pass ASR with diarization, timestamps, and hotwords covers the full workflow without external chunking chains; the BitNet variant opens up GPU-free deployments.
  • Product teams building live voice: VibeVoice-Realtime-0.5B is designed so assistants can start speaking from the first token.
  • Conversational audio production: the 90-minute, 4-speaker TTS-1.5B covers dialogue prototyping; community ComfyUI integrations plug it into visual/audio pipelines.
  • Researchers and ML speech teams: citable technical reports, documented LoRA fine-tuning, and a reference architecture.
  • Apple Silicon / edge developers: the 4-bit MLX conversion and the VibeASR.cpp runtime enable local transcription without a GPU.

Resources


Note: this article combines VibeVoice’s README and official documentation, the GitHub API, technical papers on arXiv, Hacker News discussions, and Simon Willison’s writeup, consulted on August 22, 2026. Figures change over time. No recoverable Reddit evidence was obtained during this run.

Comments