¿Quién creó VibeVoice-ASR?

VibeVoice-ASR fue publicado por Microsoft en Hugging Face.

VibeVoice-ASR

Name: VibeVoice-ASR
Author: Microsoft

Modelo ASR multilingüe de 8.7B parámetros de Microsoft compatible con transcripción y diarización de hablantes en más de 60 idiomas.

VibeVoice-ASR

VibeVoice-ASR is a unified speech-to-text model designed to handle 60-minute long-form audio in a single pass, generating structured transcriptions containing Who (Speaker), When (Timestamps), and What (Content), with support for Customized Hotwords and over 50 languages.

➡️ Code: microsoft/VibeVoice
➡️ Demo: VibeVoice-ASR-Demo
➡️ Report: VibeVoice-ASR Technical Report
➡️ Finetuning: Finetuning
➡️ vLLM: vLLM-VibeVoice-ASR

🔥 Key Features

🕒 60-minute Single-Pass Processing: Unlike conventional ASR models that slice audio into short chunks (often losing global context), VibeVoice ASR accepts up to 60 minutes of continuous audio input within 64K token length. This ensures consistent speaker tracking and semantic coherence across the entire hour.
👤 Customized Hotwords: Users can provide customized hotwords (e.g., specific names, technical terms, or background info) to guide the recognition process, significantly improving accuracy on domain-specific content.
📝 Rich Transcription (Who, When, What): The model jointly performs ASR, diarization, and timestamping, producing a structured output that indicates who said what and when.
🌍 Multilingual & Code-Switching Support: It supports over 50 languages, requires no explicit language setting, and natively handles code-switching within and across utterances. Language distribution can be found here.

Evaluation

Installation and Usage

Please refer to GitHub README.

Language Distribution

License

This project is licensed under the MIT License.

Contact

This project was conducted by members of Microsoft Research. We welcome feedback and collaboration from our audience. If you have suggestions, questions, or observe unexpected/offensive behavior in our technology, please contact us at VibeVoice@microsoft.com. If the team receives reports of undesired behavior or identifies issues independently, we will update this repository with appropriate mitigations.

Autor

Microsoft

Organización · ✓

microsoft

Detalles

Descargas637.2K

Me gusta1.2K

AccesoCódigo Abierto

Tareaautomatic-speech-recognition

Parámetros8.7B

Licenciamit

Libreríatransformers

Creado21 ene 2026

Actualizado27 ene 2026

Ver en Hugging Face

Idiomas

enzhesptdejakofrruidsvithenlplnotrtharhucacsdafaafhifietaael

Entiende todo el contexto.

Regístrate para leer casos de estudio completos, acceder a métricas detalladas y recibir todos los reportes.

Autor

Microsoft

Organización · ✓

microsoft

Detalles

Descargas637.2K

Me gusta1.2K

AccesoCódigo Abierto

Tareaautomatic-speech-recognition

Parámetros8.7B

Licenciamit

Libreríatransformers

Creado21 ene 2026

Actualizado27 ene 2026

Ver en Hugging Face

Idiomas

enzhesptdejakofrruidsvithenlplnotrtharhucacsdafaafhifietaael

Entiende todo el contexto.

Regístrate para leer casos de estudio completos, acceder a métricas detalladas y recibir todos los reportes.

Descripción del Modelo

VibeVoice-ASR

🔥 Key Features

Evaluation

Installation and Usage

Language Distribution

License

Contact