Awesome Multi-Speaker ASR
A curated list of papers, datasets, challenges, and tools for multi-speaker automatic speech recognition (ASR), especially overlapped speech recognition, speaker-attributed ASR, target-speaker ASR, and meeting transcription. Feel free to contribute!
Contents
Reviews & Surveys
2025
- (Arxiv 2025) Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio. [Paper]
End-to-End Speaker Diarization and Recognition
A unified model jointly learns speaker diarization and speech recognition to produce who spoke when and what, without combining independently trained speaker diarization and ASR systems at inference time.
2026
- (ArXiv 2026) VibeVoice-ASR-Streaming Technical Report. [Paper] [Code]
- (AAAI 2026) SpeakerLM: End-to-End Versatile Speaker Diarization and Recognition with Multimodal Large Language Models. [Paper]
- (ArXiv 2026) VIBEVOICE-ASR Technical Report. [Paper] [Code]
- (ArXiv 2026) TagSpeech: End-to-End Multi-Speaker ASR and Diarization with Fine-Grained Temporal Grounding. [Paper]
- (ArXiv 2026) MOSS Transcribe Diarize: Accurate Transcription with Speaker Diarization. [Paper]
2025
- (ArXiv 2025) Train Short, Infer Long: Speech-LLM Enables Zero-Shot Streamable Joint ASR and Diarization on Long Audio (JEDIS-LLM). [Paper]
- (ICML 2025) Sortformer: Seamless Integration of Speaker Diarization and ASR by Bridging Timestamps and Tokens. [Paper] [Code]
Earlier
- (ICASSP 2024) One Model to Rule Them All? Towards End-to-End Joint Speaker Diarization and Speech Recognition. [Paper]
- (ArXiv 2024) SA-SOT: Speaker-Aware Serialized Output Training for Multi-Talker ASR. [Paper]
- (Odyssey 2024) On Speaker Attribution with SURT. [Paper]
- (ICASSP 2024) Speaker Mask Transformer for Multi-Talker Overlapped Speech Recognition. [Paper]
- (ICASSP 2022) Transcribe-to-Diarize: Neural Speaker Diarization for Unlimited Number of Speakers Using End-to-End Speaker-Attributed ASR. [Paper]
- (INTERSPEECH 2021) End-to-End Speaker-Attributed ASR with Transformer. [Paper]
- (INTERSPEECH 2020) Joint Speaker Counting, Speech Recognition, and Speaker Identification for Overlapped Speech of Any Number of Speakers. [Paper]
Cascaded Speaker Diarization and Recognition
Speaker diarization and speech recognition are performed by separately trained modules, with timestamps used to associate transcripts with speakers. Speech separation may be added before diarization or ASR; some systems also use an LLM for correction, refinement, or diarization-guided recognition.
2026
- (ArXiv 2026) Mitigating Speaker Leakage in Cascaded Multi-talker ASR with Diarization-based Transcript Correction. [Paper]
- (ArXiv 2026) DM-ASR: Diarization-aware Multi-speaker ASR with Large Language Models. [Paper]
2025
- (INTERSPEECH 2025) Overlap-Adaptive Hybrid Speaker Diarization and ASR-Aware Observation Addition for MISP 2025 Challenge. [Paper]
- (Speech Communication 2025) LLM-Based Speaker Diarization Correction: A Generalizable Approach. [Paper] [Code]
- (ArXiv 2025) Diarization-Aware Multi-Speaker Automatic Speech Recognition via Large Language Models. [Paper]
Earlier
- (INTERSPEECH 2024) DiarizationLM: Speaker Diarization Post-Processing with Large Language Models. [Paper] [Code]
- (ArXiv 2023) AutoPrep: An Automatic Preprocessing Framework for In-the-Wild Speech Data. [Paper]
- (INTERSPEECH 2023) WhisperX: Time-Accurate Speech Transcription of Long-Form Audio. [Paper] [Code]
- (INTERSPEECH 2022) Speaker Conditioned Acoustic Modeling for Multi-Speaker Conversational ASR. [Paper]
- (SLT 2021) Integration of Speech Separation, Diarization, and Recognition for Multi-Speaker Meetings: System Description, Comparison, and Analysis. [Paper]
End-to-End Speaker-Attributed ASR
These systems directly produce who said what but do not necessarily output a complete speaker diarization timeline. They include serialized-output, target-speaker, multi-output, and separation-aware models.
2026
- (ArXiv 2026) SE-DiCoW: Self-Enrolled Diarization-Conditioned Whisper. [Paper]
2025
- (ASRU 2025) Serialized Output Prompting for Large Language Model-Based Multi-Talker Speech Recognition. [Paper] [Code]
- (ICASSP 2025) Disentangling Speakers in Multi-Talker Speech Recognition with Speaker-Aware CTC. [Paper] [Code]
- (ICASSP 2025) DiCoW: Diarization-Conditioned Whisper for Target Speaker Automatic Speech Recognition. [Paper] [Code]
Earlier
- (ArXiv 2024) Alignment-Free Training for Transducer-based Multi-Talker ASR. [Paper]
- (ArXiv 2024) Target Speaker ASR with Whisper. [Paper]
- (ArXiv 2024) Large Language Model Can Transcribe Speech in Multi-Talker Scenarios with Versatile Instructions. [Paper]
- (SLT 2024) Serialized Speech Information Guidance with Overlapped Encoding Separation for Multi-Speaker Automatic Speech Recognition. [Paper] [Code]
- (ArXiv 2024) Advancing Multi-Talker ASR Performance with Large Language Models. [Paper]
- (ArXiv 2024) Empowering Whisper as a Joint Multi-Talker and Target-Talker Speech Recognition System. [Paper]
- (ICASSP 2024) Cross-Speaker Encoding Network for Multi-Talker Speech Recognition. [Paper] [Code]
- (INTERSPEECH 2023) BA-SOT: Boundary-Aware Serialized Output Training for Multi-Talker ASR. [Paper]
- (ICASSP 2023) A Sidecar Separator Can Convert a Single-Talker Speech Recognition System to a Multi-Talker One. [Paper]
- (INTERSPEECH 2022) Streaming Multi-Talker ASR with Token-Level Serialized Output Training. [Paper]
- (INTERSPEECH 2020) Multi-Talker ASR for an Unknown Number of Sources: Joint Training of Source Counting, Separation and ASR. [Paper]
- (INTERSPEECH 2020) End-to-End Multi-Speaker Speech Recognition with Transformer. [Paper]
- (INTERSPEECH 2020) Serialized Output Training for End-to-End Overlapped Speech Recognition. [Paper] [Data]
- (INTERSPEECH 2019) End-to-End Multi-Speaker Speech Recognition Using Speaker Embeddings and Transfer Learning. [Paper]
- (ACL 2018) A Purely End-to-End System for Multi-Speaker Speech Recognition. [Paper]
- (ICASSP 2018) End-to-End Multi-Speaker Speech Recognition. [Paper]
- (ICASSP 2017) Recognizing Multi-Talker Speech with Permutation Invariant Training. [Paper]
Cascaded Speaker-Attributed ASR
These systems combine separately optimized recognition and speaker-attribution components, without requiring a complete speaker diarization timeline in the final output.
Earlier
- (INTERSPEECH 2024) SOT Triggered Neural Clustering for Speaker Attributed ASR. [Paper]
- (ArXiv 2022) A Comparative Study on Speaker-Attributed Automatic Speech Recognition in Multi-Party Meetings. [Paper]
Datasets & Challenges
| Dataset / challenge |
Year |
Language |
Setting |
Paper / homepage |
| NOTSOFAR-1 |
2024 |
English |
Real and simulated distant office meetings |
Paper · Homepage |
| CHiME-8 DASR |
2024 |
Multidomain |
Array-agnostic distant ASR and diarization |
Paper · Homepage |
| AliMeeting |
2022 |
Mandarin |
Real meetings, near- and far-field arrays |
Paper · Data |
| AISHELL-4 |
2021 |
Mandarin |
Real meeting speech, eight-channel array |
Paper · Data |
| CHiME-6 |
2020 |
English |
Real dinner-party conversations, multi-array |
Paper · Homepage |
| LibriCSS |
2020 |
English |
Replay-recorded, continuous and partially overlapped |
Paper · Data |
| LibriSpeechMix |
2020 |
English |
Simulated one-, two-, and three-speaker mixtures |
Data |
| LibriMix |
2020 |
English |
Simulated two-/three-speaker, clean/noisy mixtures |
Paper · Data |
| WSJ0-2mix / WSJ0-3mix |
2016 |
English |
Simulated, fully overlapped mixtures |
Deep Clustering |
| AMI Meeting Corpus |
2005 |
English |
Real meetings, close- and far-field microphones |
Paper · Homepage |
Evaluation & Toolkits
- MeetEval: meeting transcription metrics including cpWER, ORC-WER, MIMO-WER, and time-constrained variants. [Paper]
- ESPnet: end-to-end speech processing toolkit with recipes for meeting and overlapped speech recognition.
- Kaldi LibriCSS recipe: diarization and ASR baselines for LibriCSS.
- NVIDIA NeMo: speech recognition and diarization toolkit containing speaker-aware and Sortformer-related components.
- icefall SURT recipe: reproducible streaming unmixing and recognition transducer recipe.
Common metrics:
- WER / CER: content recognition error, usually without speaker attribution.
- cpWER: concatenated minimum-permutation WER; concatenates each speaker’s words and finds the best reference/hypothesis speaker assignment.
- ORC-WER: optimal reference combination WER; useful when the number of output streams differs from the number of reference speakers.
- MIMO-WER: evaluates flexible multi-input multi-output meeting transcription while preserving utterance order.
- tcpWER: time-constrained minimum-permutation WER; additionally checks whether matched words occur at plausible times.
- SA-WER / speaker-dependent WER: jointly reflects transcription and speaker-attribution errors; definitions can differ across papers.
Contributing
Contributions are welcome. Please open an issue or pull request and use the following format:
- (VENUE YEAR) **Paper Title.** [[Paper](PAPER_URL)] [[Code](CODE_URL)] [[Data](DATA_URL)]
Please place a paper in its primary category, use the official paper/project link when possible, and keep entries in reverse chronological order.