Publication
ICASSP 2024 Speech Signal Improvement Challenge
Microsoft Research Blog
Research Focus: Week of April 1, 2024
In this issue: New research helps COMET embrace African languages; FeatUp improves deep features, a computer vision research cornerstone; LLMs in the Imaginarium: Tool Learning through Simulated Trial and Error; Benchmarking LLMs across languages and…
Publication
Training Audio Captioning Models without Audio
Project
CoVoMix
Advancing Zero-shot Speech Generation for Human-like Multi-talker Conversation We introduce CoVoMix: Conversational Voice Mixture Generation, a novel model for zero-shot, human-like, multi-speaker, multi-round dialogue speech generation. In addition, we devise a comprehensive set of metrics…