AI & Computational Science

AI converts voices instantly without training on speaker samples

How the science connects

Deep learningGenerative model

AI Insight

MeanVoiceFlow2 is a new voice conversion system that jointly optimizes a flow-based conversion module with an efficient content encoder to achieve fast, one-step zero-shot voice conversion. The model uses conversion distillation from MeanVoiceFlow combined with diffusion-GAN training techniques to improve speech quality and speaker disentanglement. Experimental results demonstrate that MeanVoiceFlow2 achieves approximately 9 times faster inference speed than its predecessor MeanVoiceFlow while maintaining comparable speaker similarity and delivering higher perceptual quality.


This advancement enables real-time voice conversion applications with significantly reduced computational requirements, making zero-shot voice cloning more practical for deployment in resource-constrained environments such as mobile devices or interactive applications. The improved efficiency could expand accessibility of voice conversion technology for assistive communication tools, entertainment, and content creation.


⚠️ Preprint – Noch nicht peer-reviewed

Dieser Artikel wurde noch nicht von unabhängigen Experten begutachtet. Die Ergebnisse sind vorläufig und sollten mit Vorsicht interpretiert werden.

Abstract: Flow-matching approaches to voice conversion (VC) have gained attention owing to their high speech quality and strong speaker similarity. Among them, one-step models such as MeanVoiceFlow are particularly attractive because they enable efficient inference; however, their reliance on a computationally intensive content encoder remains a bottleneck. We therefore propose MeanVoiceFlow2, a framework that jointly optimizes a flow-based conversion module and a computationally efficient content encoder. The model is trained through conversion distillation using MeanVoiceFlow and the reconstruction of real data. We further incorporate diffusion-GAN training with sample mixing and teacher-guided conditioning augmentation to enhance realism and disentanglement. Experiments on zero-shot VC showed that MeanVoiceFlow2 achieved higher perceptual quality and approximately $9times$ faster inference than MeanVoiceFlow while maintaining comparable speaker similarity. Audio samples are available at https://www.kecl.ntt.co.jp/people/kaneko.takuhiro/projects/meanvoiceflow2/.

Source: MeanVoiceFlow2: Joint Optimization of Mean Flow and Content Encoder for Fast One-Step Zero-Shot Voice Conversion