Daily Paper Cast

🤗 Upvotes: 30 | cs.SD, cs.CL, cs.MM

Authors:
Ziyang Ma, Zhikang Niu, Wenming Tu, Tianrui Wang, Ruiqi Yan, Junxi Liu, Yanru Huo, Nickk Huang, Yang Liu, Qicong Xie, Zeyu Xie, Hui Wang, Haitao Li, Zixuan Jiang, Yalin Li, Jie Fang, Yifan Duan, Zeyue Tian, Guangzheng Li, Haina Zhu, Shuyi Wang, Jinwen Wang, Mingyu Cui, Tian Tan, Auden, Sen Liang, Steve Yves, Shan Yang, Liefeng Bo, Zilong Zheng, Kai Yu, Eng-Siong Chng, Xie Chen

Title:
AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing

Arxiv:
http://arxiv.org/abs/2609.08936v1

Abstract:
We introduce AuK, an open-source foundational model that unifies speech generation and editing through a common interface of natural-language instructions and audio context. To support this broad capability set, we construct approximately 3.03 billion instruction--audio instances and 1.95 million hours of effective supervision across five task families: speech generation, content editing, enhancement and separation, paralinguistic editing, and acoustic editing. AuK combines a multimodal large language model for semantic conditioning, an VAE jointly trained on speech, general audio, and music for acoustic conditioning, and a hybrid rectified-flow Transformer that performs dual-stream MMDiT blocks followed by unified single-stream DiT blocks for generation. Training begins with generation-only warm-up and proceeds to joint generation--editing pre-training. We then apply complementary post-training strategies: human-feedback preference optimization for open-ended editing and reward-based reinforcement learning for speech generation. To reduce inference cost, we further distill the model with consistency initialization and task-routed Decoupled DMD. The resulting AuK-Flash performs 4-step inference without classifier-free guidance and achieves a 4.5 wall-clock speedup over the full model under matched conditions. Experiments demonstrate leading performance on zero-shot and instruction-controlled speech generation and general instruction-guided editing, while remaining competitive on signal-level restoration tasks. We release both the source code and model weights to support reproducibility and further research.

What is Daily Paper Cast?

We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.com

Creator:
Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/
Gengyu Wang, LLM ML, http://wanggengyu.com

Listen on:
Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXL
Apple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236

Cover Image by Kawen Kuang https://kawen.art

Evan: Welcome to Daily Paper Cast!

Ashley: Today's paper is from the Hugging Face daily paper list of September 9, 2026.

It has received 30 upvotes.

Evan: The paper we're discussing is titled 'AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing'.

Ashley: It's authored by Ziyang Ma and Zhikang Niu, with Xie Chen as the corresponding author.

They are affiliated with Tencent HY.

Evan: So, Ashley, can you give us a bit of background on the development of speech generation systems?

Ashley: Sure, Evan.

Recent advances in speech generation have moved beyond traditional text-to-speech models to encompass capabilities like zero-shot voice cloning, instruction-controlled synthesis, and flexible speech editing.

These capabilities have emerged to meet the complex demands of real-world applications where users might need to synthesize speech with specific styles, replace parts of utterances, or alter attributes like emotion or accent.

Evan: Interesting.

But it sounds like current systems might still face some issues when these capabilities are used together in practical settings.

How does this new model, AuK, address these challenges?

Ashley: Exactly, Evan.

Fragmentation has been a significant obstacle, where using separate models for different tasks can lead to a disjointed user experience and duplicated development efforts.

AuK is designed to unify these capabilities into a single model that can interpret free-form natural-language instructions and generate or edit speech accordingly.

Evan: Got it.

So, unifying these different functions into one system is quite the challenge.

What makes this especially difficult?

Ashley: There are several key challenges.

First, each task has fundamentally different output constraints.

For example, generating new speech from text is very different from editing specific regions of an existing audio clip.

Then, there's the issue of varying conditioning interfaces—some tasks need only text, while others require both text and audio context.

Lastly, the methods for supervision and evaluation differ greatly between tasks, making it hard to create a cohesive training and assessment framework.

Evan: And how exactly does AuK tackle these issues?

Ashley: AuK addresses these challenges by integrating several components.

It uses a multimodal large language model for semantic conditioning, an audio variational autoencoder, or VAE, for acoustic conditioning, and a hybrid rectified-flow transformer to perform dual-stream and single-stream processing for generation and editing tasks.

This architecture allows for flexible adaptation to various tasks through a unified interface.

Evan: That sounds quite robust.

What’s the overall goal or contribution of the AuK model?

Ashley: The primary objective of AuK is to create a single, open-source model that can handle a wide range of speech tasks without requiring task-specific adaptations.

AuK is designed to work with natural-language instructions and audio context to perform tasks such as speech generation, content editing, enhancement and separation, paralinguistic editing, and acoustic editing.

To achieve this, the training process includes both generation and editing pre-training, followed by specialized post-training strategies like human-feedback preference optimization and reward-based reinforcement learning.

Evan: Wow, that's a lot to take in.

So what are some of the claimed benefits and performance metrics detailed in the paper?

Ashley: The paper reports that AuK demonstrates leading performance on zero-shot and instruction-controlled speech generation and general instruction-guided editing, while remaining competitive on signal-level restoration tasks.

The model performs well in terms of intelligibility, speaker similarity, and perceptual quality across a variety of languages and tasks.

Moreover, the release of both the source code and model weights is aimed at supporting further research and reproducibility within the community.

Evan: That wraps up the introduction section.

Evan: Alright, Ashley, let's dive into the methods section.

First off, how does the paper describe the data construction for training the AuK model?

Ashley: The data construction is quite extensive.

The authors organized the pre-training corpus into five task families: speech generation, acoustic editing, paralinguistic editing, content editing, and enhancement and separation.

These tasks share a common interface, which consists of a natural-language instruction, optional input audio, and a target waveform.

Evan: Interesting.

Let’s break that down a bit.

What’s involved in speech generation training?

Ashley: Speech generation training has two parts: transcript-free zero-shot text-to-speech, or TTS, and instruction TTS.

In the zero-shot setting, the model learns to synthesize speech using reference speech samples without transcripts.

This allows the model to clone a speaker’s voice from a short audio clip.

For instruction TTS, the model uses detailed descriptive captions to control the synthesized speech attributes such as emotion, clarity, and timbre, without requiring reference audio.

Evan: Got it.

And what about the acoustic editing task?

Ashley: Acoustic editing focuses on modifying low-level attributes like speaking rate, pitch, and loudness while preserving the linguistic content and speaker identity.

The model is trained using pairs of source and transformed utterances, covering various speaking rates, loudness offsets, and pitch shifts.

Evan: That makes sense.

And paralinguistic editing, what does it entail?

Ashley: Paralinguistic editing involves adjusting attributes that affect how something is spoken, rather than what is being said.

This includes modifying emotions, timbre, accents, and other non-verbal vocalizations.

The training data for these tasks uses expressive speech samples and aligns different attributes with natural-language instructions to guide the model.

Evan: And how do they handle content editing?

Ashley: Content editing allows the model to insert, delete, or replace spoken or sung content while maintaining the original speaker's identity and prosody.

They achieve this by masking specific intervals of the waveform and using a large language model to generate operator-specific annotations and target transcripts.

Evan: Comprehensive indeed.

What about the enhancement and separation tasks?

Ashley: Enhancement and separation tasks teach the model to handle complex acoustic scenes.

This includes removing noise, isolating speakers, or enhancing audio quality based on natural-language instructions.

The training data involves various types of degradations and complex multi-speaker scenarios to prepare the model for diverse real-world applications.

Evan: Alright, shifting to the model design, can you explain the architecture of AuK?

Ashley: Definitely.

The AuK model combines three main components: a multimodal large language model (MLLM) for semantic conditioning, an audio variational autoencoder (VAE) for acoustic conditioning, and a hybrid rectified-flow Transformer for generation and editing.

The MLLM processes the textual instructions and optional audio context to provide a semantic condition.

The VAE is responsible for mapping input audio to a latent space for conditioning and waveform reconstruction.

Evan: And how does the Transformer fit into this architecture?

Ashley: The Transformer backbone is designed with a hybrid approach.

The initial layers are dual-stream MMDiT blocks that process semantic and acoustic streams simultaneously, allowing for interaction between different modalities.

The later layers are single-stream DiT blocks that refine the combined sequence and predict the target latent.

Evan: That’s intricate.

How about the training process?

Ashley: Training is structured into two main stages.

The first stage is a generation-only warm-up that helps the model stabilize on text-to-speech alignment and synthesis quality.

The second stage involves joint optimization of generation and editing tasks using a shared flow-matching objective.

Evan: And what are the post-training strategies they apply?

Ashley: They employ two post-training strategies.

For editing tasks, they collect human feedback on various editing requests and perform flow-based preference optimization.

For speech generation, they use Flow-GRPO, a reinforcement learning approach, with automatic rewards for content correctness, speaker similarity, and instruction-style consistency to refine the model further.

Evan: How do they handle the model’s inference cost?

Ashley: To reduce inference costs, they distill the model with techniques like consistency initialization and task-routed Decoupled DMD.

This results in AuK-Flash, a lightweight version of the model, which performs a 4-step inference without classifier-free guidance, achieving a substantial speedup.

Evan: What’s the performance like for the distilled model compared to the full model?

Ashley: Experiments show that AuK-Flash retains broad generation and editing capabilities while significantly reducing inference time.

It manages to achieve leading performance on several benchmarks for speech generation and general instruction-guided editing, although it slightly trails the full model in some detailed metrics.

Evan: That wraps up the method section.

Evan: Let's dive into the experiments and results.

How did AuK perform in the various evaluations?

Ashley: The performance of AuK was evaluated on multiple tasks, including speech generation, speech editing, and signal-level restoration tasks.

The comprehensive evaluation covers zero-shot and instruction-based speech generation, general instruction-guided speech editing, and various enhancement and separation benchmarks.

Evan: Starting with speech generation, what were the key findings?

Ashley: For speech generation, AuK was assessed using the Seed-TTS-Eval and InstructTTSEval benchmarks.

On Seed-TTS-Eval, which measures linguistic accuracy and speaker preservation, AuK achieved an average word error rate, or WER, of 2.65% and a speaker similarity score, or SIM, of 0.795.

These metrics were collected across English and Chinese datasets, including harder Chinese texts.

Evan: And how did the distilled model, AuK-Flash, fare in comparison?

Ashley: AuK-Flash remained competitive with an average WER of 2.85% and a SIM of 0.790, demonstrating that while it is optimized for speed, it still retains much of the model's performance benefits.

Evan: What about instruction-controlled synthesis?

Ashley: InstructTTSEval focuses on control over acoustic parameters and descriptive styles.

Here, AuK achieved the highest accuracy in Chinese at 83.37%, while for English, AuK-Flash tied for the best score with 82.40%.

These results highlight the model’s capability to follow detailed speech synthesis instructions effectively.

Evan: Shift gears to speech editing.

How did AuK perform in this domain?

Ashley: For speech editing, AuK was evaluated on three benchmarks: MMAE-Speech, SpeechEditBench, and Ming-Freeform-Audio-Edit.

In the MMAE-Speech benchmark, AuK achieved the highest instruction-following rate (IFR) of 48.23% and content retention (CR) of 88.11%.

Meanwhile, AuK-Flash recorded the highest exact-match rate (EMR) of 13.85%.

Evan: And how does it fare in more granular evaluations like SpeechEditBench and Ming-Freeform-Audio-Edit?

Ashley: AuK topped the metrics for content, emotion, prosody, and acoustic editing in the SpeechEditBench, significantly outpacing the other models.

In the more detailed Ming-Freeform-Audio-Edit, it led in semantic editing, achieving the lowest word error rate across both Chinese and English sets, and demonstrated superior performance in insertion, deletion, and substitution tasks.

Evan: That’s quite comprehensive.

What are the results in the enhancement and separation tasks?

Ashley: In DNS Challenge 2020 for speech enhancement, AuK-Flash achieved the highest perceived quality score (UTMOS) of 4.05 with the AuK model leading in dWER at 2.66%.

Over at CHiME-4, AuK-Flash achieved the lowest WER at 7.84%.

In speaker separation tasks like Libri2Mix, both AuK variants achieved top speaker similarity scores, with AuK-Flash excelling in overall perceptual quality.

Evan: And how is it with speech super-resolution tasks?

Ashley: For VCTK-SR, AuK-Flash achieved the best overall scores in perceptual quality and recognition accuracy with an overall UTMOS score of 4.05 and a WER of 2.92%.

AuK performed better in preserving speaker identity with a SIM score of 0.97.

Evan: Any additional insights or noteworthy observations from their experiments?

Ashley: Yes, they observed emergent capabilities such as cross-task and cross-lingual transfer.

For instance, the model could transfer learned transformations like whispering from editing tasks to generation tasks, even performing well on de-accenting tasks for English, even though the training primarily focused on Chinese dialects.

Evan: That covers the results and findings.

That’s the end of the Experiment section.

Evan: Now let's delve into the related work section.

Ashley, how does AuK fit into the existing landscape of speech generation and editing models?

Ashley: AuK builds on several advancements across different areas of speech processing.

Traditional text-to-speech models have evolved significantly, starting with models focused on high-quality speech synthesis, like Seed-TTS, which aimed at versatility and high-quality generation.

Evan: And how has the field moved towards more advanced capabilities, such as zero-shot voice cloning and instruction-controlled synthesis?

Ashley: The field has seen innovations like zero-shot voice cloning, where models like CosyVoice and VoxCPM can clone a voice from just a few seconds of audio.

InstructTTS models, like those reported in the InstructTTSEval benchmark, are designed to generate speech based on descriptive text instructions, offering fine-grained control over attributes like emotion and prosody.

Evan: What about speech editing capabilities?

Ashley: For speech editing, models like CosyEdit and Audio-Omni have set benchmarks.

They can modify specific aspects of audio while preserving the majority of the original content.

This includes tasks like changing tone, inserting pauses, or altering vocal expression.

These models typically rely on datasets that pair original audio with modified versions to learn the changes required.

Evan: How has unified modeling been approached in the past?

Ashley: Unified models have been a topic of increasing interest.

Ming-UniAudio, for instance, attempted to bring multiple audio processing tasks under a single framework.

This model could process diverse audio tasks like enhancement, separation, and even cross-modal tasks such as adding background music.

However, these models sometimes struggle with task-specific optimizations, reducing the effectiveness of their outputs.

Evan: Does AuK leverage any advancements from non-audio-specific models?

Ashley: AuK incorporates techniques from large language models used in natural language processing.

Models like GPT-3 and its successors, which utilize transformers for in-context learning and conditioning, have influenced AuK’s architecture in handling multimodal inputs for semantic conditioning.

Evan: What kind of influence has the development of audio VAEs had on models like AuK?

Ashley: Audio variational autoencoders (VAEs) have significantly influenced models like AuK by providing a structured latent space that captures the nuances of the audio signal.

Previous iterations like BigVGAN have demonstrated that you can achieve high-quality audio reconstruction.

These principles are applied in AuK’s acoustic conditioning, ensuring the generated or edited speech retains high fidelity.

Evan: What are some of the unique challenges that AuK aims to solve that prior models didn’t fully address?

Ashley: AuK aims to more cohesively integrate generation and editing tasks within a single model, ensuring scalability and efficiency in both training and inference.

Previous models often required different systems for generation versus editing, leading to inefficiencies and fragmented user experiences.

By unifying these tasks, AuK can provide a smoother and more consistent user experience while reducing the overall computational overhead.

Evan: I see.

How does this compare to approaches in general AI research?

Ashley: In general AI research, especially in areas like computer vision and natural language processing, there's been a significant push towards creating versatile, unified models that can handle a wide array of tasks.

For instance, general-purpose transformers have been used successfully in both fields to achieve state-of-the-art results across multiple benchmarks.

AuK follows a similar philosophy by leveraging large-scale, multimodal training to achieve robustness and versatility in speech tasks.

Evan: Thanks for the comprehensive rundown, Ashley.

That concludes the Related Work section.

Evan: Alright, Ashley, let's wrap this up.

What are the key contributions and takeaways from the AuK paper?

Ashley: Sure, Evan.

The key contribution of the AuK model is its ability to unify speech generation and editing capabilities within a single, open-source system.

This is a significant step forward from previous models that often handled these tasks separately.

AuK leverages a multimodal large language model for semantic conditioning and an audio VAE for acoustic conditioning, integrated through a hybrid rectified-flow Transformer.

These components allow the model to handle a diverse range of tasks efficiently.

Evan: And the extensive training process seems to be a big part of why it's effective, right?

Ashley: Exactly.

The training process includes both pre-training and post-training strategies tailored to the unique requirements of generation and editing.

They utilize human feedback for editing optimization and reinforcement learning for generation refinement, ensuring high performance across various tasks.

Evan: It’s fascinating how they managed to maintain high-quality results while also improving inference speed with AuK-Flash.

Ashley: AuK-Flash retains the broad capabilities of the full model but significantly reduces inference time, making it more efficient without compromising much on performance.

It’s a great example of balancing capability with efficiency.

Evan: Well, that's it for today's episode.

Thanks for a great discussion, Ashley.

Ashley: Thank you, Evan.

And thank you to our listeners for tuning in.

Evan: Don't forget to join us again next time for another deep dive into the latest AI research.

Be sure to subscribe to Daily Paper Cast to stay up-to-date with cutting-edge developments.

Until next time!