Daily Paper Cast

🤗 Upvotes: 121 | cs.CV, cs.LG

Authors:
Jintao Zhang, Kai Jiang, Jintao Chen, Xu Wang, Deyuan Liu, Jungang Li, Dechuang Chen, Ming Lin, Jingjiang Zhou, Haopeng Jin, Qi Jia, Xiaohang Wang, Yaole Wang, Zhanqiang Zhang, Ran Li, Zhengkun Huang, Shuyue Xiong, Yuji Wang, Zikun Dai, Hui He, Yang Luo, Mang Ning, Weiqi Feng, Chengyang Ye, Xinyue Lin, Min Zhao, Hongzhou Zhu, Hengkai Tan, Zeyuan Wang, Chendong Xiang, Kaiwen Zheng, Zhijie Deng, Fan Bao, Jianfei Chen, Jun Zhu

Title:
Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation

Arxiv:
http://arxiv.org/abs/2609.11638v1

Abstract:
We present Vidu S2, which comprises Vidu S2-Avatar, a real-time interactive digital-character model, and Vidu S2-Editing, a real-time video editing model. Moreover, we explore the feasibility of real-time spatial video generation for both Vidu S2-Avatar and Vidu S2-Editing. Compared with Vidu S1, Vidu S2-Avatar supports real-time 720p video generation, generation with dynamic references that can be updated at any moment, and stronger instruction following, such as dancing. Vidu S2-Editing supports editing a video stream in real time, including style rendering, clothing replacement, character replacement, and background replacement. Experiments show that Vidu S2 outperforms all baselines. A playable online demo is available at https://vidu.com/vidu-stream.

What is Daily Paper Cast?

We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: dailypapercast.ai@gmail.com

Creator:
Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/
Gengyu Wang, LLM ML, http://wanggengyu.com

Listen on:
Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXL
Apple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236

Cover Image by Kawen Kuang https://kawen.art

Evan: Welcome to Daily Paper Cast.

Ashley: Today, we're diving into a paper from Hugging Face's daily paper list from September 15, 2026, which has received 121 upvotes.

Evan: The title of the paper is 'Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation.'

Ashley: This paper comes from Tsinghua University and Shengshu Technology, with lead authors Jintao Zhang and Kai Jiang, and the corresponding author is Zhijie Deng.

Evan: Alright Ashley, let’s dive right into the introduction.

What’s the main demand they're addressing with Vidu S2?

Ashley: Recent video generation models, such as Sora and Veo, have shown impressive capabilities in generating high-quality videos.

However, these models generally follow an offline, one-shot generation paradigm.

Users have to input a prompt, wait for minutes or even tens of minutes, and receive the complete video only after the generation finishes.

Evan: So, the user can’t interact with the video generation process in real-time?

Ashley: Exactly.

During this offline process, the user passively waits, unable to interact or adjust the video generation dynamically.

This works for pre-generated content but falls short for interactive visual experiences like live streaming, face-to-face communication, and gaming.

Evan: So how big is the demand for real-time interactive video generation compared to offline-generated content?

Ashley: It's significantly higher.

The authors use a hypothetical example: If each user has an average demand of around 50% for both real-time interactive and offline-generated content, the generation demand for real-time content would be much greater due to its unique interactive quality.

Evan: Now, how does Vidu S2 aim to address this gap?

Ashley: Vidu S2 is built on the foundation of its predecessor, Vidu S1, and it moves forward in two significant directions: Vidu S2-Avatar and Vidu S2-Editing.

Vidu S2-Avatar enhances real-time interactive digital characters to support higher resolution, dynamic references, and stronger instruction following, even for complex actions like dancing.

Evan: And what about Vidu S2-Editing?

Ashley: Vidu S2-Editing can edit video streams in real-time, including style transfer, virtual try-on, character replacement, and background replacement.

Essentially, it brings a new level of interactivity and editing flexibility to streaming video.

Evan: That sounds like a big step forward.

Did they include any new experiments or benchmarks?

Ashley: Yes, they did.

The authors have conducted experiments showing that Vidu S2 outperforms all existing baselines.

Importantly, there's also an online demo available for users to experience these capabilities firsthand.

Evan: Interesting.

It sounds like they're going beyond just improving algorithms to making the whole user experience more dynamic and interactive.

Ashley: Precisely.

They are not only advancing the technical aspects but also ensuring that the technology can be experienced in practical, real-world scenarios, making it more accessible and user-friendly for interactive applications.

Evan: Well, that wraps up our deep dive into the introduction of this paper.

Ashley: Stay tuned as we get into the methodologies and experiments in a bit.

Evan: Welcome back, everyone.

Let's jump into the methods behind Vidu S2.

Ashley: Alright, Evan.

The methods are detailed and multifaceted, covering the Vidu S2-Avatar and Vidu S2-Editing components.

Let's start with Vidu S2-Avatar, the real-time interactive digital-character model.

Evan: Sounds good.

How does Vidu S2-Avatar improve upon its predecessor, Vidu S1?

Ashley: Vidu S2-Avatar makes four key advancements over Vidu S1.

First, it introduces Self-Replay Forcing, a novel method to improve the quality and data efficiency by replaying the generated trajectory in a single gradient-enabled causal pass.

This allows gradients to propagate across segments while avoiding the issue of history detachment.

Evan: Interesting.

And the second advancement?

Ashley: The second improvement is the upgrade from 540p resolution to 720p, keeping the frame rate between 25 and 42 frames per second.

A lightweight Refiner component is used to upscale the video in real-time, providing clearer and more detailed digital characters.

Evan: What about the other two advancements?

Ashley: Third, the instruction-following capability has been greatly enhanced.

Vidu S2-Avatar can now follow a wider range of instructions, including complex actions like dancing.

This is achieved by incorporating more diverse training data, including solo dance videos and 2D/3D animations.

Ashley: Finally, dynamic reference interaction has been introduced.

Users can now update reference images at any point during the video stream.

For instance, the digital character can pick up a new object, change clothing, or switch backgrounds seamlessly.

Evan: So, it sounds like Vidu S2-Avatar not only enhances video quality but also the interactive experience.

How did they handle the data preparation for this?

Ashley: Great question.

The data preparation pipeline for Vidu S2-Avatar is comprehensive.

Videos from diverse sources, such as talking heads, film, TV, solo dance, and animations, are processed through several stages: clipping, filtering, speech processing, captioning, and embedding.

Additionally, they have enhanced operators for high-clarity selection and background stabilization to ensure superior training data quality.

Evan: And how do they filter and select these high-clarity videos?

Ashley: They use a multidimensional hybrid selection framework to evaluate and select videos not just based on resolution but also on frame rates, codec, pixel format, bit depth, and bitrate.

This ensures that the selected videos are of the highest clarity suitable for 720p generation.

Evan: That explains a lot.

So, what about Vidu S2-Editing?

Ashley: Vidu S2-Editing focuses on real-time video editing capabilities.

It can perform style transfers, virtual try-ons, character replacements, and background replacements, all in real-time.

Each of these functions utilizes frame-aligned attention to maintain the exact motion and timing of the original video, ensuring coherence and fluidity in the edited output.

Evan: That sounds impressive.

How do they prepare the data for these editing tasks?

Ashley: For data preparation, they filtered videos through the Vidu S2-Avatar pipeline but applied stricter criteria to ensure high quality.

They constructed four disjoint subsets for the different editing tasks, each comprising 200,000 videos.

The goal was to cover a wide range of visual styles and content for training.

Evan: And how do they handle style transfer specifically?

Ashley: For style transfer, they first reconstruct a video from its surface-normal representation and a reference image using a model like NormalCrafter.

They then replace the reference image with a stylized one to create a consistent stylized video that mimics the spatial structure and motion of the original.

Evan: It seems like both components aim to offer robust real-time interaction and editing capabilities.

Ashley: Exactly.

Vidu S2 also explores the feasibility of real-time spatial video generation where the output can be experienced in virtual reality (VR).

For instance, Vidu S2-Avatar can generate synchronized left- and right-eye views, enhancing immersive VR experiences.

Evan: That's quite forward-thinking.

How do they ensure low-latency during these complex processes?

Ashley: They employ several optimization strategies like efficient attention mechanisms, kernel fusion, launch optimization, and multi-GPU parallelism.

For instance, quantized linear-layer accelerations and custom CUDA kernels reduce memory usage and computational load, ensuring real-time, low-latency performance.

Evan: Sounds like they've covered all bases to maintain performance without compromising on quality.

Ashley: Precisely.

The proposed methods in the paper are designed to ensure that real-time video generation and editing are not only feasible but also efficient and high quality, making interactive visual experiences more accessible.

Evan: Well, that wraps up the Method section.

Stay tuned as we dive into the experiments and results next.

Evan: Welcome back!

Let's delve into the experiments and the results they achieved with Vidu S2.

Ashley: Sure, Evan.

They evaluated Vidu S2 on two main tasks: streaming digital character generation with Vidu S2-Avatar and real-time video editing with Vidu S2-Editing.

Evan: Alright, let's start with Vidu S2-Avatar.

How did they set up the experiments for digital character generation?

Ashley: For digital character generation, they used the StreamAV-Bench, a comprehensive benchmark suite that validates several aspects of streaming audio-video performance.

They also ran an internal benchmark to compare Vidu S2-Avatar against commercial systems.

Evan: What specific metrics did they measure?

Ashley: The metrics included Visual Aesthetics, Visual Quality, Production Quality, Audio Quality, Audio–Visual Alignment, Audio–Visual Synchronization, Audio Instruction Fulfillment, Subject Consistency, and Background Consistency.

Each metric evaluates different attributes such as visual appeal, motion naturalness, audio fidelity, and consistency over long durations.

Evan: And how did Vidu S2-Avatar perform on these benchmarks?

Ashley: Vidu S2-Avatar outperformed all other systems across every reported metric.

For example, it scored highest in Visual Quality and Production Quality, indicating that its visual and audio output were top-notch.

Notably, it also excelled in Audio–Visual Synchronization and Instruction Fulfillment, which are crucial for interactive applications.

Evan: That’s quite comprehensive!

How did they ensure the results were reliable?

Ashley: They employed standardized protocols for the benchmarks and used aggregated results from multiple evaluators to ensure consistency.

Every system received the same inputs and followed identical evaluation procedures to maintain fairness.

Evan: What about the results from their internal benchmarks?

Ashley: Internally, they conducted duration-stratified evaluations to see how Vidu S2-Avatar performs over time, from short clips of 10 seconds to longer streams of up to 90 seconds.

Vidu S2-Avatar consistently maintained high ratings across all durations, particularly excelling in long-horizon stability and temporal consistency.

Evan: Fascinating.

Let’s now talk about the video editing task.

How did they set up the experiments for Vidu S2-Editing?

Ashley: For Vidu S2-Editing, they used a variety of public benchmarks including OpenVE-Bench, Sparkle-Bench, and RefVIE-Bench.

These benchmarks evaluate instruction-guided video editing, reference-conditioned editing, and unpaired virtual try-ons.

They also constructed an internal long-horizon test set.

Evan: What metrics were used for evaluating video editing?

Ashley: The metrics included Global Style for overall appearance, Background Change, Instruction Compliance, Visual Quality, Foreground Motion Preservation, and other specific criteria like Reference Fidelity and Temporal Consistency.

Evan: And how did Vidu S2-Editing fare in these evaluations?

Ashley: Vidu S2-Editing achieved the highest scores in almost all evaluated metrics.

For instance, it scored 4.71 in Global Style and 4.14 in Background Change on OpenVE-Bench, surpassing even the strongest offline models like Bernini-R 14B.

It also excelled in maintaining coherence and fluidity between edited and original video sequences.

Evan: What about the results from the internal tests?

Ashley: In the internal tests, Vidu S2-Editing was preferred by evaluators for its overall quality, video quality, temporal consistency, and semantic adherence.

It outperformed other commercial systems like Decart-Lucy2.5 and XMax-X2.0.

Evan: So, it seems Vidu S2-Editing was able to seamlessly integrate new elements while maintaining the integrity and quality of the original video.

Ashley: Exactly.

The experiments highlight Vidu S2’s capability to not only generate but also edit videos in real-time, all while maintaining high visual and audio fidelity.

Evan: It’s impressive to see how comprehensive and robust their evaluation process was.

Anything else notable from the experiments?

Ashley: Yes, they also demonstrated Vidu S2's capability in generating spatial videos for VR headsets.

They evaluated this by producing synchronized left- and right-eye views, enhancing the immersive experience without introducing noticeable latency.

Evan: That's quite forward-thinking.

It's clear they've thought through both the technical details and the practical applications.

Ashley: The experiments validate that Vidu S2 is ready for real-world applications in areas like live streaming, virtual events, and interactive entertainment.

Evan: Well, that wraps up the Experiment section.

Stay with us as we continue to explore other aspects of this intriguing paper.

Evan: Now, let's move onto the related work that Vidu S2 builds upon.

Ashley, could you kick things off for us?

Ashley: Of course, Evan.

The related work section of the paper provides a comprehensive overview of the landscape of video generation and editing technologies.

They discuss several key models and frameworks that paved the way for Vidu S2.

Evan: What are some of the primary models they compare Vidu S2 to?

Ashley: They start by mentioning recent video generation models like Sora, Veo, Wan, and Seedance.

These models have shown strong capabilities in generating high-quality videos but typically follow an offline, one-shot paradigm.

Evan: What are the limitations of these traditional models?

Ashley: As I mentioned earlier, the offline generation paradigm requires users to input a prompt and wait for the entire video to be generated, which can take a significant amount of time.

This process doesn't allow for real-time interaction or adjustments.

Evan: Got it.

So, how did the researchers aim to improve this with the Vidu series?

Ashley: Vidu S1, the predecessor of Vidu S2, was one of the earliest attempts to provide real-time interactive video generation.

It utilized TurboDiffusion and TurboServe to enable continuous user interaction, but had limitations in resolution and instruction-following capabilities.

Evan: And how does Vidu S2 improve on those limitations?

Ashley: Vidu S2 takes several steps forward by enhancing resolution to 720p, adding dynamic reference capabilities, and improving instruction adherence, even for complex actions like dancing.

These advancements are built on the foundations laid by TurboDiffusion and TurboServe.

Evan: That's interesting.

What other models are mentioned in the related work?

Ashley: The paper also references the Self-Forcing model, which addresses the train-test gap in autoregressive video diffusion.

This model conditions each video segment on chunks it generated itself, although it still faced issues with efficiency and quality due to a lack of gradient flow through the self-generated history.

Evan: How does Vidu S2 address these concerns?

Ashley: Vidu S2 introduces Self-Replay Forcing, which re-noises and replays the student-generated trajectory in a single gradient-enabled pass.

This ensures that gradients can flow across segments, improving data efficiency and training quality.

Evan: What about the real-time video editing capabilities?

Are there any notable related works there?

Ashley: Yes, the paper discusses various existing video editing models like Bernini, SCAIL-2, and CoinVE-Edit.

These models offer robust editing capabilities but often operate offline and can't handle the real-time demands that Vidu S2-Editing addresses.

Evan: So, what makes Vidu S2-Editing stand out in the landscape of video editing models?

Ashley: Vidu S2-Editing distinguishes itself by enabling real-time edits with frame-aligned attention.

This ensures that the edits maintain the same motion and timing as the original video, making the result appear seamless and natural.

Evan: It sounds like these contributions are addressing some significant gaps in the field.

Is there anything else notable in the related work about how Vidu S2 compares to competitors?

Ashley: Indeed, the paper also explores real-time spatial video generation and editing aimed at enhancing VR experiences.

This is something not extensively covered in earlier models and shows Vidu S2's forward-thinking approach.

Evan: It's fascinating to see how these models evolved and where Vidu S2 fits within the broader context of video generation and editing technologies.

Ashley: The related work effectively highlights the advancements made by Vidu S2 while acknowledging the foundational work of previous models in the field.

Evan: Alright, that brings us to the end of the related work section.

Stay with us as we delve into more details of this fascinating paper.

Evan: Welcome back!

Let's summarize the key contributions and takeaways from the Vidu S2 paper.

Ashley: To recap, Vidu S2, which includes Vidu S2-Avatar and Vidu S2-Editing, makes significant advancements in real-time interactive video generation and editing.

Evan: For Vidu S2-Avatar, the improvements focus on enhancing video quality to 720p, supporting dynamic references, and improving the instruction-following capabilities for complex actions like dancing.

Ashley: And for Vidu S2-Editing, the emphasis is on real-time video editing capabilities.

This includes style transfer, virtual try-on, character replacement, and background replacement, all performed seamlessly with frame-aligned attention to maintain coherence and fluidity.

Evan: The data preparation and optimization methods are thorough, ensuring high clarity and stability of videos.

They used diverse video sources and sophisticated selection frameworks to gather the best training data.

Ashley: Their experiments demonstrated that Vidu S2 surpasses existing baselines in all evaluated metrics for both digital character generation and video editing.

Notably, they showed impressive results in visual quality, motion naturalness, and temporal consistency.

Evan: Exactly.

The work also explores the potential for real-time spatial video generation, making significant strides towards enhancing immersive VR experiences.

Ashley: Overall, Vidu S2 is not just an incremental improvement.

It represents a significant leap in both the functionality and usability of real-time video generation and editing technologies, making interactive visual experiences more accessible and practical.

Evan: That’s all for today's discussion on 'Daily Paper Cast.' Thanks for tuning in!

Ashley: We hope you enjoyed our deep dive into Vidu S2.

Make sure to join us in future episodes as we continue to explore the latest research in AI, NLP, and more.

Evan: Until next time, keep questioning, keep learning, and stay curious.

Goodbye!

Ashley: Goodbye, everyone!