Human4DiT: 360-degree Human Video Generation with 4D Diffusion Transformer

编辑：广东客 | 分类：CV / XR | 2024年11月12日

Note: We don't have the ability to review paper

PubDate: May 2024

Teams: Tsinghua University

Writers: Ruizhi Shao, Youxin Pang, Zerong Zheng, Jingxiang Sun, Yebin Liu

PDF: Human4DiT: 360-degree Human Video Generation with 4D Diffusion Transformer

Abstract

We present a novel approach for generating 360-degree high-quality, spatio-temporally coherent human videos from a single image. Our framework combines the strengths of diffusion transformers for capturing global correlations across viewpoints and time, and CNNs for accurate condition injection. The core is a hierarchical 4D transformer architecture that factorizes self-attention across views, time steps, and spatial dimensions, enabling efficient modeling of the 4D space. Precise conditioning is achieved by injecting human identity, camera parameters, and temporal signals into the respective transformers. To train this model, we collect a multi-dimensional dataset spanning images, videos, multi-view data, and limited 4D footage, along with a tailored multi-dimensional training strategy. Our approach overcomes the limitations of previous methods based on generative adversarial networks or vanilla diffusion models, which struggle with complex motions, viewpoint changes, and generalization. Through extensive experiments, we demonstrate our method’s ability to synthesize 360-degree realistic, coherent human motion videos, paving the way for advanced multimedia applications in areas such as virtual reality and animation.

本文链接：https://paper.nweon.com/16092

Human4DiT: 360-degree Human Video Generation with 4D Diffusion Transformer

您可能还喜欢...

最新AR/VR行业分享

最新AR/VR专利

最新AR/VR行业招聘

Human4DiT: 360-degree Human Video Generation with 4D Diffusion Transformer

您可能还喜欢...

An advanced QER selection algorithm for 360-degree VR video streaming

Real-time Cross-modal Cybersickness Prediction in Virtual Reality

A Real-Time Markerless Augmented Reality Framework Based on SLAM Technique

最新AR/VR行业分享

最新AR/VR专利

最新AR/VR行业招聘