HuMo
An open-source framework for generating human-centric videos conditioned on multimodal inputs including text, images, and audio with precise subject preservation and audio-visual synchronization.
Overview
HuMo is a state-of-the-art human-centric video generation framework that synthesizes videos of humans by collaboratively conditioning on multiple modalities: text prompts, reference images, and audio. It addresses key challenges in multimodal video synthesis such as maintaining subject identity and synchronizing audio with visual facial movements. HuMo introduces a two-stage progressive training paradigm with task-specific strategies, including minimal-invasive image injection for subject preservation and a novel focus-by-predicting mechanism to enhance audio-visual synchronization. The framework also features a time-adaptive Classifier-Free Guidance strategy during inference to allow flexible and fine-grained multimodal control. HuMo surpasses previous specialized methods by unifying these capabilities into a single collaborative multimodal conditioning model.