🎭 EchoMimicV3

Audio/Text-driven Human Animation Model

Upload a portrait image and an audio file to generate animated talking head video.

Parameter Recommended Range
Audio CFG 2.0 - 3.0 (higher = better lip sync)
Text CFG 3.0 - 6.0 (higher = better prompt following)
Steps 20-25

Requirements: NVIDIA GPU with 24GB+ VRAM (A100 or RTX 4090 recommended)


Model: EchoMimicV3 by Ant Group

Technical Details
  • Parameters: 1.3B
  • Base Model: Wan2.1-Fun-1.3B-InP
  • Audio Encoder: wav2vec2-base-960h
  • License: Apache-2.0