Choose image motion or an audio-driven avatar
Hosted image-to-video animates a still image without audio input. For speech-driven lip sync, follow the local Avatar setup below.
LongCat-Video-Avatar 1.5 is Meituan's current audio-driven avatar release. It is not the same checkpoint or audio stack as the December 2025 version: v1.5 replaces Wav2Vec2 with Whisper-Large, supports single- and multi-stream audio, and uses an 8-step distilled inference path. The official tasks are audio + text to video, audio + text + image to video, and video continuation.
For an overview of the separate audio-driven model, see the LongCat Video Avatar guide. That page documents local inference; it is not a hosted audio-upload tool.
This page separates the official Python path from Kijai's community ComfyUI wrapper. That distinction matters because a wrapper workflow, quantized checkpoint, or block-swap recipe can change independently from the official model.
Pick the right route first
| Goal | Start here | Trade-off |
|---|---|---|
| Animate one still image | Hosted image-to-video | Image and motion prompt only; no audio upload or speech-driven lip sync |
| Animate a portrait to recorded speech | Official local Avatar setup | Audio-driven generation; requires your own compatible GPU environment |
| Reproduce the official release | Official LongCat-Video repository | Closest to upstream, but requires a CUDA/Python environment and large weights |
| Build a node workflow | Community ComfyUI path | The linked example uses the original Avatar audio stack, not official 1.5 |
| Run a production batch | Pin either path to a tested commit | You own capacity, retries, output review, and upgrade testing |
What the hosted tool supports
The hosted LongCat image-to-video tool accepts a still image and a motion prompt and returns a video file. It offers regular 480p/720p and Distilled 480p choices, with MP4 as the default output. It does not offer a LongCat-Video-Avatar model, audio upload, or lip sync driven by your recording. Animating a portrait here cannot validate an Avatar audio workflow.
Sign-in and sufficient credits are required to submit a hosted generation. Check the displayed credit cost before submitting; model, frame count, and account eligibility still determine the available run. Current plans apply to the hosted service, not to running the open weights on your own GPU.
LongCat Video is an independent hosted service, not Meituan's official model site. Hosted uploads follow the site's privacy policy; the local and community workflows below run in the environment you configure.
What changed in Avatar 1.5
- The audio encoder is Whisper-Large rather than Wav2Vec2.
- The official release describes 8-step inference through DMD2-based distillation.
- It covers realistic people, stylized characters, animals, multi-person scenes, and object interaction.
- Single-stream and multi-stream audio are both supported.
These are upstream claims, not a promise that every portrait, language, device, or duration will be artifact-free. Test your own face angles, speech pace, hands, occlusion, and long sequences before choosing it for production.
Official installation path
Use the official repository when reproducibility matters:
git clone --single-branch --branch main https://github.com/meituan-longcat/LongCat-Video
cd LongCat-Video
conda create -n longcat-video python=3.10
conda activate longcat-video
# Install the PyTorch/CUDA combination for your machine, then:
pip install ninja psutil packaging
pip install flash_attn==2.7.4.post1
pip install -r requirements.txt
conda install -c conda-forge librosa ffmpeg
pip install -r requirements_avatar.txtDownload the LongCat-Video-Avatar-1.5 weights linked by the repository. Do not silently substitute the original Avatar checkpoint: its audio encoder and inference path differ.
Use the official single-person run_demo_avatar_single_audio_to_video.py example with --model_type avatar-v1.5 and --use_distill. For a portrait plus audio, select --stage_1=ai2v and provide the example's image, audio, and prompt through --input_json. The script defaults to the original Avatar version unless you select 1.5; downloading newer weights alone does not select the newer audio path. Follow the official inference examples for the complete launch command and GPU configuration.
The official configuration enables FlashAttention-2 by default and documents FlashAttention-3 or xFormers as alternatives after installation. Treat an attention-backend failure as an environment problem; do not catch it and continue with an unknown configuration.
Community ComfyUI path
Kijai's current wrapper keeps LongCat nodes and example workflows on its main branch:
cd ComfyUI/custom_nodes
git clone https://github.com/kijai/ComfyUI-WanVideoWrapper.git
cd ComfyUI-WanVideoWrapper
pip install -r requirements.txt
git rev-parse HEADRecord the commit printed by the last command with every saved workflow. The linked LongCatAvatar_audio_image_to_video_example_01.json is an original Avatar workflow: checked on September 13, 2026, it loads LongCat-Avatar_bf16.safetensors and a Wav2Vec2 audio encoder. It is not an official Avatar 1.5/Whisper example. Use its matching weights and dependencies; do not substitute 1.5 weights into that graph. For 1.5, use the official path above unless your pinned wrapper version documents a matching Whisper workflow.
The wrapper maintainer calls the project a personal, perpetually work-in-progress sandbox. That is useful context, not a criticism: update it in a separate copy, run one known input, and keep the last working commit so a node rename cannot break all jobs at once.
Local Avatar input checklist that prevents wasted runs
- Use a sharp portrait with the full head visible and no face-covering object.
- Start with one speaker and a short, clean WAV file before multi-person audio.
- Keep the audio sample rate and channel layout consistent across tests.
- Use a prompt that describes identity, framing, clothing, background, and restrained motion.
- Save the model filename, wrapper commit, workflow JSON, seed, dimensions, duration, and audio filename.
VRAM and duration: measure, do not guess
The official Avatar 1.5 pages do not publish a universal “works on 8GB/12GB” minimum. Community FP8, GGUF, CPU offload, and block-swap workflows can reduce peak VRAM, but the result depends on checkpoint format, resolution, frame count, attention backend, system RAM, and wrapper version.
Use a bounded test ladder:
- Run the repository's known example at its documented settings.
- Run 3–5 seconds with your own portrait and audio.
- Record peak VRAM, system RAM, wall time, and output dimensions.
- Change only one variable at a time.
- Increase duration only after lip sync and identity remain stable.
“Supports long video” does not mean an infinite clip has zero drift. Continuation is a repeated generation operation; inspect every join for face changes, color shifts, hand artifacts, background jumps, and audio offset.
Failure triage
| Symptom | Likely boundary | First action |
|---|---|---|
| Nodes are missing | Wrapper install or stale node definitions | Restart ComfyUI and verify the pinned wrapper commit |
| Checkpoint will not load | Original Avatar and Avatar 1.5 files were mixed | Verify the exact model filename and upstream source |
| CUDA out of memory | Resolution, frames, model format, or insufficient offload | Reduce one dimension, shorten the clip, then measure again |
| Lips lag behind audio | Audio preprocessing, frame timing, or wrong encoder path | Test clean single-speaker audio on the official v1.5 example |
| Identity changes at joins | Continuation overlap or difficult source framing | Shorten the continuation and use a clearer reference image |
| First run works, update breaks | Moving community dependencies | Return to the recorded commit and upgrade in a separate copy |
Production acceptance check
Before publishing, watch the complete result with sound and verify lip timing, identity, hands, teeth, eye motion, scene continuity, audio rights, portrait consent, and disclosure requirements. Do not infer commercial rights from “open source” alone; review the model license plus the rights attached to every input and third-party dependency.
Primary sources
- Meituan LongCat-Video official repository
- LongCat-Video-Avatar-1.5 official model card
- Kijai ComfyUI-WanVideoWrapper
- Community original Avatar example workflow (Wav2Vec2)
- fal hosted LongCat Distilled image-to-video input and output schema
To animate a still image without audio input, open hosted image-to-video. To animate a portrait to speech, use the local Avatar paths above. Keep official and community setups in separate environments so one dependency upgrade cannot break both.