LongCat Video Avatar 1.5 ComfyUI Setup

Jan 23, 2026

Choose image motion or an audio-driven avatar

Hosted image-to-video animates a still image without audio input. For speech-driven lip sync, follow the local Avatar setup below.

LongCat-Video-Avatar 1.5 is Meituan's current audio-driven avatar release. It is not the same checkpoint or audio stack as the December 2025 version: v1.5 replaces Wav2Vec2 with Whisper-Large, supports single- and multi-stream audio, and uses an 8-step distilled inference path. The official tasks are audio + text to video, audio + text + image to video, and video continuation.

For an overview of the separate audio-driven model, see the LongCat Video Avatar guide. That page documents local inference; it is not a hosted audio-upload tool.

This page separates the official Python path from Kijai's community ComfyUI wrapper. That distinction matters because a wrapper workflow, quantized checkpoint, or block-swap recipe can change independently from the official model.

Pick the right route first

GoalStart hereTrade-off
Animate one still imageHosted image-to-videoImage and motion prompt only; no audio upload or speech-driven lip sync
Animate a portrait to recorded speechOfficial local Avatar setupAudio-driven generation; requires your own compatible GPU environment
Reproduce the official releaseOfficial LongCat-Video repositoryClosest to upstream, but requires a CUDA/Python environment and large weights
Build a node workflowCommunity ComfyUI pathThe linked example uses the original Avatar audio stack, not official 1.5
Run a production batchPin either path to a tested commitYou own capacity, retries, output review, and upgrade testing

What the hosted tool supports

The hosted LongCat image-to-video tool accepts a still image and a motion prompt and returns a video file. It offers regular 480p/720p and Distilled 480p choices, with MP4 as the default output. It does not offer a LongCat-Video-Avatar model, audio upload, or lip sync driven by your recording. Animating a portrait here cannot validate an Avatar audio workflow.

Sign-in and sufficient credits are required to submit a hosted generation. Check the displayed credit cost before submitting; model, frame count, and account eligibility still determine the available run. Current plans apply to the hosted service, not to running the open weights on your own GPU.

LongCat Video is an independent hosted service, not Meituan's official model site. Hosted uploads follow the site's privacy policy; the local and community workflows below run in the environment you configure.

What changed in Avatar 1.5

  • The audio encoder is Whisper-Large rather than Wav2Vec2.
  • The official release describes 8-step inference through DMD2-based distillation.
  • It covers realistic people, stylized characters, animals, multi-person scenes, and object interaction.
  • Single-stream and multi-stream audio are both supported.

These are upstream claims, not a promise that every portrait, language, device, or duration will be artifact-free. Test your own face angles, speech pace, hands, occlusion, and long sequences before choosing it for production.

Official installation path

Use the official repository when reproducibility matters:

git clone --single-branch --branch main https://github.com/meituan-longcat/LongCat-Video
cd LongCat-Video

conda create -n longcat-video python=3.10
conda activate longcat-video

# Install the PyTorch/CUDA combination for your machine, then:
pip install ninja psutil packaging
pip install flash_attn==2.7.4.post1
pip install -r requirements.txt
conda install -c conda-forge librosa ffmpeg
pip install -r requirements_avatar.txt

Download the LongCat-Video-Avatar-1.5 weights linked by the repository. Do not silently substitute the original Avatar checkpoint: its audio encoder and inference path differ.

Use the official single-person run_demo_avatar_single_audio_to_video.py example with --model_type avatar-v1.5 and --use_distill. For a portrait plus audio, select --stage_1=ai2v and provide the example's image, audio, and prompt through --input_json. The script defaults to the original Avatar version unless you select 1.5; downloading newer weights alone does not select the newer audio path. Follow the official inference examples for the complete launch command and GPU configuration.

The official configuration enables FlashAttention-2 by default and documents FlashAttention-3 or xFormers as alternatives after installation. Treat an attention-backend failure as an environment problem; do not catch it and continue with an unknown configuration.

Community ComfyUI path

Kijai's current wrapper keeps LongCat nodes and example workflows on its main branch:

cd ComfyUI/custom_nodes
git clone https://github.com/kijai/ComfyUI-WanVideoWrapper.git
cd ComfyUI-WanVideoWrapper
pip install -r requirements.txt
git rev-parse HEAD

Record the commit printed by the last command with every saved workflow. The linked LongCatAvatar_audio_image_to_video_example_01.json is an original Avatar workflow: checked on September 13, 2026, it loads LongCat-Avatar_bf16.safetensors and a Wav2Vec2 audio encoder. It is not an official Avatar 1.5/Whisper example. Use its matching weights and dependencies; do not substitute 1.5 weights into that graph. For 1.5, use the official path above unless your pinned wrapper version documents a matching Whisper workflow.

The wrapper maintainer calls the project a personal, perpetually work-in-progress sandbox. That is useful context, not a criticism: update it in a separate copy, run one known input, and keep the last working commit so a node rename cannot break all jobs at once.

Local Avatar input checklist that prevents wasted runs

  1. Use a sharp portrait with the full head visible and no face-covering object.
  2. Start with one speaker and a short, clean WAV file before multi-person audio.
  3. Keep the audio sample rate and channel layout consistent across tests.
  4. Use a prompt that describes identity, framing, clothing, background, and restrained motion.
  5. Save the model filename, wrapper commit, workflow JSON, seed, dimensions, duration, and audio filename.

VRAM and duration: measure, do not guess

The official Avatar 1.5 pages do not publish a universal “works on 8GB/12GB” minimum. Community FP8, GGUF, CPU offload, and block-swap workflows can reduce peak VRAM, but the result depends on checkpoint format, resolution, frame count, attention backend, system RAM, and wrapper version.

Use a bounded test ladder:

  1. Run the repository's known example at its documented settings.
  2. Run 3–5 seconds with your own portrait and audio.
  3. Record peak VRAM, system RAM, wall time, and output dimensions.
  4. Change only one variable at a time.
  5. Increase duration only after lip sync and identity remain stable.

“Supports long video” does not mean an infinite clip has zero drift. Continuation is a repeated generation operation; inspect every join for face changes, color shifts, hand artifacts, background jumps, and audio offset.

Failure triage

SymptomLikely boundaryFirst action
Nodes are missingWrapper install or stale node definitionsRestart ComfyUI and verify the pinned wrapper commit
Checkpoint will not loadOriginal Avatar and Avatar 1.5 files were mixedVerify the exact model filename and upstream source
CUDA out of memoryResolution, frames, model format, or insufficient offloadReduce one dimension, shorten the clip, then measure again
Lips lag behind audioAudio preprocessing, frame timing, or wrong encoder pathTest clean single-speaker audio on the official v1.5 example
Identity changes at joinsContinuation overlap or difficult source framingShorten the continuation and use a clearer reference image
First run works, update breaksMoving community dependenciesReturn to the recorded commit and upgrade in a separate copy

Production acceptance check

Before publishing, watch the complete result with sound and verify lip timing, identity, hands, teeth, eye motion, scene continuity, audio rights, portrait consent, and disclosure requirements. Do not infer commercial rights from “open source” alone; review the model license plus the rights attached to every input and third-party dependency.

Primary sources

To animate a still image without audio input, open hosted image-to-video. To animate a portrait to speech, use the local Avatar paths above. Keep official and community setups in separate environments so one dependency upgrade cannot break both.