Model Category

Lip Sync Models — 15 Ways to Match Audio to a Face

9 named models across two input modes — animate a portrait from audio, or re-sync the lips on an existing video. 15 models in total, no content filters.

15 Lip Sync Models9 Named Across Two Input Modes

Lip sync models drive a talking video from an audio track, either animating a still portrait or re-syncing the lips on an existing clip. The category ships 15 models in total, 9 of them named with dedicated endpoints split across image-based and video-based modes. This page covers the models themselves — see the Lip Sync Studio page for the generation workflow.

Lip Sync Models

15 Lip Sync Models, One Open Generative AI Catalog

Every model below is listed by the input mode it takes — a portrait image plus audio, or an existing video plus audio. The Lip Sync Studio page covers how a generation run is put together; this page stays on the models.

infinite-talk

Portrait image + audio → talking video, 480p or 720p.

wan-2.2-speech-to-video

Portrait image + audio → talking video, 480p or 720p.

ltx-2.3-lipsync

Portrait image + audio → talking video, up to 1080p.

ltx-2-19b-lipsync

Portrait image + audio → talking video, up to 1080p.

sync-lipsync

Re-sync the lips on an existing video with new audio.

latentsync

Video + audio → lipsync video.

creatify-lipsync

Video + audio → lipsync video.

veed-lipsync

Video + audio → lipsync video.

infinite-talk-v2v

Video + audio → lipsync video, 480p or 720p.

Hosted vs. run it locally

Hosted Today, With No Local Lip Sync Path

Open Generative AI documents local inference for image and video generation only, so these models run through the hosted API — the same one the rest of the catalog uses.

Hosted

All 15 lip sync models — 9 of them named, across image-based and video-based modes — run through the hosted API with no local setup.

Local Inference

Local inference (sd.cpp and Wan2GP) is documented for image and video generation only; the source project lists no local path for lip sync, so it runs hosted today.

Compare

Three Lip Sync Models, Side by Side

A quick look at three representative models in the catalog — one per input mode, plus the highest-resolution option.

ModelInput modeResolutions
Infinite TalkPortrait image + audio480p, 720p
LTX 2.3 LipsyncPortrait image + audio480p, 720p, 1080p
Sync LipsyncExisting video + audioNot specified

Swipe to see the full table

FAQ

Lip Sync Model FAQ

15 in total; 9 of them are named with dedicated endpoints across two input modes — 4 image-based models and 5 video-based models.

Explore more

More Open Generative AI Model Categories

Lip sync pairs closely with the categories below. For the generation workflow itself — choosing an input mode, preparing the audio track, running the job — read the Lip Sync Studio page instead.

Lip Sync a Video With Open Generative AI — Free, Right Now

15 lip sync models, no content filters, no subscription. No account required to try it.

Try Free
Open Generative AI images and videos are generated by AI. Review results before using them for any purpose that requires accuracy or authenticity.