Create a Multi-Speaker LipDub

Use a Multi-speaker project whenever more than one face appears at any point in your footage— even when only one person speaks. This workflow lets you review individual face detections, train only the people you want to lip-sync, assign a separate audio file to each speaker, and control where lip-sync is applied. A single project can contain multiple scenes and videos, allowing training data for the same person to be shared across related footage.

Before you begin

A Multi-speaker project gives you more control than the other LipDub project types, but it also requires more preparation.

Before starting:

  • Prepare one clean, accurately timed audio file for each speaker.
  • Preserve all silence so every line occurs at the correct source-video timecode.
  • Include only the dialogue intended for that speaker.
  • Make sure the source and supplemental footage clearly show each person speaking.
  • Confirm that you have permission to use each person’s voice and likeness.
  • Review the detailed media and training recommendations in Prepare Your Content.

Important: Multi-speaker projects do not translate dialogue or generate new speech. You must prepare and upload the finished audio for every speaker.

LipDub removes the original dialogue and replaces it with the audio files you provide.


Understand the Multi-speaker workflow

A Multi-speaker project follows four stages:

Stage

What you will do

1. Upload Videos

Add the source videos and any supplemental footage needed for actor training.

2. Label Actors

Review detected faces, correct mistakes, and assign detections to the appropriate people.

3. Train AI

Train a lip-sync model for each person whose mouth movements you want to change.

4. Upload Audio & Generate

Assign audio to each speaker, choose where lip-sync should be applied, and generate the result.

LipDub uses the term actor for each person whose face is labeled and trained. An actor may appear in several videos or scenes within the same project.


Create a Multi-speaker project

  1. Open Projects.
  2. Select Create Project.
  3. Choose Multi-speaker project.
  4. Select Continue.
  5. Enter a descriptive project name.
  6. Select the language spoken in the source footage.
  7. Select the languages you expect to LipDub.
  8. Select Create.

The selected languages become separate tabs in the project. You can add another language later from Upload Audio & Generate.

Note: Language tabs organize your uploaded audio; LipDub does not create the translated audio for this workflow. Many teams also use the tabs like mini folders to organize alternate audio versions.


Step 1: Upload the videos

The Scenes panel on the left organizes the videos in your project.

You can use scenes to separate:

  • Individual videos
  • Episodes or chapters
  • Different shots
  • Alternate edits
  • Groups of related deliverables

Project organization is flexible, so use the structure that best matches your production.

Add a scene

  1. Select the + beside Scenes.
  2. Name the new scene.
  3. Open the scene before uploading its footage.

Scenes and source files can be renamed as your project evolves.

Upload source footage

  1. Open 1. Upload Videos.
  2. Under Lipdub footage, select Upload videos.
  3. Add the source video or videos you want to generate.
  4. Wait until each clip has been processed successfully.
  5. Select a video thumbnail to preview it in the player.

You will choose the individual video to generate later in the workflow.

Add supplemental training footage

Use Upload additional footage when the source video does not contain enough clear speaking footage for one or more actors.

Supplemental footage:

  • Contributes to actor training
  • Is not included in the generated output
  • Should show the actor actively speaking with their lips visible
  • Should match the source footage as closely as possible

Detailed training-footage recommendations are covered in Prepare Your Content.


Step 2: Label the actors

After the footage has been processed, open 2. Label Actors.

LipDub groups the detected faces into potential actors. Your job is to confirm that each group contains only detections of the same person.

You only need to label and train people whose lips you intend to change.

Review an actor’s detections

  1. Select an actor at the top of the page.
  2. Assign or confirm the appropriate actor label.
  3. Expand Face detections.
  4. Review every thumbnail assigned to that actor.
  5. Select any detection that belongs to someone else.
  6. Select Remove selected.

Incorrect detections can reduce training quality by teaching the model from the wrong face.

Restore a discarded detection

LipDub may place uncertain detections under Discarded face detections.

When a discarded detection belongs to the selected actor:

  1. Expand Discarded face detections.
  2. Select the correct thumbnails.
  3. Select Assign selected.
  4. Choose the appropriate actor.
  5. Select Confirm.

Repeat this process for every person you intend to lip-sync.

Use the same actor across multiple videos

When the same person appears in multiple source videos or scenes, assign their detections to the same actor label.

This allows LipDub to combine the relevant footage and use one actor model across the videos in that Multi-speaker project.

Important: Actor training is currently project-specific. Training cannot be shared between different project types.


Step 3: Train the actors

Open 3. Train AI after the face detections have been reviewed.

  1. Select the checkbox beside each actor you want to train.
  2. Select Train.
  3. Choose a training-quality level.
  4. Confirm the training request.

You can select several actors and train them in parallel.

Choose a training model

Model

Recommended use

Turbo

Quick previews using LipDub’s generic model. Turbo performs no actor training and does not create a reusable actor model.

Flash

Faster training with good quality for social media and everyday content.

Premium

Longer training for high-resolution and higher-quality results.

Ultra

The highest-quality option for film, television, and premium productions.

Higher-quality training generally takes longer. LipDub displays the current credit estimate before training begins.

Review the training status

Each actor displays their current status.

  • Actor has been trained means the actor is ready for lip-sync.
  • The tracks for this actor have been modified means the face detections changed after training and the actor should be trained again.
  • Training could not be completed means the attempt failed and should be retried.

Use Refresh status to check an active training job.

An actor must be successfully trained before you can apply lip-sync to their whole video or selected regions. An untrained actor must be set to Not at all before the result can be generated.

Reuse training for an identical video

When you upload a video that LipDub recognizes as identical to one previously trained, the platform may ask whether you want to reuse the existing training data.

You can:

  • Reuse the previous training, or
  • Continue and train the video normally

Selecting a different training model also allows you to create new training instead of reusing the existing model.


Step 4: Prepare the speaker audio

Before uploading, create one separate audio file for each person who should be heard in the finished video.

Each file should:

  • Contain only that person’s dialogue
  • Begin at the same timecode as the source video
  • Preserve all silence before, between, and after their lines
  • Place every line exactly where it should occur
  • Exclude any dialogue belonging to another speaker

For example, when a speaker’s first line begins 20 seconds into the video, their file should normally contain 20 seconds of silence before that line.

Important: LipDub does not automatically reposition each line. The timing in your uploaded file determines when the dialogue is heard.

If two speaker files contain dialogue at the same time, both voices will be heard at the same time. Preserve silence carefully to prevent unintended overlapping dialogue.

Audio is still added when a speaker is off-screen, facing away from the camera, or otherwise not available for visual lip-sync.

Audio-length behavior

  • When a speaker’s audio is longer than the source video, it is cut off at the end of the video.
  • When a speaker’s audio is shorter than the source video, the speaker is silent after their audio ends and no further lip-sync is applied.

Upload and assign the audio

Open 4. Upload Audio & Generate.

Select the source video

Under Select your video, choose the source video you want to generate.

The video player lets you review the source footage while preparing the speaker assignments.

Select a language or version

Under Choose your language, select the tab for the audio set you want to configure.

Select the + to add another language to the project.

Each language tab stores its own audio assignments and generation setup. You can configure and start generations for several tabs without waiting for each previous generation to finish.

Upload one file for each speaker

Under Upload audio stems:

  1. Find the actor who should receive the audio.
  2. Select Upload audio.
  3. Add that actor’s prepared audio file.
  4. Preview the uploaded file using the audio player.
  5. Confirm that the duration and timing are correct.
  6. Repeat for every speaker who should be heard.

You can also use the audio-file menu to select a file that has already been uploaded for that actor and language.


Choose where lip-sync is applied

For each actor, choose one of the following options.

Option

Behavior

For the whole video

Applies visual lip-sync throughout the selected video wherever that actor is visible.

Selected regions

Applies visual lip-sync only within the timecode ranges you enter. Outside those ranges, the original facial movement remains.

Not at all

Does not apply lip-sync and makes that actor silent in the generated version.

These settings apply separately to each actor in that specific generation.

Use Selected regions

To lip-sync only part of the video:

  1. Choose Selected regions for the actor.
  2. Select Add region.
  3. Enter the Start timecode.
  4. Enter the End timecode.
  5. Add additional regions where needed.

Timecodes use the following format:

Hours : Minutes : Seconds : Frames

You can delete a region using the trash icon beside it.

LipDub blends between the generated lip-sync and the original footage at the boundaries of each region.

Important: Selected regions control only the actor’s visual lip-sync. They do not trim or limit the uploaded audio. The actor’s complete audio file is still included in the generated video, including outside the selected regions.

Even when you select only a single frame for lip-sync, the complete uploaded audio file will still play. Prepare or edit the audio itself when only part of it should be heard.


Generate the video

After every actor has the correct audio and lip-sync setting:

  1. Select Generate result.
  2. Enter a name for the generated video.
  3. Choose an output resolution.
  4. Configure the AI-generated-video option where appropriate.
  5. Review the cost breakdown.
  6. Select Generate.

Choose an output resolution

Resolution

Output behavior

Standard — Up to HD

Produces output up to 1080p. A 4K source generated with Standard will produce a 1080p result.

Pro — Up to 4K

Supports output up to 4K.

LipDub shows the current credit estimate for the selected resolution before generation.

Optimize AI-generated footage

Turn on Optimize for AI generated video when your source footage was created with an AI video tool.

This uses a version of LipDub’s model tuned for AI-generated footage. It is generally recommended for AI source video.

Leave it off for conventional human footage unless you are comparing different results. It may reduce visual texture quality on real human shots.


Review and download the result

Completed and in-progress generations appear under Preview and download result.

When a result is ready:

  1. Select the version from the results menu.
  2. Play the complete video.
  3. Check each speaker’s dialogue, timing, and lip-sync.
  4. Review the transitions into and out of any selected regions.
  5. Confirm that off-screen dialogue and overlapping sections behave as intended.
  6. Select Download to save the finished video.
  7. Use the thumbs-up or thumbs-down controls under Good result? to provide quality feedback.

You can return to the audio assignments, select another language or set of files, and generate additional versions without uploading or training the source footage again.

Processing taking longer than expected? If one processing, training, or generation action remains in progress for more than two hours, contact support@lipdub.ai. Failed training and generation attempts are normally refunded automatically.


Last updated: 7/20/26, 5:35 PM