Create a Multi-Speaker LipDub
Before you begin
A Multi-speaker project gives you more control than the other LipDub project types, but it also requires more preparation.
Before starting:
- Prepare one clean, accurately timed audio file for each speaker.
- Preserve all silence so every line occurs at the correct source-video timecode.
- Include only the dialogue intended for that speaker.
- Make sure the source and supplemental footage clearly show each person speaking.
- Confirm that you have permission to use each person’s voice and likeness.
- Review the detailed media and training recommendations in Prepare Your Content.
Important: Multi-speaker projects do not translate dialogue or generate new speech. You must prepare and upload the finished audio for every speaker.
LipDub removes the original dialogue and replaces it with the audio files you provide.
Understand the Multi-speaker workflow
A Multi-speaker project follows four stages:
Stage | What you will do |
|---|---|
1. Upload Videos | Add the source videos and any supplemental footage needed for actor training. |
2. Label Actors | Review detected faces, correct mistakes, and assign detections to the appropriate people. |
3. Train AI | Train a lip-sync model for each person whose mouth movements you want to change. |
4. Upload Audio & Generate | Assign audio to each speaker, choose where lip-sync should be applied, and generate the result. |
LipDub uses the term actor for each person whose face is labeled and trained. An actor may appear in several videos or scenes within the same project.
Create a Multi-speaker project
- Open Projects.
- Select Create Project.
- Choose Multi-speaker project.
- Select Continue.
- Enter a descriptive project name.
- Select the language spoken in the source footage.
- Select the languages you expect to LipDub.
- Select Create.
The selected languages become separate tabs in the project. You can add another language later from Upload Audio & Generate.
Note: Language tabs organize your uploaded audio; LipDub does not create the translated audio for this workflow. Many teams also use the tabs like mini folders to organize alternate audio versions.
Step 1: Upload the videos
The Scenes panel on the left organizes the videos in your project.
You can use scenes to separate:
- Individual videos
- Episodes or chapters
- Different shots
- Alternate edits
- Groups of related deliverables
Project organization is flexible, so use the structure that best matches your production.
Add a scene
- Select the + beside Scenes.
- Name the new scene.
- Open the scene before uploading its footage.
Scenes and source files can be renamed as your project evolves.
Upload source footage
- Open 1. Upload Videos.
- Under Lipdub footage, select Upload videos.
- Add the source video or videos you want to generate.
- Wait until each clip has been processed successfully.
- Select a video thumbnail to preview it in the player.
You will choose the individual video to generate later in the workflow.
Add supplemental training footage
Use Upload additional footage when the source video does not contain enough clear speaking footage for one or more actors.
Supplemental footage:
- Contributes to actor training
- Is not included in the generated output
- Should show the actor actively speaking with their lips visible
- Should match the source footage as closely as possible
Detailed training-footage recommendations are covered in Prepare Your Content.
Step 2: Label the actors
After the footage has been processed, open 2. Label Actors.
LipDub groups the detected faces into potential actors. Your job is to confirm that each group contains only detections of the same person.
You only need to label and train people whose lips you intend to change.
Review an actor’s detections
- Select an actor at the top of the page.
- Assign or confirm the appropriate actor label.
- Expand Face detections.
- Review every thumbnail assigned to that actor.
- Select any detection that belongs to someone else.
- Select Remove selected.
Incorrect detections can reduce training quality by teaching the model from the wrong face.
Restore a discarded detection
LipDub may place uncertain detections under Discarded face detections.
When a discarded detection belongs to the selected actor:
- Expand Discarded face detections.
- Select the correct thumbnails.
- Select Assign selected.
- Choose the appropriate actor.
- Select Confirm.
Repeat this process for every person you intend to lip-sync.
Use the same actor across multiple videos
When the same person appears in multiple source videos or scenes, assign their detections to the same actor label.
This allows LipDub to combine the relevant footage and use one actor model across the videos in that Multi-speaker project.
Important: Actor training is currently project-specific. Training cannot be shared between different project types.
Step 3: Train the actors
Open 3. Train AI after the face detections have been reviewed.
- Select the checkbox beside each actor you want to train.
- Select Train.
- Choose a training-quality level.
- Confirm the training request.
You can select several actors and train them in parallel.
Choose a training model
Model | Recommended use |
|---|---|
Turbo | Quick previews using LipDub’s generic model. Turbo performs no actor training and does not create a reusable actor model. |
Flash | Faster training with good quality for social media and everyday content. |
Premium | Longer training for high-resolution and higher-quality results. |
Ultra | The highest-quality option for film, television, and premium productions. |
Higher-quality training generally takes longer. LipDub displays the current credit estimate before training begins.
Review the training status
Each actor displays their current status.
- Actor has been trained means the actor is ready for lip-sync.
- The tracks for this actor have been modified means the face detections changed after training and the actor should be trained again.
- Training could not be completed means the attempt failed and should be retried.
Use Refresh status to check an active training job.
An actor must be successfully trained before you can apply lip-sync to their whole video or selected regions. An untrained actor must be set to Not at all before the result can be generated.
Reuse training for an identical video
When you upload a video that LipDub recognizes as identical to one previously trained, the platform may ask whether you want to reuse the existing training data.
You can:
- Reuse the previous training, or
- Continue and train the video normally
Selecting a different training model also allows you to create new training instead of reusing the existing model.
Step 4: Prepare the speaker audio
Before uploading, create one separate audio file for each person who should be heard in the finished video.
Each file should:
- Contain only that person’s dialogue
- Begin at the same timecode as the source video
- Preserve all silence before, between, and after their lines
- Place every line exactly where it should occur
- Exclude any dialogue belonging to another speaker
For example, when a speaker’s first line begins 20 seconds into the video, their file should normally contain 20 seconds of silence before that line.
Important: LipDub does not automatically reposition each line. The timing in your uploaded file determines when the dialogue is heard.
If two speaker files contain dialogue at the same time, both voices will be heard at the same time. Preserve silence carefully to prevent unintended overlapping dialogue.
Audio is still added when a speaker is off-screen, facing away from the camera, or otherwise not available for visual lip-sync.
Audio-length behavior
- When a speaker’s audio is longer than the source video, it is cut off at the end of the video.
- When a speaker’s audio is shorter than the source video, the speaker is silent after their audio ends and no further lip-sync is applied.
Upload and assign the audio
Open 4. Upload Audio & Generate.
Select the source video
Under Select your video, choose the source video you want to generate.
The video player lets you review the source footage while preparing the speaker assignments.
Select a language or version
Under Choose your language, select the tab for the audio set you want to configure.
Select the + to add another language to the project.
Each language tab stores its own audio assignments and generation setup. You can configure and start generations for several tabs without waiting for each previous generation to finish.
Upload one file for each speaker
Under Upload audio stems:
- Find the actor who should receive the audio.
- Select Upload audio.
- Add that actor’s prepared audio file.
- Preview the uploaded file using the audio player.
- Confirm that the duration and timing are correct.
- Repeat for every speaker who should be heard.
You can also use the audio-file menu to select a file that has already been uploaded for that actor and language.
Choose where lip-sync is applied
For each actor, choose one of the following options.
Option | Behavior |
|---|---|
For the whole video | Applies visual lip-sync throughout the selected video wherever that actor is visible. |
Selected regions | Applies visual lip-sync only within the timecode ranges you enter. Outside those ranges, the original facial movement remains. |
Not at all | Does not apply lip-sync and makes that actor silent in the generated version. |
These settings apply separately to each actor in that specific generation.
Use Selected regions
To lip-sync only part of the video:
- Choose Selected regions for the actor.
- Select Add region.
- Enter the Start timecode.
- Enter the End timecode.
- Add additional regions where needed.
Timecodes use the following format:
Hours : Minutes : Seconds : Frames
You can delete a region using the trash icon beside it.
LipDub blends between the generated lip-sync and the original footage at the boundaries of each region.
Important: Selected regions control only the actor’s visual lip-sync. They do not trim or limit the uploaded audio. The actor’s complete audio file is still included in the generated video, including outside the selected regions.
Even when you select only a single frame for lip-sync, the complete uploaded audio file will still play. Prepare or edit the audio itself when only part of it should be heard.
Generate the video
After every actor has the correct audio and lip-sync setting:
- Select Generate result.
- Enter a name for the generated video.
- Choose an output resolution.
- Configure the AI-generated-video option where appropriate.
- Review the cost breakdown.
- Select Generate.
Choose an output resolution
Resolution | Output behavior |
|---|---|
Standard — Up to HD | Produces output up to 1080p. A 4K source generated with Standard will produce a 1080p result. |
Pro — Up to 4K | Supports output up to 4K. |
LipDub shows the current credit estimate for the selected resolution before generation.
Optimize AI-generated footage
Turn on Optimize for AI generated video when your source footage was created with an AI video tool.
This uses a version of LipDub’s model tuned for AI-generated footage. It is generally recommended for AI source video.
Leave it off for conventional human footage unless you are comparing different results. It may reduce visual texture quality on real human shots.
Review and download the result
Completed and in-progress generations appear under Preview and download result.
When a result is ready:
- Select the version from the results menu.
- Play the complete video.
- Check each speaker’s dialogue, timing, and lip-sync.
- Review the transitions into and out of any selected regions.
- Confirm that off-screen dialogue and overlapping sections behave as intended.
- Select Download to save the finished video.
- Use the thumbs-up or thumbs-down controls under Good result? to provide quality feedback.
You can return to the audio assignments, select another language or set of files, and generate additional versions without uploading or training the source footage again.
Processing taking longer than expected? If one processing, training, or generation action remains in progress for more than two hours, contact support@lipdub.ai. Failed training and generation attempts are normally refunded automatically.