Replace Dialogue in a Single-Speaker Video
Before you begin
A Dialogue Replacement project is intended for footage with one visible face.
Important: Use a Multi-speaker project whenever more than one face appears at any point in the source video— even when only one person speaks.
Before starting:
- Review the video and audio recommendations in Prepare Your Content.
- Use footage with a clear, unobstructed view of the speaker’s face.
- Aim to include at least 10 seconds of the speaker talking clearly.
- Use the cleanest replacement audio available.
- Confirm that you have permission to use and clone the speaker’s voice and likeness.
Important: Your uploaded or generated dialogue becomes the video’s complete final audio track. The source music, ambience, sound effects, and original dialogue are not retained.
A Dialogue Replacement project uses one source video. After that video is processed and trained, you can use New dub to create multiple audio and lip-sync versions from the same footage.
Create a Dialogue Replacement project
- Open Projects.
- Select Create Project.
- Choose Dialogue Replacement project.
- Select Continue.
- Enter a descriptive project name.
- Select Create.
- Select Upload a video and add your source video.
LipDub will process the footage and train the visible speaker. You can leave the page while processing continues.
Limited testing: A library of LipDub AI avatars may be made available for qualifying Dialogue Replacement use cases. Contact support@lipdub.ai to discuss access.
Create a new dub
After the source video is ready:
- Open the project.
- Select New dub under Your renders.
- Choose how you want to provide the new dialogue.
Option | Use it when |
|---|---|
Upload Audio | You already have a recorded performance that should become the final audio. |
Replace Dialogue | You want LipDub to generate speech from a written script. |
For automatic video translation, use a Translation project instead.
Option 1: Upload replacement audio
Use Upload Audio when the new dialogue has already been recorded.
- Open the Upload Audio tab.
- Drag the file into the upload area or select Choose file.
- Select Continue after the file has uploaded.
LipDub accepts most common audio files. You can also upload an MP4 or MOV containing the audio instead of extracting it first.
Your file should contain only the audio you want in the finished video. Any background noise, music, sound effects, or unintended speech in the file will be included in the result.
How audio length affects the result
- When the replacement audio extends beyond the source video, it is cut off at the end of the video.
- When the replacement audio ends before the source video, the output is shortened when the replacement audio ends.
Review the timing in Studio before generating when the new dialogue should begin somewhere other than the start of the video.
Option 2: Generate audio from a script
Use Replace Dialogue when you want LipDub to generate the new speech.
- Open the Replace Dialogue tab.
- Under Target Language, select the language in which the script should be spoken.
- Enter the dialogue exactly as you want it spoken. The script field currently supports up to 3,000 characters.
- Choose a voice option:
- Choose a voice
- Clone the speaker’s voice
- Review the estimated credit usage.
- Select Continue.
LipDub will generate the audio and add it to Your renders.
Choose a voice
Select Choose a voice to use a voice from the LipDub voice library.
Select Switch to open the library. From there, you can:
- Search for a voice
- Filter voices by gender and age
- Preview a voice
- Select Choose voice
Connected ElevenLabs voices may also appear under My ElevenLabs voices when the integration has been configured.
Clone the speaker’s voice
Select Clone the speaker’s voice to generate the script using the vocal identity from the source video.
The source should contain at least 10 seconds of clear speech. Thirty seconds or more is recommended for better voice similarity.
Only clone a voice you have permission to use. LipDub may restrict recognized celebrity voices in certain cases.
Preview the generated audio
When the audio is ready, LipDub provides three options:
- Review in Studio
- Proceed to LipDubbing
- Download audio
Select Proceed to LipDubbing when the generated audio already starts at the correct time and does not require further review.
Select Review in Studio when you need to preview the audio against the video or change where it begins.
Position the audio in Studio
Studio displays the video, the replacement audio, and a timeline.
The Audio Information panel shows:
- The selected audio version
- Audio duration
- Language
- Generation type
Use the Select Audio menu to switch between audio versions created within the project.
Adjust the start time
The timeline contains two tracks:
- Target Audio contains the new dialogue.
- Original Audio contains the source audio for reference.
To reposition the new dialogue:
- Select the audio clip on the Target Audio track.
- Drag it to the time where the dialogue should begin.
- Play the video from shortly before the new start point.
- Confirm that the speech begins at the intended moment.
Note: The Original Audio track is provided as a timing reference. It is not mixed into the final Dialogue Replacement output.
Use the timeline zoom controls when you need to make a more precise adjustment.
Create or select another audio version
From Studio, you can:
- Select another version from Select Audio
- Download the selected audio
- Select Generate new audio to upload another file or generate another script
This lets you compare several performances while continuing to use the same trained source video.
Generate the lip-synced video
When the selected audio and timing are ready:
- Select Generate LipDub in Studio.
When skipping Studio, select Proceed to LipDubbing from the audio preview instead. - Review the options in Generate Summary.
- Adjust Expressiveness if needed.
- Enable Optimize for AI generated video when appropriate.
- Review the estimated credit usage.
- Select Generate.
Set expressiveness
Expressiveness controls how strongly the speaker’s lips and jaw react to the new dialogue.
Setting | Result |
|---|---|
1-2: Lower settings | Less pronounced movement and less emphasis on each syllable |
3: Default | Balanced lip and jaw movement |
4-5: Higher settings | More open, expressive, and reactive movement |
Expressiveness is a visual preference and generally does not increase the likelihood of artifacts. You can create additional versions at different settings to compare them.
Optimize AI-generated footage
Enable Optimize for AI generated video when the source video was created with an AI video tool.
The setting uses a version of the LipDub model tuned for AI source footage. It is recommended for AI-generated video, although you can generate versions with the setting on and off to compare them.
Leave it off for conventional human footage unless you are experimenting with different results. It may reduce visual texture quality on real human video.
Review and download the result
Each audio or generated video version appears under Your renders.
When the LipDub is complete:
- Select the version you want to review.
- Use Original and Lipdubbed to compare the source and generated videos.
- Watch the full result and check:
- Dialogue timing
- Pronunciation and voice quality
- Lip and jaw movement
- The beginning and end of the output
- Select Download result to save the finished video.
- Use the thumbs-up or thumbs-down controls under Good result? to provide quality feedback.
Select New dub to create another version from the same source video. The existing speaker training will be reused.
Legacy SRT upload
Direct SRT upload is no longer enabled by default. It may be enabled by special request for qualifying Dialogue Replacement use cases.
Contact support@lipdub.ai with details about your intended workflow.
Processing taking longer than expected? If one processing, training, audio-generation, or final-generation action remains in progress for more than two hours, contact support@lipdub.ai. Technical failures are normally refunded automatically.