Replace Dialogue in a Single-Speaker Video

Use a Dialogue Replacement project to replace the speech in a video containing one visible face. Upload finished dialogue or write a script for LipDub to generate in a selected voice, then review its timing before creating the lip-synced video. Once the source video has been processed and trained, you can create additional dialogue versions without uploading or training that video again.

Before you begin

A Dialogue Replacement project is intended for footage with one visible face.

Important: Use a Multi-speaker project whenever more than one face appears at any point in the source video— even when only one person speaks.

Before starting:

  • Review the video and audio recommendations in Prepare Your Content.
  • Use footage with a clear, unobstructed view of the speaker’s face.
  • Aim to include at least 10 seconds of the speaker talking clearly.
  • Use the cleanest replacement audio available.
  • Confirm that you have permission to use and clone the speaker’s voice and likeness.

Important: Your uploaded or generated dialogue becomes the video’s complete final audio track. The source music, ambience, sound effects, and original dialogue are not retained.

A Dialogue Replacement project uses one source video. After that video is processed and trained, you can use New dub to create multiple audio and lip-sync versions from the same footage.


Create a Dialogue Replacement project

  1. Open Projects.
  2. Select Create Project.
  3. Choose Dialogue Replacement project.
  4. Select Continue.
  5. Enter a descriptive project name.
  6. Select Create.
  7. Select Upload a video and add your source video.

LipDub will process the footage and train the visible speaker. You can leave the page while processing continues.

Limited testing: A library of LipDub AI avatars may be made available for qualifying Dialogue Replacement use cases. Contact support@lipdub.ai to discuss access.


Create a new dub

After the source video is ready:

  1. Open the project.
  2. Select New dub under Your renders.
  3. Choose how you want to provide the new dialogue.

Option

Use it when

Upload Audio

You already have a recorded performance that should become the final audio.

Replace Dialogue

You want LipDub to generate speech from a written script.

For automatic video translation, use a Translation project instead.


Option 1: Upload replacement audio

Use Upload Audio when the new dialogue has already been recorded.

  1. Open the Upload Audio tab.
  2. Drag the file into the upload area or select Choose file.
  3. Select Continue after the file has uploaded.

LipDub accepts most common audio files. You can also upload an MP4 or MOV containing the audio instead of extracting it first.

Your file should contain only the audio you want in the finished video. Any background noise, music, sound effects, or unintended speech in the file will be included in the result.

How audio length affects the result

  • When the replacement audio extends beyond the source video, it is cut off at the end of the video.
  • When the replacement audio ends before the source video, the output is shortened when the replacement audio ends.

Review the timing in Studio before generating when the new dialogue should begin somewhere other than the start of the video.


Option 2: Generate audio from a script

Use Replace Dialogue when you want LipDub to generate the new speech.

  1. Open the Replace Dialogue tab.
  2. Under Target Language, select the language in which the script should be spoken.
  3. Enter the dialogue exactly as you want it spoken. The script field currently supports up to 3,000 characters.
  4. Choose a voice option:
    • Choose a voice
    • Clone the speaker’s voice
  5. Review the estimated credit usage.
  6. Select Continue.

LipDub will generate the audio and add it to Your renders.

Choose a voice

Select Choose a voice to use a voice from the LipDub voice library.

Select Switch to open the library. From there, you can:

  • Search for a voice
  • Filter voices by gender and age
  • Preview a voice
  • Select Choose voice

Connected ElevenLabs voices may also appear under My ElevenLabs voices when the integration has been configured.

Clone the speaker’s voice

Select Clone the speaker’s voice to generate the script using the vocal identity from the source video.

The source should contain at least 10 seconds of clear speech. Thirty seconds or more is recommended for better voice similarity.

Only clone a voice you have permission to use. LipDub may restrict recognized celebrity voices in certain cases.


Preview the generated audio

When the audio is ready, LipDub provides three options:

  • Review in Studio
  • Proceed to LipDubbing
  • Download audio

Select Proceed to LipDubbing when the generated audio already starts at the correct time and does not require further review.

Select Review in Studio when you need to preview the audio against the video or change where it begins.


Position the audio in Studio

Studio displays the video, the replacement audio, and a timeline.

The Audio Information panel shows:

  • The selected audio version
  • Audio duration
  • Language
  • Generation type

Use the Select Audio menu to switch between audio versions created within the project.

Adjust the start time

The timeline contains two tracks:

  • Target Audio contains the new dialogue.
  • Original Audio contains the source audio for reference.

To reposition the new dialogue:

  1. Select the audio clip on the Target Audio track.
  2. Drag it to the time where the dialogue should begin.
  3. Play the video from shortly before the new start point.
  4. Confirm that the speech begins at the intended moment.

Note: The Original Audio track is provided as a timing reference. It is not mixed into the final Dialogue Replacement output.

Use the timeline zoom controls when you need to make a more precise adjustment.

Create or select another audio version

From Studio, you can:

  • Select another version from Select Audio
  • Download the selected audio
  • Select Generate new audio to upload another file or generate another script

This lets you compare several performances while continuing to use the same trained source video.


Generate the lip-synced video

When the selected audio and timing are ready:

  1. Select Generate LipDub in Studio.
    When skipping Studio, select Proceed to LipDubbing from the audio preview instead.
  2. Review the options in Generate Summary.
  3. Adjust Expressiveness if needed.
  4. Enable Optimize for AI generated video when appropriate.
  5. Review the estimated credit usage.
  6. Select Generate.

Set expressiveness

Expressiveness controls how strongly the speaker’s lips and jaw react to the new dialogue.

Setting

Result

1-2: Lower settings

Less pronounced movement and less emphasis on each syllable

3: Default

Balanced lip and jaw movement

4-5: Higher settings

More open, expressive, and reactive movement

Expressiveness is a visual preference and generally does not increase the likelihood of artifacts. You can create additional versions at different settings to compare them.

Optimize AI-generated footage

Enable Optimize for AI generated video when the source video was created with an AI video tool.

The setting uses a version of the LipDub model tuned for AI source footage. It is recommended for AI-generated video, although you can generate versions with the setting on and off to compare them.

Leave it off for conventional human footage unless you are experimenting with different results. It may reduce visual texture quality on real human video.


Review and download the result

Each audio or generated video version appears under Your renders.

When the LipDub is complete:

  1. Select the version you want to review.
  2. Use Original and Lipdubbed to compare the source and generated videos.
  3. Watch the full result and check:
    • Dialogue timing
    • Pronunciation and voice quality
    • Lip and jaw movement
    • The beginning and end of the output
  4. Select Download result to save the finished video.
  5. Use the thumbs-up or thumbs-down controls under Good result? to provide quality feedback.

Select New dub to create another version from the same source video. The existing speaker training will be reused.


Legacy SRT upload

Direct SRT upload is no longer enabled by default. It may be enabled by special request for qualifying Dialogue Replacement use cases.

Contact support@lipdub.ai with details about your intended workflow.

Processing taking longer than expected? If one processing, training, audio-generation, or final-generation action remains in progress for more than two hours, contact support@lipdub.ai. Technical failures are normally refunded automatically.


Last updated: 7/20/26, 5:21 PM