Translate and Lip-Sync a Video
Before you begin
A Translation project is intended for footage with one visible face.
Important: Use a Multi-speaker project whenever more than one face appears at any point in the video— even when only one person speaks.
Before starting:
- Review the video and audio guidance in Prepare Your Content.
- Use source audio with clear, understandable speech.
- Provide at least 10 seconds of speech when cloning the speaker’s voice. Thirty seconds or more is recommended.
- Confirm that you have permission to use and clone the speaker’s voice and likeness.
Important: The generated translation becomes the video’s final audio track. LipDub does not currently retain the source music, ambience, or sound effects in a Translation project.
Create a Translation project
- From Home, select Create Translation.
You can also open Projects, select Create Project, and choose Translation project. - Upload the video you want to translate.
- Choose whether to enable lip-sync:
- Turn on Enable lip sync? to translate the speech and synchronize it to the speaker’s mouth.
- Leave Enable lip sync? off to use Pure Dubbing. LipDub will add the translated audio without changing the original mouth movements.
- When lip-sync is enabled, select a lip-sync model.
Model | Recommended use |
|---|---|
Turbo | Fast previews. Turbo uses LipDub’s generic base model and does not train or create a reusable actor model. |
Flash | Faster training with good quality, particularly for social and everyday content. |
Premium | Longer training for high-resolution and higher-quality results. |
Ultra | The highest-quality option for film, television, and other premium productions. |
The estimated credit cost is shown in LipDub before you continue.
- Under What’s your video’s original language?, select the language spoken in the source video.
- Under What language do you want to translate into?, select the target language and available language style.
- Choose the voice option you want to use. To preserve the original speaker’s vocal identity, select Voice clone the speakers.
- Review any optional translation settings under Advanced Options.
- Select Upload & Translate.
LipDub will upload the video, transcribe the source speech, create the translation, generate the translated voice, and prepare the project for review. You can leave the page while processing continues.
Choose a language style
Some target languages offer more than one style. These styles use different voice models.
Language style | General guidance |
|---|---|
Style 1 | Usually provides stronger similarity to the original speaker’s voice. |
Style 2 | Usually provides a more native-sounding accent in the target language. |
This is a general guide rather than a strict rule. Results vary by speaker and language, so you may generate another version with a different style to compare them.
Both styles use LipDub’s market-leading translation and pacing system.
Add translation context and vocabulary
Open Advanced Options during project setup to provide additional guidance.
Add a context prompt
Use Add a context prompt to explain the video’s purpose, audience, setting, or preferred tone.
For example:
- A casual conversation between friends
- A formal company presentation
- A technical product demonstration
- A promotional video using friendly, conversational language
Keep the instruction focused on information that helps LipDub understand how the content should be translated.
Add custom vocabulary
Use Add custom vocabulary for names, brands, products, abbreviations, or specialized terms that may otherwise be transcribed incorrectly.
Enter terms as a comma-separated list, such as:
LipDub, Acme Cloud, Product X, CES
Custom vocabulary helps LipDub recognize the terms in the source speech. You can review and refine how those terms are translated later in the editor.
Review the translation
When the translation is ready, select Review your translation.
The Translation editor displays:
- The source transcript
- The translated transcript
- The duration of each source and translated segment
- Pacing guidance for each segment
- A video preview
- Separate Target Audio and Source Audio tracks on the timeline
Review the video from beginning to end before generating the final result.
Check the source transcript
Start with the source transcript. An incorrect source word can produce an incorrect translation.
For each segment:
- Play the source audio.
- Compare the audio with the source text.
- Correct any transcription errors.
- Confirm names, brands, numbers, and specialized terminology.
Check the translated transcript
Review the translated text for accuracy, tone, and meaning.
You can edit a translated segment directly when you need to:
- Correct a translation
- Preserve a brand or product name
- Change the tone
- Improve cultural relevance
- Adjust wording for pronunciation or pacing
After changing or regenerating a segment, replay its translated audio before continuing.
Resolve pacing warnings
LipDub compares the duration of each translated segment with the time available in the source video.
A segment marked Well paced should fit naturally within its available time.
When a segment is too fast or too slow, LipDub displays a pacing warning and may offer Rephrase?
Select Rephrase? to create another version that better fits the available time.
Note: Rephrasing is contextually aware. It may adapt the wording for the target language and culture rather than produce a strictly literal translation. Always review the new text to confirm that the intended meaning has been preserved.
LipDub will not overlap consecutive speech segments from the same speaker. If translated speech extends beyond the end of the source video, the audio is cut off at the end of the video. Resolve pacing warnings before generating the final result.
Review and adjust timing
Use the timeline at the bottom of the editor to compare the source and translated audio.
- Target Audio shows the generated translated speech.
- Source Audio shows the original speech.
- Select a segment to move the playhead to that part of the video.
- Replay the surrounding sections to check transitions and pauses.
- Reposition a target-audio segment when its dialogue needs to begin earlier or later.
Preserve intentional silence and natural spacing between phrases. After making timing changes, play through the transition rather than reviewing the edited segment in isolation.
Edit a transcript using a CSV
For larger transcript changes, select Download Transcript.
Always edit the CSV exported from the current Translation editor:
- Download the current transcript.
- Edit only the text in Column B.
- Do not change the column headers or index numbers.
- Save the file as a CSV.
- Select Upload Transcript to replace the transcript in the editor.
Warning: External transcript files may contain different segment boundaries. Changing the exported headers or index numbers will cause the upload to fail.
Complete transcript-file requirements are covered in Prepare Your Content.
Use Translation Memory
Translation Memory helps you maintain preferred wording for names, terminology, and recurring phrases.
You can create, edit, and delete entries from Translation Memory in the main navigation.
Translation Memory is currently applied during review:
- Open the Advanced Translation editor.
- Select the relevant translated segment.
- Review the suggested modification from Translation Memory.
- Approve the suggestion to apply it to that segment.
Translation Memory entries are not currently applied automatically to every translated segment. Review and approve them individually where relevant.
Generate the lip-synced video
When the translation, audio, pacing, and timing are ready, select Proceed to lip sync.
The Render settings window provides the following controls.
Set expressiveness
The Expressiveness slider controls how strongly the lips and jaw react to the generated speech.
Setting | Result |
|---|---|
1-2: Lower settings | Less pronounced mouth movement and less emphasis on each syllable |
3: Default | Balanced lip and jaw movement |
4-5: Higher settings | More open, expressive, and reactive lip and jaw movement |
Expressiveness is a visual preference. You can create another version at a different setting to compare the results.
Add captions
Turn on Add Captions to burn the translated transcript into the generated video.
Current captions use:
- White text
- A black outline
- A standard fixed caption style
Caption appearance cannot currently be customized.
Optimize AI-generated footage
Turn on Optimize for AI generated video when the source was created using an AI video tool.
This uses a version of LipDub’s model tuned for AI-generated footage. It is generally recommended for all AI source video, although you can generate versions with the setting on and off to compare them.
Leave this setting off for conventional human footage unless you are testing alternatives. It may reduce visual texture quality on real human video.
Start the generation
Review the credit estimate shown in the window, then select Generate.
LipDub displays the project’s status while the video is generating. You may leave the page and return later.
Review and download the result
Completed and in-progress translations appear under Your renders.
When a result is ready:
- Select the language version you want to review.
- Use the Original and Lipdubbed controls to compare the source and generated videos.
- Watch the complete result, including the beginning and end of each spoken segment.
- Select Download result to save the finished video.
Check:
- Translation accuracy
- Voice consistency
- Lip and jaw movement
- Segment timing
- Pacing
- Captions, when enabled
- Visual transitions around edits and jump cuts
To create another language or style from the same source project, select Create translation. You do not need to upload the source video again.
Processing taking longer than expected? If one processing, training, or generation action remains in progress for more than two hours, contact support@lipdub.ai.
Supported languages
LipDub currently supports the following languages. Some target languages provide more than one language style.
Afrikaans | English (Scottish) | Javanese | Punjabi |
Arabic | English (Welsh) | Kannada | Romanian |
Armenian | Estonian | Kazakh | Russian |
Assamese | Filipino | Kirghiz (Kyrgyz) | Serbian |
Azerbaijani | Finnish | Korean | Sindhi |
Belarusian | French (France) | Latvian | Slovak |
Bengali | French (Canadian) | Lingala | Slovenian |
Bosnian | Galician | Lithuanian | Somali |
Bulgarian | Georgian | Luxembourgish | Spanish (Mexican) |
Cambodian (Khmer) | German | Macedonian | Spanish (Spain) |
Catalan | Greek | Malay | Swahili |
Cebuano | Gujarati | Malayalam | Swedish |
Chichewa | Hausa | Marathi | Tamil |
Chinese (Cantonese) | Hebrew | Nepali | Telugu |
Chinese (Mandarin) | Hindi | Norwegian | Thai |
Croatian | Hungarian | Pashto | Turkish |
Czech | Icelandic | Persian (Farsi) | Ukrainian |
Danish | Indonesian | Polish | Urdu |
Dutch | Irish | Portuguese (Brazil) | Vietnamese |
English | Italian | Portuguese (Portugal) | Welsh |
English (Irish) | Japanese |