Prepare Your Content

The quality of your source files directly affects the quality of your LipDub. For the best results, use clear, consistent video with an unobstructed view of each speaker’s face, and provide clean, accurately timed audio where required. This guide covers the video, training footage, audio, transcript, and CSV requirements that apply before you begin a project.

General upload requirements

LipDub AI can process short clips as well as long-form content.

  • Maximum file size: 15 GB per file
  • Video duration: Approximately 10 seconds to more than 120 minutes
    • LipDub has successfully processesed video clips ranging from 0.5 seconds to 150 minutes.
  • Supported video files: MOV and MP4
  • Supported audio: WAV, MP3, AAC
    • You can also upload an MP4 or MOV as an audio source instead of extracting its audio first

Important: Use a Multi-speaker project whenever more than one face appears at any point in the video, even when only one person speaks.


Prepare your source video

Supported video specifications

Specification

Supported

File format

MOV and MP4

Video codec

H.264; Apple ProRes 422, 422 HQ, 4444, and 4444 XQ

Frame rate

Constant 23.976, 24, 25, 29.97, or 30 FPS

Color space

sRGB and Rec. 709

Resolution

SD through 4K, including 480p, 720p, 1080p, 2K, and 4K

Pixel aspect ratio

1:1

Bit depth

1-, 2-, 4-, and 8-bit are supported. Higher bit depths are accepted, but the face region is processed at 8-bit before being returned to the source bit depth.

LipDub returns the same file type you upload and preserves the source video’s aspect ratio, frame rate, and metadata.

When an output-quality option is available:

  • Standard produces output up to 1080p. A 4K source generated with Standard will produce a 1080p result.
  • Pro supports output up to 4K.

Files that require conversion

The following video types are not supported and should be converted before upload:

  • Variable Frame Rate footage
  • Interlaced footage
  • Anamorphic footage
  • Image sequences, including EXR sequences
  • Videos containing multiple video streams
  • Footage with a non-square pixel aspect ratio, such as 2:1

Important: Phone cameras commonly record using a Variable Frame Rate. Convert phone footage to a Constant Frame Rate before uploading it to LipDub.

HDR, Rec. 2020, and Rec. 2100 footage may be accepted, but LipDub cannot guarantee that the output colors will exactly match the source. For the most predictable color, use SDR footage in sRGB or Rec. 709.


Keep all footage consistent

Video specifications should remain consistent across the source footage and any supplemental footage used to train a speaker.

Where possible, match the following:

  • Resolution
  • Frame rate
  • Codec and file format
  • Color space
  • Color grade
  • Lighting and exposure
  • Camera angle and framing

Mixing substantially different footage increases the likelihood of visible artifacts. For example, avoid:

  • Combining ungraded and heavily color-graded footage
  • Mixing footage with noticeably different color casts
  • Using 1080p training footage for a subject shown in a 4K source video
  • Combining clips with very different lighting or camera angles

For a Multi-speaker project containing several scenes, keep the visual treatment as consistent as the production allows.


Frame each speaker clearly

For ideal source and training footage:

  • Use clear, well-lit footage.
  • Position the speaker primarily facing the camera.
  • Keep the full face, jaw, chin, mouth, and lips visible.
  • Avoid extreme close-ups. As a general guide, the face should occupy no more than approximately one-third of the frame.
  • Keep hair and other objects away from the mouth and jaw area.
  • Avoid framing where part of the face repeatedly leaves the screen.

Think of the ideal face position like a passport photo: the face is clearly visible, reasonably front-facing, and unobstructed.

Hair, hands, microphones, masks, props, or other objects covering the lips may make the result less reliable. Extreme close-ups can also take longer to process and may be more prone to visual glitches or artifacts.

Motion and edited footage

The subject and camera do not need to remain still. LipDub is designed to work with:

  • Moving speakers
  • Moving cameras
  • Different shot types
  • Edited footage
  • Jump cuts

Jump cut footage is supported, but translated phrases may have slightly different timing in each language. As a result, a cut placed tightly between words in the source may not feel identical in every translated version.

AI-generated and animated footage

LipDub supports AI-generated source footage. Select Optimize for AI generated video when generating this type of content.

This setting uses a version of the model tuned for AI source footage. It should generally remain off for real human footage because it can reduce visual texture quality.

LipDub may also work with some animated characters, although its strongest and most consistent results are produced with photorealistic human video.


Prepare supplemental training footage

Supplemental actor training is available in a Multi-speaker project. It allows you to provide additional footage of a person and share that training data across multiple videos or scenes within the project.

Training footage should match the source footage as closely as possible, including its lighting, resolution, frame rate, color, camera angle, and general appearance.

Most importantly, the footage must show the subject actively speaking with their lips visible. Footage of a person silently appearing on screen does not provide useful training data.

Use the following targets:

Tootal amount of visible speaking footage

Guidance

10 to 30 seconds

May work in some cases, but results are less predictable

30+ seconds

Practical minimum for dependable training

1 minute

Recommended for stronger results

Up to 5 minutes

Can provide additional coverage and quality

More than 5 minutes

Usually produces diminishing improvements

Some users record LipDub’s lip-shape script to capture a wider range of speech sounds and mouth shapes. This can help the model understand the individual nuances of how the subject speaks.

In a Multi-speaker project, you only need to label and train the faces you intend to lip-sync.


Prepare clean audio

Across every workflow, use the cleanest audio available. Noise, unintended dialogue, music, effects, and other material in your input will affect the output.

For best results:

  • Use clear speech with minimal background noise.
  • Avoid clipping, distortion, echo, and heavy processing.
  • Remove any dialogue you do not want included.
  • Keep timing and silence intact when the audio must align with an existing video.
  • Check that the audio starts at the intended timecode.

LipDub accepts most common audio file types. You can also upload an MP4 or MOV when the exact audio is already contained in a video and you do not want to extract it separately.

LipDub can synchronize supplied audio in any spoken language. It can also work with singing and fictional languages by generating the corresponding human mouth movements.

Audio for voice cloning

Provide at least 10 seconds of clear speech to clone a voice. For better results, provide 30 seconds or more.

Only clone a voice or likeness that you have permission to use. LipDub may restrict the cloning of recognized celebrity voices. Contact support@lipdub.ai if a permitted use case is being restricted.


Understand how audio is used in each project

Audio behavior differs by project type.

Project type

How audio is handled

Translation project

LipDub generates translated speech from the source video. Uploaded replacement audio is not currently supported. The translated audio becomes the final audio track, so the source music, ambience, and sound effects are not retained.

Pure Dubbing

Uses the translated audio from a Translation project without changing the subject’s mouth movements. The translated audio becomes the final audio track.

Personalization project

The original audio mix is preserved. LipDub changes only the selected personalized words or phrases.

Dialogue Replacement project

Uploaded or generated dialogue becomes the complete final audio track. Original music, sound effects, and ambience are not retained.

Multi-speaker project

The original dialogue is removed and replaced with the audio supplied for each trained speaker. Prepare one separate audio file per speaker.

Dialogue Replacement audio length

  • When the replacement audio is longer than the video, it is cut off at the end of the video.
  • When the replacement audio is shorter than the video, the video is trimmed to match the audio.

Multi-speaker audio preparation

For each speaker:

  • Upload one separate audio file.
  • Include only the dialogue intended for that person.
  • Preserve all silence in the file.
  • Place each line at the exact timecode where it should be heard.
  • Do not combine several speakers into one file.

Audio is still included when a speaker is off-screen or turned away from the camera.

When the audio is longer than the video, it is cut off at the end. When it is shorter, lip-sync stops after the supplied audio ends.

Important: In a Multi-speaker generation, Selected Regions controls where visual lip-sync is applied. It does not trim the uploaded audio. The speaker’s full audio file is still added to the generated video, including outside the selected regions. Selecting No Lip Sync makes that speaker silent.


Prepare an edited transcript

Translation projects include an editor for reviewing and changing the transcript.

When editing a transcript outside LipDub, always begin with the CSV exported from the Translation editor. Files created separately may contain different segments and cannot be matched reliably to the project.

Follow these rules:

  1. Download the current Transcript CSV from the editor.
  2. Edit only the text in Column B.
  3. Keep the existing column headers unchanged.
  4. Keep every index number unchanged.
  5. Save the edited file as a CSV.
  6. Upload it to overwrite the transcript in the editor.

Warning: Changing the column headers or index numbers will cause the upload to fail.

Direct SRT upload is no longer enabled by default. Legacy SRT upload may be made available for qualifying Dialogue Replacement use cases by contacting support@lipdub.ai.


Prepare a Personalization CSV

A Personalization project uses a CSV to create many versions of the same video.

Use a simple, unformatted CSV with:

  • One plain-text column header for each field
  • One recipient or output version per row
  • A required email column
  • A value in every cell
  • No blank rows or blank values

A CSV can contain up to 10 personalized variables. There is no campaign-size limit, and a single sheet may contain thousands of rows.

Example:

email

first_name

company

jamie@example.com

Jamie

Acme

alex@example.com

Alex

Northwind

taylor@example.com

Taylor

Contoso

The email address does not need to be valid or deliverable, but the column must exist and contain a value in every row. LipDub does not email recipients or send videos to these addresses. The field is used only to organize outputs and make exported video links easier to use in an external email platform.

For a name or word that requires a specific pronunciation, enter a phonetic spelling in the CSV value. Preview the result before generating the entire campaign.


Pre-upload checklist

Before starting a project, confirm that:

  • You selected a Multi-speaker project when multiple faces appear.
  • Each file is no larger than 15 GB.
  • The video uses MOV or MP4 with a supported codec.
  • The footage uses a supported Constant Frame Rate.
  • Source and training footage have consistent visual specifications.
  • Each face, mouth, jaw, and chin is clearly visible.
  • Supplemental footage shows the subject actively speaking.
  • Voice-cloning material contains at least 10 seconds of clear speech.
  • Multi-speaker audio is separated into one accurately timed file per speaker.
  • Your transcript was created from the CSV exported by LipDub.
  • Every Personalization CSV cell contains data.
  • You have permission to use and clone every voice and likeness in the project.


Last updated: 7/17/26, 10:11 PM