F
Frank Houbre
Guest
I generated the same four-second line of dialogue nine times before I noticed what I was actually doing wrong. The line was right. The face was right, locked from a reference sheet I built for exactly this reason. What kept breaking was the timing: the model either cut the actor off half a word early, or it let the shot run two full seconds past the last syllable with nothing happening on screen. I kept blaming the model. The model was not the problem. I was generating video first and hoping the dialogue would fit whatever length I asked for.
AI video generation runs on fixed clip lengths, not on however long your line takes to say, so the fix is to lock the audio’s exact duration first and build the shot around that number. In 2026, a single generation from Veo 3.1 caps at 8 seconds in standard mode, Kling 3.0 and Seedance 2.0 cap around 15 seconds, Runway Gen-4 caps at 16, and Seedance 2.5 currently holds the record at roughly 30 seconds per clip, according to Higgsfield’s own 2026 breakdown of generation times across models. None of those numbers care how long your character takes to say I never asked you to come back. That’s on you to measure first.
Because the model has no idea what your character is about to say until it is already generating pixels. You type a duration, or you pick a clip-length tier, and the model fills that window with motion and, increasingly, with its own synthesized voice performance. If your real dialogue is faster or slower than what the model assumes, one of two things happens: the line gets clipped mid-word, or the shot pads itself with a beat of silence that reads as a mistake in the edit.
This is a timing problem, not a prompting problem. You can rewrite the prompt fifty different ways and it will not fix a shot that was never built around the actual length of the performance.
Prelay is the standard workflow in most Western animation: the voice actor records the dialogue first, in a booth, often across several takes, and animators build the mouth shapes, timing, and physical performance around that finished audio. Edge Studio’s glossary of voice-over production terms describes it plainly: the voice comes first specifically so the animation has something fixed to match.
Animators did not adopt prelay because it is elegant. They adopted it because animating blind, then trying to force a voice performance into a shot that was never built for it, produces worse timing every time. AI video generation has the exact same problem prelay solved ninety years ago, except now the animator is a model with a hard length cap, and every wasted attempt costs real money.
Here is the workflow I run now for any shot with a spoken line, built after that nine-take mistake on a Lost Garden scene:
The tooling for step four got considerably better in 2026. ByteDance’s Seed Audio 1.0, announced in July 2026, generates multi-character dialogue, sound effects, ambience, and music together in a single pass, with distinct voices per speaker and roughly two minutes of usable output per request. Feed that finished track into a compatible video model and the picture is generated to match a timeline that already exists, instead of the other way around.
I had already built the habit of generating video silent and adding sound afterward for most shots, and that is still the right call for anything without dialogue. But a two-character exchange in my animated series kept failing specifically because I was applying that same silent-first habit to a scene where timing was the content. Once I recorded the exchange as scratch audio first, timed it at just under 11 seconds, and generated the shot against that fixed length instead of guessing, it rendered clean on the second attempt instead of the tenth. The fix was not a better prompt. It was giving the model a number instead of a hope.
In the shot-planning stage of my AI-assisted production workflow, I now log a measured audio duration next to the shot description before generation starts, the same way earlier habits logged the model and seed used per shot. It is a small field, but it is the one that used to be a guess.
Does audio-first only matter for AI-generated voices?
No. It matters just as much if a human actor is recording the real performance. What changes is the reference: with AI voice generation you can regenerate a line quickly to test a different pace before locking it, which a live actor session does not allow as cheaply.
What if my video tool does not accept an audio reference at all?
You can still measure the audio’s exact length by hand and pick the closest clip-length tier your model offers, then trim dead air in editing afterward. It is less precise than a native audio reference, but it is still built around a real number instead of a guess.
Does this replace normal sound design and mixing?
No. Audio-first fixes pacing during generation. You still run a full sound pass afterward for music, effects, and final mix levels once the picture is locked.
How much time does this actually save?
On the Lost Garden shot above, it cut a nine-attempt problem down to two. Your mileage depends on how dialogue-dependent the scene is, but any shot where a spoken line drives the pacing is a candidate.
AI video generation runs on fixed clip lengths, not on however long your line takes to say, so the fix is to lock the audio’s exact duration first and build the shot around that number. In 2026, a single generation from Veo 3.1 caps at 8 seconds in standard mode, Kling 3.0 and Seedance 2.0 cap around 15 seconds, Runway Gen-4 caps at 16, and Seedance 2.5 currently holds the record at roughly 30 seconds per clip, according to Higgsfield’s own 2026 breakdown of generation times across models. None of those numbers care how long your character takes to say I never asked you to come back. That’s on you to measure first.
Why does an AI-generated dialogue shot keep running long or short?
Because the model has no idea what your character is about to say until it is already generating pixels. You type a duration, or you pick a clip-length tier, and the model fills that window with motion and, increasingly, with its own synthesized voice performance. If your real dialogue is faster or slower than what the model assumes, one of two things happens: the line gets clipped mid-word, or the shot pads itself with a beat of silence that reads as a mistake in the edit.
This is a timing problem, not a prompting problem. You can rewrite the prompt fifty different ways and it will not fix a shot that was never built around the actual length of the performance.
What is “prelay,” and why has traditional animation used it for decades?
Prelay is the standard workflow in most Western animation: the voice actor records the dialogue first, in a booth, often across several takes, and animators build the mouth shapes, timing, and physical performance around that finished audio. Edge Studio’s glossary of voice-over production terms describes it plainly: the voice comes first specifically so the animation has something fixed to match.
Animators did not adopt prelay because it is elegant. They adopted it because animating blind, then trying to force a voice performance into a shot that was never built for it, produces worse timing every time. AI video generation has the exact same problem prelay solved ninety years ago, except now the animator is a model with a hard length cap, and every wasted attempt costs real money.
The order of recording and animating changes the final result. Recording the performance first gives it room to set the pace. Recording after means squeezing a performance into a shape that was already decided without it.
How do you actually build an audio-first AI video shot?
Here is the workflow I run now for any shot with a spoken line, built after that nine-take mistake on a Lost Garden scene:
- Write the line, then say it out loud on a timer, at the pace your character would actually deliver it, not the pace you read silently in your head. Silent reading is almost always faster than performed dialogue.
- Record or generate the real audio track before touching a video tool. A rough voice memo works for pacing; a generated line from a voice tool works even better because you keep the exact take. Either way, you now have a file with a real, measured duration, not a guess.
- Match that duration to the clip-length tier that actually fits it, rounding up to the nearest option your model offers rather than trying to compress a 12-second line into an 8-second generation.
- Feed the audio in as a reference where the tool supports it, so the video generation performs to that exact track instead of guessing at pacing on its own.
- Generate the shot once, against the locked track, and only regenerate if something in the picture is wrong, not because the timing drifted, since the timing is no longer a variable.
The tooling for step four got considerably better in 2026. ByteDance’s Seed Audio 1.0, announced in July 2026, generates multi-character dialogue, sound effects, ambience, and music together in a single pass, with distinct voices per speaker and roughly two minutes of usable output per request. Feed that finished track into a compatible video model and the picture is generated to match a timeline that already exists, instead of the other way around.
What this actually fixed on AI-generated animation
I had already built the habit of generating video silent and adding sound afterward for most shots, and that is still the right call for anything without dialogue. But a two-character exchange in my animated series kept failing specifically because I was applying that same silent-first habit to a scene where timing was the content. Once I recorded the exchange as scratch audio first, timed it at just under 11 seconds, and generated the shot against that fixed length instead of guessing, it rendered clean on the second attempt instead of the tenth. The fix was not a better prompt. It was giving the model a number instead of a hope.
In the shot-planning stage of my AI-assisted production workflow, I now log a measured audio duration next to the shot description before generation starts, the same way earlier habits logged the model and seed used per shot. It is a small field, but it is the one that used to be a guess.
Common mistakes to avoid
- Estimating duration by reading the line silently. Performed dialogue almost always runs longer than silent reading, sometimes by 30% or more once pauses and emphasis are added.
- Picking the clip-length tier before you know the line’s real duration, then trimming the dialogue to fit the tier instead of the other way around.
- Assuming every shot needs this. Pure action, establishing shots, and anything without a spoken line do not need audio-first treatment; forcing it everywhere just slows you down for no benefit.
- Throwing away the scratch take. Even a rough voice-memo recording is useful reference data. Keep it attached to the shot, not just the final polished line.
FAQ
Does audio-first only matter for AI-generated voices?
No. It matters just as much if a human actor is recording the real performance. What changes is the reference: with AI voice generation you can regenerate a line quickly to test a different pace before locking it, which a live actor session does not allow as cheaply.
What if my video tool does not accept an audio reference at all?
You can still measure the audio’s exact length by hand and pick the closest clip-length tier your model offers, then trim dead air in editing afterward. It is less precise than a native audio reference, but it is still built around a real number instead of a guess.
Does this replace normal sound design and mixing?
No. Audio-first fixes pacing during generation. You still run a full sound pass afterward for music, effects, and final mix levels once the picture is locked.
How much time does this actually save?
On the Lost Garden shot above, it cut a nine-attempt problem down to two. Your mileage depends on how dialogue-dependent the scene is, but any shot where a spoken line drives the pacing is a candidate.