Transcribe the source
FFmpeg prepares the audio and multilingual Whisper turns the original speech into timed subtitles.
A Finder-based workflow that transcribes source speech, adapts it into concise Russian and assembles a timed voice track into the final video.
Problem → solution
A literal Russian translation often becomes too long for the original speech slot. The pipeline adapts language, timing and generated voice as one coordinated process.
FFmpeg prepares the audio and multilingual Whisper turns the original speech into timed subtitles.
Russian lines preserve meaning, use a controlled glossary and shorten selectively when timing demands it.
Voice segments are timed, joined, mixed, normalised and attached to the final video.
Subtitle structure is validated, overlong cues receive targeted retries, and per-line audio caching avoids rebuilding finished segments.
Intended for material I own or have permission to adapt. The Russian track uses a synthetic AI voice and is presented as such.