Guide

Multi-Language Voiceover Versions: One Script, Editable Translations, Per-Language Narration (2026)

Turn one narration script into English, Chinese, Spanish, and Japanese versions: editable translation drafts, one narration track per language, and one-language-at-a-time preview and export.

Beat2Cut TeamAugust 13, 20264 min read

If the same video needs to reach viewers in several languages, subtitles are the cheap answer — but what audiences actually prefer is narration they can listen to in their own language. Beat2Cut turns one narration script into per-language voiceover versions: the script is translated into an editable draft per language (English, Chinese, Spanish, and Japanese are in the current catalog), you review and correct each draft, and only then is a narration track generated for each language. Preview and export carry one selected language at a time — the same picture, a different voice.

The load-bearing word in that paragraph is editable. Here is why that constraint exists, and exactly how the workflow runs.

Multi-language subtitles are not multi-language narration

Subtitles ask the viewer to read while watching. That works for some content — and fails for exactly the formats where narration matters most: tutorials where the viewer's eyes are on the demonstration, ambient content watched half-attentively, anything consumed on a phone at arm's length.

A narrated version, by contrast, is a separate deliverable per language: same picture timeline, same music, but the voice track speaks Spanish in one export and Japanese in another. Until recently that meant hiring voice talent per language or maintaining parallel project files. The economics changed once script-to-voice generation became reliable — what did not change is translation quality, which is why the review step cannot be skipped.

Why "auto-dubbing" keeps getting pushback

Platforms have started auto-dubbing videos for viewers, and creator complaints about it are remarkably consistent — three failures, over and over:

  1. Mistranslations go live with a voice attached. A wrong subtitle is embarrassing; a wrong spoken sentence sounds authoritative and is much worse.
  2. It is on by default. Viewers hear a synthetic dub the creator never approved, sometimes without an obvious way back to the original.
  3. The creator cannot fix the text. The pipeline goes straight from machine translation to synthetic speech, with no human hands in between.

Beat2Cut's version of this feature makes the opposite trade: the translated draft is a mandatory, visible, editable stop. There is no button that turns a machine translation into audio without showing it to you first. If that costs a few minutes of review per language, that is the feature working as intended.

How it works, step by step

The whole flow lives in the Script pane of the Studio timeline — it is not a separate product or page.

1. Write the source script once

Write or paste your narration next to the shots as usual. If you have already generated the source-language narration, nothing about it changes.

2. Pick target languages

Choose from the catalog languages — currently English, Chinese, Spanish, and Japanese (the source language itself is generated the normal way, so it is not offered as a target). Only languages that actually have voices in the catalog are offered.

3. Get a quote before anything runs

Before translation starts you see what it will cost in credits: translation is priced by script length × number of target languages, and the later speech step is estimated alongside it. If your balance will not cover it, the run is blocked up front rather than stopping half-done. Failed translations are not charged.

4. Review and edit each draft

Each language comes back as an editable text field. Fix terminology, shorten sentences that ran long, localize an idiom the machine translated literally. Drafts are saved with the project in your browser, so a page refresh does not lose your edits. An empty draft cannot be sent to speech — every language you generate is text you have seen.

5. Generate one narration track per language

On confirm, each language runs through the same narration engine as regular generation — the speech itself is synthesized by cloud neural voice models, it requires sign-in, and it is billed on the text after your edits. Languages run one at a time; each one lands on its own speaker track (the Japanese track is labeled 日本語, the Spanish one Español). The source-language clips are never touched.

6. Preview and export one language at a time

A version selector in the top bar picks the active language. The preview plays only that language's narration; the finished MP4, the WAV/MP3 audio mixdown, and SRT/VTT subtitles all follow the same selection. Export once per language you need — with multiple versions in a project, the export filename is tagged with the language (.es, .ja) so deliverables do not overwrite each other.

What it costs, in plain terms

StepWhat you pay
Beat detection & timeline editingFree & unlimited (after sign-in)
Scene detection, silence removal, cut-list delivery2 credits on Free, included with Pro
Translation into draftsCredits, by script length × number of target languages — quoted first
Narration per languageCredits, same as regular narration, on your edited text
SRT/VTT subtitles, WAV/MP3 mixdownFree
Finished MP4 per language version10 credits each on Free; included with Pro

Free accounts receive 100 credits a month, refilled on the 1st (non-accumulating), enough to translate and voice a short script. Pro provides 10,000 credits every 5 hours with unlimited detection and deliveries.

What v1 deliberately does not do

Honesty about scope beats a feature list that overpromises:

  • No lip-sync dubbing — this replaces a narration track, it does not re-animate a speaking face or duck an on-camera voice underneath a dub.
  • No automatic transcription of the source script from the video — the script is something you write (that is also what makes it worth translating).
  • No one-click bundle export of all languages — you select a language and export it, then switch and export the next.
  • No picking the speech vendor per language — you choose a voice and a language, not a provider console.

FAQ

Which languages are supported?

The translation catalog currently targets English, Chinese, Spanish, and Japanese — the languages with voices available in the catalog. The voice list itself is filterable by language, and the catalog grows as voices are added.

Can I skip reviewing the translation?

No, and that is intentional. The draft is the safety mechanism: only text that exists in an editable field — text you could have read and changed — can be sent for speech generation.

What happens if I cancel mid-generation?

Cancelling stops the narration job currently running (like cancelling a regular generation, work already synthesized may still be billed per the normal rules). Languages that have not started yet are simply not run. Failed translations are never charged.

Does my video get uploaded for this?

No. The picture, music, and editing stay in the browser. What goes to the cloud is the script text — for translation and for speech synthesis by the third-party providers — plus your generation settings; generated audio is stored with your account so your project can reload it.


Publish in more than one language: open the Studio script pane, write the script once, and let each market hear it natively — with your edits, not the machine's first guess.

#voiceover#translation#multilingual#AI narration

Ready to Try Beat2Cut?

Create beat-synced cuts, auto-split your footage into scenes, and add AI narration — all on one timeline in your browser.