DocsCore Features

Dubbing

Create dubbed outputs by combining translation, voice generation, and timing-aware delivery.

What you get
Dubbed video master • Optional burned-in subtitles • History entry
What it costs

1,800 credits per minute of source videoA 10-minute video costs 18,000 credits.

  • The most expensive thing on the platform — check the length before you submit.
  • Paid plans include dubbing minutes; credits are only spent once those are used up.
  • Not available on the Free plan.

What it looks like

Dubbing in SonicVox

Listen

Every clip below is the same sentence, generated by SonicVox — so what changes between them is the voice, not the writing.

SonicVox turns your script into natural speech — with the pacing, emphasis, and character you would expect from a studio recording.
  • Emma Carter — English

    A neutral English read, the kind most narration and product walkthroughs start from.

  • Marcus Grand — English

    A deeper English delivery — the same sentence, so you are comparing voice rather than writing.

  • So-young Kim — Korean voice

    A non-English voice reading the same line, to hear accent carry across the identical script.

Speech engines

Five engines, one API. Pick by what the job needs — the closest voice match, the widest language coverage, or the lowest latency.

Signature

Default

Most natural — best voice match

Our flagship engine. The most natural-sounding speech with the closest match to your selected voice. Best for narration, ads, and any premium read in 10 major languages.

Languages
10
Cloned voices
Yes
Streaming
No

en, zh, ja, ko, de, fr, ru, pt, es, it

Studio

Ultra-realistic, expressive

Our newest studio-grade engine — ultra-realistic clones with an expressiveness dial, in 23 languages. Great when you want maximum realism and a touch of drama.

Languages
23
Cloned voices
Yes
Streaming
No

ar, da, de, el, en, es, fi, fr, he, hi, it, ja, ko, ms, nl, no, pl, pt, ru, sv, sw, tr, zh

Expressive

Lively, dynamic delivery

An expressive engine with naturally dynamic intonation and pacing — a characterful alternative to Signature for narration and dramatic reads. Multilingual.

Languages
16
Cloned voices
Yes
Streaming
No
Emotion
Yes

de, el, en, es, fi, fr, hu, it, ja, ko, nl, pl, pt, ru, tr, zh

Multilingual

30+ languages, incl. Urdu, Hindi, Arabic

The broadest language coverage — 30+ languages and dialects, including ones the other engines can't speak. Renders a designed voice from its written description rather than from a recording, so it can't reproduce a cloned voice. Best for global content.

Languages
30+
Cloned voices
From a description
Streaming
No

Classic

Fast streaming clone

A fast, lightweight clone engine with low-latency streaming. Mirrors your reference's pacing. Best for quick drafts and real-time use.

Languages
5
Cloned voices
Yes
Streaming
Yes

en, zh, ja, ko, yue

Key facts

Streaming: Classic only
The low-latency endpoints are backed by one engine; the others render a whole take before returning.
Cloned voices: 4 of 5 engines
An engine that renders from a written description cannot reproduce a recording-based clone.

Overview

Dubbing is the workflow for replacing or adding spoken audio to existing media. It typically combines transcription, translation, voice generation, review, transcript export, and editable resource operations before final export.

Quickstart

1

Import the source

Start from the original media or transcript you want to dub.

2

Prepare language and voice choices

Select the target language and the destination voice or voice profile.

3

Generate dubbed output

Create the replacement speech and review sync, pacing, and intelligibility.

Best practices

  • Use Studio when the dubbing job needs manual assembly or layered sound design.
  • Check translated scripts for sentence length before generation to avoid pacing issues.

FAQs

How accurate is the lip-sync?

There is no lip-sync. Dubbing replaces the audio track and leaves the original video frames untouched, so mouth movement still follows the source performance. It suits voice-over, narration, and talking-head content more than tight close-ups.

Does the dubbed voice sound like the original talent?

Yes. We clone the on-camera speaker's voice from the source audio and use that voice for every dubbed language. Audiences hear the same person.

Which video formats are supported?

MP4, MOV, AVI and MKV for video, plus audio-only MP3, WAV, FLAC and M4A. WebM works when you paste a remote video URL, but direct WebM uploads are rejected.

Can I review the translation before it renders?

After the first render, yes. The side-by-side editor opens once a job has produced a result: adjust the translated lines and re-render from there. There is no pre-render review step.

Are subtitles included?

Yes. Frame-accurate SRT/VTT in every dubbed language, plus optional burned-in subtitles in the rendered video.

How long does dubbing take?

Hours per video for typical pieces — same-day for shorts. Batch hundreds overnight.

Can I dub into multiple languages at once?

One language per run. Pick your dub language and render it, then add further languages one at a time from the job afterwards — each one renders separately.

Is consent required to use someone's voice?

Yes. The on-camera talent must have given written consent for their voice to be cloned and used in dubbed productions.

Was this page helpful?
Dubbing | SonicVox Docs | SonicVox Docs