DocsCore Features

Voice Cloning

Create reusable custom voices from short recordings and reuse them across the SonicVox stack.

What you get
WAV download • History entry • Reusable asset in your workspace
What it costs

900 credits per cloned voice you enroll

  • Charged once when the voice is created, not per use.
  • Speaking with a cloned voice afterwards is billed at the normal Text to Speech rate.
  • Requires a Starter plan or above; your plan also caps how many custom voices you can keep.

What it looks like

Voice Cloning in SonicVox

Listen

Every clip below is the same sentence, generated by SonicVox — so what changes between them is the voice, not the writing.

SonicVox turns your script into natural speech — with the pacing, emphasis, and character you would expect from a studio recording.
  • Emma Carter — English

    A neutral English read, the kind most narration and product walkthroughs start from.

  • Marcus Grand — English

    A deeper English delivery — the same sentence, so you are comparing voice rather than writing.

  • So-young Kim — Korean voice

    A non-English voice reading the same line, to hear accent carry across the identical script.

Speech engines

Five engines, one API. Pick by what the job needs — the closest voice match, the widest language coverage, or the lowest latency.

Signature

Default

Most natural — best voice match

Our flagship engine. The most natural-sounding speech with the closest match to your selected voice. Best for narration, ads, and any premium read in 10 major languages.

Languages
10
Cloned voices
Yes
Streaming
No

en, zh, ja, ko, de, fr, ru, pt, es, it

Studio

Ultra-realistic, expressive

Our newest studio-grade engine — ultra-realistic clones with an expressiveness dial, in 23 languages. Great when you want maximum realism and a touch of drama.

Languages
23
Cloned voices
Yes
Streaming
No

ar, da, de, el, en, es, fi, fr, he, hi, it, ja, ko, ms, nl, no, pl, pt, ru, sv, sw, tr, zh

Expressive

Lively, dynamic delivery

An expressive engine with naturally dynamic intonation and pacing — a characterful alternative to Signature for narration and dramatic reads. Multilingual.

Languages
16
Cloned voices
Yes
Streaming
No
Emotion
Yes

de, el, en, es, fi, fr, hu, it, ja, ko, nl, pl, pt, ru, tr, zh

Multilingual

30+ languages, incl. Urdu, Hindi, Arabic

The broadest language coverage — 30+ languages and dialects, including ones the other engines can't speak. Renders a designed voice from its written description rather than from a recording, so it can't reproduce a cloned voice. Best for global content.

Languages
30+
Cloned voices
From a description
Streaming
No

Classic

Fast streaming clone

A fast, lightweight clone engine with low-latency streaming. Mirrors your reference's pacing. Best for quick drafts and real-time use.

Languages
5
Cloned voices
Yes
Streaming
Yes

en, zh, ja, ko, yue

Key facts

Streaming: Classic only
The low-latency endpoints are backed by one engine; the others render a whole take before returning.
Cloned voices: 4 of 5 engines
An engine that renders from a written description cannot reproduce a recording-based clone.

Overview

Voice Cloning turns a sample recording into a reusable voice identity. Once a clone is ready, it can be selected in Text to Speech, Studio, and other generation flows that accept a voice ID.

Quickstart

1

Prepare a clean sample

Aim for about 10 seconds of a single speaker, with low background noise and consistent mic quality. Anything from 3 to 60 seconds is accepted, but 10 is where every engine is satisfied and nothing you recorded is discarded.

2

Upload or record

Create the clone from an uploaded file or capture the source directly in the browser if supported in your environment.

3

Name the voice

Use a durable naming pattern so the voice is easy to find later in dropdowns and project editors.

4

Use the clone downstream

After processing, select the cloned voice in TTS or Studio to generate with the new persona.

Who this is for

Ideal users

  • Teams standardizing a narrator or host voice across multiple projects
  • Creators building recurring character or persona-driven content
  • Agent builders who need a branded voice identity for live interactions

Before you start

  • A clean single-speaker recording of about 10 seconds (3–60 seconds accepted), with low room noise
  • Permission to clone and use the voice for the intended workflow
  • A naming convention for voices so the clone is easy to reuse later

Settings

SettingDescriptionValues
Sample qualityCleaner recordings improve identity retention and reduce artifacting.
Voice nameThis label becomes the main way the clone is selected elsewhere in the app.

Use cases

Brand voice consistency

Create one approved voice identity and use it across product walkthroughs, launch assets, and support content.

Character-based productions

Use cloned voices as stable characters for dialogue scenes, prototypes, and serialized content.

Live agent personality

Pair a cloned voice with the Agents Platform so the deployed experience sounds aligned with your product or brand.

Examples

Brand narrator

Clone a consistent host voice for tutorials and onboarding content.

Character prototype

Create a reusable test voice for interactive dialogue and concept videos.

Best practices

  • Keep raw source recordings archived so you can recreate higher-quality clones later.
  • Create internal naming conventions for ownership, language, and intended use.
  • Validate consent before cloning third-party voices.

Troubleshooting

The clone does not sound close enough to the source

Likely cause

Most quality issues come from noisy recordings, inconsistent mic distance, or source clips with multiple speakers.

What to do

Use a shorter but cleaner sample, trim silence, remove background noise first if needed, and keep the sample focused on one speaker only.

My team cannot tell which clone is the approved one

Likely cause

Clones were named casually during experiments and the same voice now appears under multiple ambiguous labels.

What to do

Adopt a durable naming scheme such as owner-language-purpose-version and retire draft clones that should not be used in production.

FAQs

How much audio do I need for a useful clone?

About 10 seconds. The upload accepts 3 to 60 seconds and up to 25MB, but the Classic engine reads only the first 10 seconds of a reference, so a longer clip buys you nothing there. A short clean sample beats a long noisy one — prioritise clarity, one speaker, and a stable recording environment.

Where can I use a cloned voice after it is created?

Use the clone in Text to Speech, Studio, and other SonicVox workflows that expose the reusable voice selector.

How long does cloning take?

Seconds. Zero-shot cloning means there's no model training step — drop in 10 seconds of clean audio and the clone is ready for production immediately.

How much reference audio do I need?

10 seconds of clean speech is enough for a high-quality clone. Longer clips are accepted, but the Classic engine only reads the first 10 seconds, so extra length is discarded rather than used.

Does the clone work in other languages?

Yes — in 23 languages (25 counting Cantonese and Hungarian), your clone speaks with your own tone, cadence, and identity, from a single recording. Beyond those, our Multilingual engine covers 30+ more languages, but it renders a voice from a written description rather than from your recording, so it doesn't carry your identity — we tell you in the studio rather than swapping voices on you.

Who owns the cloned voice?

You do. Voices you clone from your own consented recordings are yours to use commercially under our standard terms. SonicVox never uses your voice for training without explicit, written consent.

Is consent required to clone someone's voice?

Yes. You must have the speaker's explicit, written consent before cloning their voice. SonicVox enforces consent verification for production use, and we ban accounts that violate it.

How do you prevent voice fraud?

Generated audio is watermarked, identity verification is required for clone creation, and consent records are stored with each voice. We work with regulators on emerging anti-fraud standards.

Can I clone multiple voices?

Yes. Manage a library of clones, brand voices, and characters. Role-based access controls who can use which voice, with full audit trails.

What audio quality should the reference be?

Clean speech with minimal background noise — laptop mic in a quiet room is fine. We auto-detect and warn about poor source quality before you commit to a clone.

Detailed guide

Long-form notes, richer formatting, and implementation context for teams that need more than the quickstart.

Deep dive
Rich formatted reference
Use this section for implementation nuance, workflow depth, and operational guidance that does not fit in a simple checklist.

Voice Cloning Guide

Learn how to clone voices in SonicVox to create personalized text-to-speech.

What is Voice Cloning?

Voice cloning creates a reusable digital voice identity from a short reference sample.

Once the clone is ready, you can use it across SonicVox workflows such as Text to Speech, Studio, and other voice-enabled tools that accept a voice ID.

Start with the cleanest sample you have. A short, high-quality recording usually performs better than a long noisy one.

Requirements

For best results, prepare your sample before you upload it.

RequirementWhy it matters
About 10 seconds of speechThe number comes from the engines, not from taste. Signature, the default, characterises a voice from about three seconds. Classic accepts a reference of 3–10 seconds and uses only the first 10 seconds of anything longer. So 10 is the point where every engine is satisfied and nothing you recorded is thrown away. Both the upload form and the API (POST /api/v1/voices) accept 3–60 seconds and up to 25MB — a 60-second sample is legal, but on Classic 50 of those seconds are discarded before the model sees them.
Clear speechCleaner audio improves identity retention and reduces artifacts.
One speaker onlyMultiple voices confuse the clone and lower quality.
WAV, MP3, FLAC, OGG, M4AThese are the formats the Clone Voice screen advertises. The file picker also accepts AAC and WebM, which is what browser recordings produce.

Before uploading, make sure the sample:

  • has low background noise
  • does not include music or crowd ambience
  • keeps a steady speaking tone
  • represents the kind of delivery you want later in generation

How to Clone a Voice

Follow this workflow from start to finish for the cleanest result.

  1. Navigate to Voice Cloning

Voice Cloning is its own entry in the sidebar, under Voice Generation. It opens at /app/voice-cloning/generate.

  1. Add your sample

The Voice Sample card has two tabs, Upload and Record. On Upload (the default) the panel reads Drop your audio file here / or click to browse files, so you can either:

  • click the drop area and choose a file
  • drag and drop your recording onto it
  • switch to the Record tab and capture about 10 seconds straight from your microphone
  1. Name the cloned voice

Fill in Voice Name with a clear label such as John's Voice, Narrator Style, or a team naming convention like Brand Host EN v1.

  1. Enter the text to speak

Text to Speak is required, not optional — the clone is created and immediately read back to you with that text. Leave it empty (or leave the voice unnamed, or the sample missing) and the button stays greyed out.

  1. Create the clone

Click Clone Voice. The screen shows Processing your voice clone… and warns This may take 30-60 seconds. Creating a clone spends credits from your balance — see What it costs on this page; speaking with the finished voice afterwards is billed at the normal Text to Speech rate.

  1. Validate the result

Use the new voice in a short Text to Speech test before committing it to a larger production workflow.

Using Cloned Voices

After the clone is created, it becomes available in multiple places inside SonicVox.

You can use it in:

  • the voice selector in Text to Speech
  • Studio Editor voice options for scripted blocks
  • API-driven workflows that reference the voice ID
  • agent and automation flows where a reusable brand voice is needed

Tips for Best Quality

Use these quality rules before you scale generation.

  1. Use professional or clean recordings whenever possible.
  2. Avoid music, effects, and room noise in the background.
  3. Choose a natural speaking tone rather than shouting, whispering, or performing exaggerated delivery.
  4. Test with short phrases first before generating a full script.
  5. Keep naming consistent so your team knows which clone is approved for production.

If the first result is not good enough, do not keep retrying the same noisy sample. Replace the source clip with a cleaner recording and create a fresh clone.

Limitations

  • Voice cloning needs a Starter plan or above. The Free plan cannot clone at all — the feature is gated, not merely limited.
  • Your plan caps how many custom voices you can keep: Starter 10, Creator 30, Pro 160, Scale 660, Business 660, Enterprise 2,000.
  • Voice clones are private to your account unless your broader product workflow exposes them elsewhere
  • Some accents, delivery styles, or recording conditions may need multiple attempts for the best result

When quality matters, treat the source sample as the foundation. A better recording usually improves the outcome more than repeated retries.

Was this page helpful?
Voice Cloning | SonicVox Docs | SonicVox Docs