DocsCore Features

Studio

Assemble TTS, SFX, imported audio, and generated clips into one production workflow.

What you get
Mixed export (MP3, WAV, or MP4 video) • Saved project you can reopen • History entry
What it costs

Charged per block per the feature each block uses

  • Assembling and editing a project is free; generating a block bills at that feature's rate.
  • Exporting renders any block that has not been generated yet, so an export can spend credits.

What it looks like

Studio in SonicVox

Listen

Every clip below is the same sentence, generated by SonicVox — so what changes between them is the voice, not the writing.

SonicVox turns your script into natural speech — with the pacing, emphasis, and character you would expect from a studio recording.
  • Emma Carter — English

    A neutral English read, the kind most narration and product walkthroughs start from.

  • Marcus Grand — English

    A deeper English delivery — the same sentence, so you are comparing voice rather than writing.

  • So-young Kim — Korean voice

    A non-English voice reading the same line, to hear accent carry across the identical script.

Speech engines

Five engines, one API. Pick by what the job needs — the closest voice match, the widest language coverage, or the lowest latency.

Signature

Default

Most natural — best voice match

Our flagship engine. The most natural-sounding speech with the closest match to your selected voice. Best for narration, ads, and any premium read in 10 major languages.

Languages
10
Cloned voices
Yes
Streaming
No

en, zh, ja, ko, de, fr, ru, pt, es, it

Studio

Ultra-realistic, expressive

Our newest studio-grade engine — ultra-realistic clones with an expressiveness dial, in 23 languages. Great when you want maximum realism and a touch of drama.

Languages
23
Cloned voices
Yes
Streaming
No

ar, da, de, el, en, es, fi, fr, he, hi, it, ja, ko, ms, nl, no, pl, pt, ru, sv, sw, tr, zh

Expressive

Lively, dynamic delivery

An expressive engine with naturally dynamic intonation and pacing — a characterful alternative to Signature for narration and dramatic reads. Multilingual.

Languages
16
Cloned voices
Yes
Streaming
No
Emotion
Yes

de, el, en, es, fi, fr, hu, it, ja, ko, nl, pl, pt, ru, tr, zh

Multilingual

30+ languages, incl. Urdu, Hindi, Arabic

The broadest language coverage — 30+ languages and dialects, including ones the other engines can't speak. Renders a designed voice from its written description rather than from a recording, so it can't reproduce a cloned voice. Best for global content.

Languages
30+
Cloned voices
From a description
Streaming
No

Classic

Fast streaming clone

A fast, lightweight clone engine with low-latency streaming. Mirrors your reference's pacing. Best for quick drafts and real-time use.

Languages
5
Cloned voices
Yes
Streaming
Yes

en, zh, ja, ko, yue

Key facts

Streaming: Classic only
The low-latency endpoints are backed by one engine; the others render a whole take before returning.
Cloned voices: 4 of 5 engines
An engine that renders from a written description cannot reproduce a recording-based clone.

Overview

Studio is the composition layer for SonicVox. Use it to combine generated speech, cloned voices, sound effects, and uploads into a single export-ready project.

Quickstart

1

Create a project

Start from an empty project or create one from a script, podcast outline, or generated source text.

2

Add blocks

Mix text blocks, sound effects, and imported audio so each element becomes part of the final timeline.

3

Generate missing speech

Render text blocks into audio and keep iterating until the pacing feels right.

4

Export the final mix

Produce a download-ready file when the timeline is approved.

Who this is for

Ideal users

  • Creators producing multi-block narrated assets
  • Teams assembling reusable voice, SFX, and imported audio into one export
  • Operators who need a lightweight browser-based assembly workflow instead of a full DAW

Before you start

  • At least one generated or uploaded asset to place on the timeline
  • A default voice choice if the project includes several text blocks
  • A naming convention for projects so drafts and final exports stay organized

Settings

SettingDescriptionValues
Block orderControls sequence and timing within the project.
Default voiceSpeeds up creation when many text blocks share the same speaker.
Export formatDetermines the final delivery container for your mix.

Use cases

Podcast-style spoken segments

Mix narration blocks, transitions, and imported intros or outros into a single project without leaving the browser.

Marketing assets with layered sound design

Combine TTS, generated effects, and hand-picked audio to create trailers, product explainers, and social clips.

Operational assembly workflow

Use Studio as the last-mile editor after core teams have already generated approved clips in the individual feature pages.

Best practices

  • Generate voice assets first, then polish arrangement in Studio.
  • Use sound effects for transitions instead of overloading the spoken track.
  • Keep projects modular so sections can be swapped without rebuilding the entire mix.

Troubleshooting

The project feels messy after a few iterations

Likely cause

Teams often add new blocks faster than they rename, regroup, or delete temporary tests.

What to do

Treat Studio like an assembly layer: keep only approved blocks, use a clear project name, and remove throwaway experiments once a section is finalized.

Exports do not match the latest timeline changes

Likely cause

Text, voice, or order changes may be made faster than all blocks are regenerated and rechecked.

What to do

Before export, confirm all changed text blocks have current audio, scrub the whole sequence once, and save the project name so the final state is obvious.

FAQs

When should I create assets outside Studio first?

Create source assets in TTS, Voice Cloning, Sound Effects, or Enhancement first whenever those workflows need several quality passes. Bring only approved material into Studio.

What is the fastest way to start a new Studio project?

Set a clear project name, choose a default voice early, add your first paragraph or import the base audio, and only then branch into SFX and polish.

How is Studio different from the standalone tools?

Studio is a full workspace for voice projects — script, casting, multi-take direction, multi-track arrangement, review, and export. The standalone tools handle one-shot operations; Studio is where teams actually build content.

Can I import scripts from Final Draft or Word?

Partly. Import .docx, .pdf, .txt, or .epub through the Doc to Speech add-on, or paste plain text straight into the editor. Final Draft .fdx is not supported — export it to .docx or paste the text. Each paragraph becomes its own block; you cast voices per block yourself, starting from the project's default voice.

Can multiple people work on the same project?

Yes, inside a workspace. Share a project with a member or a group from Workspace settings → Resources at Viewer, Editor, or Admin. Editing is not simultaneous — there is no live cursor or real-time co-editing — so treat it as passing the project between people rather than working in it at the same moment.

What happens to my project if I re-cast a character?

Pick a new voice for the character and re-render — every line they speak gets a fresh take in the new voice. Old versions are kept in history.

Can I export to my DAW?

You can hand your DAW a finished mix. Export the project as MP3 or WAV — or as MP4 / WEBM once a video is attached — for the whole project or one chapter at a time. Per-character stems and OMF/AAF session files are not available.

Does Studio support multi-language scripts?

Studio renders the language your script is written in — filter the voice picker by language and cast a native-language voice per character. Studio has no script branching: to ship the same content in several languages, use Multilingual AI, which lays your assets out as an asset x language grid and renders each language from one place.

Are revisions audit-tracked?

Yes. Every change — script edits, cast changes, render approvals — is in version history with the editor's name and timestamp.

Is there an API for Studio?

Not yet. Studio projects are created, rendered, and exported in the app — there are no public REST endpoints for projects, renders, or exports. You can subscribe an account webhook to export events to be notified when a render finishes, and drive the underlying generation over the public API.

Detailed guide

Long-form notes, richer formatting, and implementation context for teams that need more than the quickstart.

Deep dive
Rich formatted reference
Use this section for implementation nuance, workflow depth, and operational guidance that does not fit in a simple checklist.

Studio Editor Guide

Create professional audio productions with SonicVox's multi-track studio editor.

Overview

The Studio Editor is a browser-based multi-track audio editor that combines:

  • Text-to-Speech blocks
  • Sound effect blocks
  • Imported audio files
  • Timeline-based editing

Getting Started

Create a Project

  1. Navigate to Studio in the sidebar
  2. On the New Studio Project panel, type a name into Project Name
  3. Click Create Studio Project (or press Enter)

Work you already started is under the Recent Projects tab beside it.

Your plan caps how many Studio projects you can keep at once. Past the cap, creating another one is refused with a message naming the limit and your plan — delete a project you no longer need, or upgrade.

Add Content Blocks

Text Blocks

  • A new project opens on an empty script. Click Add Paragraph to write the first one
  • Press Enter inside a paragraph to start the next one — it splits the text at your cursor
  • Pick a voice for the selected paragraph in the Edit tab of the right sidebar

SFX Blocks

  • Open the Sound Effects tab in the right sidebar
  • Search the library and click an effect to drop it in, or click Generate New Sound Effect, describe the sound, and Add to Timeline

Audio Import

  • Open the Imports tab in the right sidebar
  • Under Add Audio, choose Upload for a file on disk, Record to capture from your microphone, or Text to Speech / Voice Cloning to pull in something you already generated
  • The file picker accepts any audio format your browser recognises — WAV, MP3, M4A, FLAC and so on. A single upload is capped at 100 MB, and imports count against your plan's storage quota

Timeline Features

Block Arrangement

  • Drag blocks to reorder
  • Adjust start times
  • Trim durations

Playback

  • Click Play to preview
  • Scrub through timeline
  • Real-time waveform display

Edit Panel

When you select a paragraph, the Edit panel shows:

  • Selected Paragraph - A read-only preview of the text; edit the words in the paragraph itself
  • Volume - Adjust from 0% to 200%
  • Text Style - Body text or Heading 1-4
  • Voice - Change the speaker and the model
  • Speech Status - Generation progress, with play, download, and a Text toggle that shows the same read-only preview

Generating Speech

  1. Write your paragraphs and choose their voices
  2. Click Generate All in the toolbar to render every paragraph that has no audio yet — the button counts through them as it goes
  3. Or just press Play: playback generates each paragraph as it reaches it
  4. Preview a single paragraph from Speech Status in the Edit panel

Keyboard Shortcuts

KeyAction
SpacePlay/Pause (when you are not typing in a paragraph)
EnterSplit the paragraph at the cursor — this is how you add the next one
Backspace, at the very start of a paragraphMerge it into the previous paragraph
Delete or Backspace, with a timeline clip selectedDelete that clip
Ctrl+SSave the project name (your content saves itself)
?Show the full shortcut list

Export

Click Export to combine every track into one file.

  • Audio - MP3 (the default) or WAV
  • Video - MP4 or WEBM, offered once the project has a video attached, with an optional Burn Captions toggle
  • Export Scope - the whole project, or one chapter at a time once you have chapters
  • Optional metadata (title, author, ISBN) can be written into the file

Projects built on cloned or high-fidelity voices are combined as a master WAV regardless of the format you pick, so quality holds across paragraph seams.

Tips

  1. Write in short paragraphs - Easier to edit and regenerate
  2. Preview each block - Before exporting final mix
  3. Don't hunt for a save button - Paragraphs, timing, and block settings persist as you edit; the Save button and Ctrl+S only push the project name
  4. Use SFX for transitions - Adds polish between scenes
Was this page helpful?
Studio | SonicVox Docs | SonicVox Docs