Speech Separation
Split multi-speaker recordings into separate tracks for analysis, cleanup, or downstream editing.
900 credits per job
- Flat per job. Inputs are capped at 20 minutes.
What it looks like

Overview
Speech Separation isolates speakers from mixed recordings so you can route them into cleaner editing, transcription, or restoration workflows.
Quickstart
Upload the mixed recording
Use a file that contains overlapping or alternating speakers.
Run separation
Generate the isolated tracks and wait for the service to finish processing.
Validate each track
Listen to the split outputs before using them in transcription or final editing.
Best practices
- Pair separation with STT when you need cleaner speaker-attributed transcripts.
- Keep the original mix so you can compare artifacts introduced by separation.
FAQs
How many speakers can it separate?
As many as are in the recording — every speaker gets their own track, with the count detected automatically. Accuracy is highest with up to 5 speakers; in larger groups, very similar voices may share a track.
Does it handle overlapping speech?
Yes. Two people talking at once are separated onto their own tracks. When three or more voices hit the exact same instant, that moment is kept in every involved speaker's track rather than guessed — and we tell you how much of the recording that affected.
Will the separated tracks sound natural?
Yes. Each speaker's track preserves their natural tone and timbre — no robotic artifacts.
Can I label speakers with custom names?
Yes. After separation, map auto-detected speakers to known names and the labels persist across exports.
What output formats are supported?
A separate WAV track per speaker, ready to play or download individually.
Can I stream live for real-time use?
Not yet. Separation runs on uploaded recordings today — up to 20 minutes / 25 MB for recordings with 3+ speakers — and processing is faster than real-time.
Does it work with single-mic recordings?
Yes. We can separate a single-mic recording of multiple speakers into per-speaker tracks.
Is there an API?
Not yet. Separation runs in the app today — upload a recording and download each speaker's track from the Speech Separation page. It is not part of the public API, so there is no way to submit recordings or fetch the tracks programmatically.