Podcast speaker separation turns one mixed conversation into a time-aligned WAV track for each detected voice. It is most useful when an editor has only a stereo or mono mix and needs local control over a host, guest, or interruption—not when clean isolated microphone tracks already exist.
When this workflow earns its place
Use separation for a podcast when:
- a remote platform delivered only a mixed recording;
- a guest is consistently quieter than the host;
- interruptions or backchannels need local level control;
- one participant has distracting room spill;
- the video edit needs dialogue control but no multitrack audio survived.
Skip it when every microphone already has a clean, synchronized local track. Original isolated recordings contain more reliable detail than reconstructed stems.
The editor's handoff
| Input | Seply output | Editing use |
|---|---|---|
| One podcast mix | One WAV per detected speaker | Place all files at the same timeline start |
| Known two-person episode | Two constrained speaker tracks | Balance host and guest independently |
| Panel with uncertain drop-ins | Automatically detected tracks | Identify and label tracks during review |
| Original video file | Extracted or uploaded audio for processing | Return separated WAV tracks to the video timeline |
Seply keeps the output aligned with the source. The editor should also import the original mix on a muted reference lane.
Shared Seply demo
Compare the mix with each speaker track
19 seconds · two speakers
This is the same public demonstration used on the Seply homepage. It is evidence of one example, not a universal accuracy claim.
A production-ready sequence
1. Make a short diagnostic pass
Before processing the full episode, inspect one minute containing ordinary dialogue, one interruption, and the noisiest section. The 30-credit trial can cover up to about 3 minutes of Speaker Separation, which is enough for a representative test on many shows.
2. Process the least-damaged mix
Avoid a mastered file with a loud music bed if a pre-mix dialogue export exists. Upload the authorized source, select the known speaker count when the cast is fixed, and review the credit estimate. Speaker Separation bills 1 credit per started 6 seconds.
3. Label and spot-check tracks
Listen at the beginning, middle, and end. Then inspect every marked overlap. Confirm that identity does not move between tracks and that important words remain present.
4. Edit in context
Use clip gain before compression. Lower an interruption only for the affected phrase; do not hard-gate an entire track simply because another person is speaking. A little natural leakage may sound better than abrupt silence.
5. Retain evidence until delivery
Keep the original mix and your project file. Seply stores completed source media and speaker tracks privately for 30 days, or until you delete them earlier. Download the WAV files before that window ends.
What separation will not fix
Separation also does not generate show notes or captions. For text deliverables, use Speech to Text with diarization. The guide to speaker separation vs. diarization explains the difference.
The practical decision
If you can request isolated originals, do that first. If the mix is the only surviving source, use separation to create editing control, then make the smallest changes needed for clarity. Read how to handle overlapping speech before cutting interruptions, and use the step-by-step separation guide for file and review details.
Formats and limits
- Accepted input
- WAV, MP3, M4A, AAC, FLAC, OGG, OPUS, WEBM, WMA, SPX, MP4, AVI, MOV, MKV
- Output
- One time-aligned WAV track per detected speaker
- Upload limit
- Up to 1 GB
- Billing increment
- 1 credit per 6 seconds, 1-credit minimum
Private processing
Uploads and results are available only to the signed-in account.
30-day result access
Completed source media and outputs are retained privately for 30 days, or until deleted.
Specialist infrastructure
Seply uses third-party professional processing infrastructure and private object storage to complete jobs.
Upload only recordings you own or have permission to process.