Skip to content

speaker separation · problem-solving

How to Separate Speakers in One Audio File

Separate speakers from one recording by checking the source, choosing automatic or manual speaker count, processing the file, and reviewing each WAV track.

Written by
Seply Editorial Team
Reviewed

To separate speakers in one audio file, upload an authorized recording, confirm the detected duration and credit estimate, choose automatic or manual speaker count, start processing, then listen to every time-aligned WAV track against the original mix. The review step matters: source quality and overlapping voices affect the result.

This workflow is designed for a recording that already contains two or more people in one mixed file—not for splitting music into vocals and instruments.

1. Inspect the source before uploading

Start with the least-processed version you are allowed to use. A high-quality original usually preserves more voice detail than a file that has been repeatedly compressed, noise-gated, or exported through messaging software.

Use this preflight check:

  • Confirm that at least two distinct people speak in the recording.
  • Make sure the file contains an audible audio track.
  • Prefer stable levels without clipping or aggressive background music.
  • Note sections with crosstalk, applause, echo, or off-mic speech for later review.
  • Confirm you own the recording or have permission to process it.

Seply accepts audio and compatible video files. For some videos, the browser extracts the original audio track without re-encoding before upload; otherwise the compatible container is uploaded as provided.

Formats and limits

Accepted input
WAV, MP3, M4A, AAC, FLAC, OGG, OPUS, WEBM, WMA, SPX, MP4, AVI, MOV, MKV
Output
One time-aligned WAV track per detected speaker
Upload limit
Up to 1 GB
Billing increment
1 credit per 6 seconds, 1-credit minimum

2. Choose automatic or manual speaker count

Use automatic detection when the number of voices is uncertain or when a short participant appears only once. Use a manual count when you know the recording has a fixed cast and want to constrain the result between 2 and 10 speakers.

Recording patternBetter starting choiceWhy
Host and one guestManual: 2The cast is known and stable
Group interview with drop-insAutomaticShort appearances can be easy to overlook
Panel with five named microphonesManual: 5Production notes already define the cast
Crowd, audience, or many incidental voicesAutomatic, then review carefully“Speaker” boundaries may not match editorial roles

A manual count is a constraint, not proof that every resulting track belongs perfectly to one person.

3. Review the estimate before starting

Seply reads the media duration in the browser and shows the required credits before upload. Speaker Separation bills 1 credit per started 6 seconds, with a 1-credit minimum. A full 12-minute recording therefore requires 120 credits. A 30-credit trial covers up to about 3 minutes.

The server-calculated duration is authoritative. Processing will not start if the available credit balance is below the reservation shown in the tool.

4. Let the job finish, then compare every track

Seply uploads the authorized source to private storage, submits the job to specialist processing infrastructure, and preserves the completed source and tracks privately for 30 days. You can leave the processing screen and reopen the job from account history.

When the result is ready, do three passes:

  1. Identity pass: verify that the same person remains on the same track through the recording.
  2. Overlap pass: jump to interruptions and simultaneous speech; compare each stem with the original.
  3. Artifact pass: listen for clipped words, voice leakage, metallic texture, or missing low-level speech.

The tracks remain aligned to the source timeline, so they can be placed together in an editor without manually rebuilding timing.

Shared Seply demo

Compare the mix with each speaker track

19 seconds · two speakers

0:00 / 0:19Selected Original mix

This is the same public demonstration used on the Seply homepage. It is evidence of one example, not a universal accuracy claim.

5. Export and edit conservatively

Download each separated speaker as WAV. In your audio or video editor, place the tracks at the same start time. Keep the original mix on a muted reference track so you can check edits.

A practical cleanup order is:

  1. Correct obvious routing or identity mistakes.
  2. Balance speaker levels.
  3. Reduce unwanted voice leakage only where it is distracting.
  4. Apply noise reduction and EQ after separation, using modest settings.
  5. Compare the finished passage against the original before export.

What to do when the first result is weak

Do not repeatedly apply heavy enhancement and expect lost speech to return. Instead, locate the cause:

SymptomLikely source issueUseful next action
Both voices appear in one trackSimilar timbre or dense overlapCheck the original overlap and try the known speaker count
Words disappearClipping, masking, or very low levelReturn to the earliest source export
Music leaks into tracksSpeech competes with a music bedUse a version without music if available
Room sound pulsesStrong reverb or aggressive prior denoisingReduce downstream gating; edit only affected sections

For crosstalk-specific decisions, continue with how to handle overlapping speech. If you need text rather than audio stems, compare speaker separation and diarization. For a field workflow, see interview speaker separation.

Private processing

Uploads and results are available only to the signed-in account.

30-day result access

Completed source media and outputs are retained privately for 30 days, or until deleted.

Specialist infrastructure

Seply uses third-party professional processing infrastructure and private object storage to complete jobs.

Upload only recordings you own or have permission to process.