Skip to content

speaker separation · commercial

Interview Speaker Separation for Cleaner Editorial Audio

Separate interviewer and interviewee voices from one mixed recording, review overlap and identity, then use aligned WAV tracks for a controlled editorial mix.

Written by
Seply Editorial Team
Reviewed

Interview speaker separation creates independent, synchronized audio tracks for the interviewer and interviewee from a single mixed recording. It gives an editor local control over questions, answers, interruptions, and level differences while preserving the original timeline.

This is a recovery and editing workflow for journalism, documentary production, user research, oral history, and creator interviews. It does not establish who a person is and must not replace source verification.

Start by classifying the interview

Source situationRecommended action
Two clean local microphone tracks existEdit the originals; separation is unnecessary
Only one mixed remote-call file existsSeparate interviewer and interviewee, then review
Camera file contains both voicesUpload the compatible video or its original audio track
Interpreter or producer speaks occasionallyUse automatic speaker detection or set the known total
Public-space recording contains crowd voicesExpect extra tracks or leakage; test a short section first

The strongest source is normally the earliest recording with the least compression and processing.

The interview workflow

Prepare a review map

Write down the time ranges that matter most: the opening identification, key claims, emotional statements, interruptions, and any passage likely to be quoted. These become mandatory comparison points after processing.

Separate the voices

Choose a manual count for a stable two-person interview. Choose automatic detection if another participant enters briefly or the cast is uncertain. Seply accepts common audio and video files up to 1 GB and shows the duration-based credit reservation before upload.

Verify identity and completeness

Do not label a track from a single sentence. Check multiple known passages across the recording. Listen for identity switching, leakage from the other voice, and words that are present in the original but weak in the separated output.

Edit with the original beside you

Place every WAV at the same timeline start. Keep the source mix muted but immediately available. Use separated tracks for local level changes, question removal, and overlap control; return to the original whenever a phrase sounds uncertain.

Shared Seply demo

Compare the mix with each speaker track

19 seconds · two speakers

0:00 / 0:19Selected Original mix

This is the same public demonstration used on the Seply homepage. It is evidence of one example, not a universal accuracy claim.

Editorial choices by interview type

  • Journalism: preserve wording and context. If a processed word is uncertain, verify it from the original or with the source before quoting.
  • Documentary: use separation to build a cleaner mix, but maintain room tone and natural reactions where they support the scene.
  • UX research: separation can make facilitator prompts easier to reduce, while diarized transcription is often the better artifact for analysis.
  • Oral history: retain the untouched original as the archival record; treat separated tracks as access copies.
  • YouTube interviews: align the WAV tracks under the original video and make dialogue edits before loudness finishing.

Failure cases to plan for

If interruptions are the main problem, use the marker method in how to handle overlapping speech. For complete setup and review steps, see how to separate speakers in one audio file.

When the deliverable is text

Audio tracks are useful for editors; researchers and writers may instead need searchable text, speaker labels, timestamps, and captions. In that case, use the interview transcription workflow, or run both products from the same authorized source.

Formats and limits

Accepted input
WAV, MP3, M4A, AAC, FLAC, OGG, OPUS, WEBM, WMA, SPX, MP4, AVI, MOV, MKV
Output
One time-aligned WAV track per detected speaker
Upload limit
Up to 1 GB
Billing increment
1 credit per 6 seconds, 1-credit minimum

Private processing

Uploads and results are available only to the signed-in account.

30-day result access

Completed source media and outputs are retained privately for 30 days, or until deleted.

Specialist infrastructure

Seply uses third-party professional processing infrastructure and private object storage to complete jobs.

Upload only recordings you own or have permission to process.