Interview speaker separation creates independent, synchronized audio tracks for the interviewer and interviewee from a single mixed recording. It gives an editor local control over questions, answers, interruptions, and level differences while preserving the original timeline.
This is a recovery and editing workflow for journalism, documentary production, user research, oral history, and creator interviews. It does not establish who a person is and must not replace source verification.
Start by classifying the interview
| Source situation | Recommended action |
|---|---|
| Two clean local microphone tracks exist | Edit the originals; separation is unnecessary |
| Only one mixed remote-call file exists | Separate interviewer and interviewee, then review |
| Camera file contains both voices | Upload the compatible video or its original audio track |
| Interpreter or producer speaks occasionally | Use automatic speaker detection or set the known total |
| Public-space recording contains crowd voices | Expect extra tracks or leakage; test a short section first |
The strongest source is normally the earliest recording with the least compression and processing.
The interview workflow
Prepare a review map
Write down the time ranges that matter most: the opening identification, key claims, emotional statements, interruptions, and any passage likely to be quoted. These become mandatory comparison points after processing.
Separate the voices
Choose a manual count for a stable two-person interview. Choose automatic detection if another participant enters briefly or the cast is uncertain. Seply accepts common audio and video files up to 1 GB and shows the duration-based credit reservation before upload.
Verify identity and completeness
Do not label a track from a single sentence. Check multiple known passages across the recording. Listen for identity switching, leakage from the other voice, and words that are present in the original but weak in the separated output.
Edit with the original beside you
Place every WAV at the same timeline start. Keep the source mix muted but immediately available. Use separated tracks for local level changes, question removal, and overlap control; return to the original whenever a phrase sounds uncertain.
Shared Seply demo
Compare the mix with each speaker track
19 seconds · two speakers
This is the same public demonstration used on the Seply homepage. It is evidence of one example, not a universal accuracy claim.
Editorial choices by interview type
- Journalism: preserve wording and context. If a processed word is uncertain, verify it from the original or with the source before quoting.
- Documentary: use separation to build a cleaner mix, but maintain room tone and natural reactions where they support the scene.
- UX research: separation can make facilitator prompts easier to reduce, while diarized transcription is often the better artifact for analysis.
- Oral history: retain the untouched original as the archival record; treat separated tracks as access copies.
- YouTube interviews: align the WAV tracks under the original video and make dialogue edits before loudness finishing.
Failure cases to plan for
If interruptions are the main problem, use the marker method in how to handle overlapping speech. For complete setup and review steps, see how to separate speakers in one audio file.
When the deliverable is text
Audio tracks are useful for editors; researchers and writers may instead need searchable text, speaker labels, timestamps, and captions. In that case, use the interview transcription workflow, or run both products from the same authorized source.
Formats and limits
- Accepted input
- WAV, MP3, M4A, AAC, FLAC, OGG, OPUS, WEBM, WMA, SPX, MP4, AVI, MOV, MKV
- Output
- One time-aligned WAV track per detected speaker
- Upload limit
- Up to 1 GB
- Billing increment
- 1 credit per 6 seconds, 1-credit minimum
Private processing
Uploads and results are available only to the signed-in account.
30-day result access
Completed source media and outputs are retained privately for 30 days, or until deleted.
Specialist infrastructure
Seply uses third-party professional processing infrastructure and private object storage to complete jobs.
Upload only recordings you own or have permission to process.