You have a conversation with two people, but only want the guest's answer, your own contribution, or one person's voice for a new edit. Cutting up the mixed file by hand is awkward, especially when the voices overlap.
To isolate one voice from a recording, first separate the conversation into speaker tracks. Listen to the outputs to find the person you want, then download only that person's WAV file. Seply separates the speakers before you choose which voice to keep. You do not select a person by name or submit a voice sample before processing.
This workflow is for spoken conversations such as interviews and podcasts. It is not a method for separating two singers in a song.
Hear the original conversation and the voice to keep
Listen to the mix, then keep Maya’s voice
42-second synthetic interview · original dialogue voiced with licensed TTS · actual Seply outputs from September 6, 2026. These tracks were processed with Advanced Overlap, which requires a successful purchase.
- Reading the real waveform
- Reading the real waveform
- Reading the real waveform
The names are editorial labels for this known sample, not automatic identity recognition. Download Advanced — Maya to keep her track; the original pauses remain. This shared example demonstrates one recording, not a guaranteed result.
In this example, the goal is to keep Maya, the guest, rather than the interviewer. Compare Original mix with Advanced — Maya, then use that track's download control to get her WAV. The sample is an original dialogue voiced with licensed TTS and processed through Seply on September 6, 2026. It is not a customer recording or a claim that all conversations will separate this cleanly.
The complete interview test and source conditions show both processing modes and the overlap points. These are the same shared source and output files.
1. Separate the conversation before choosing a person
Open Seply, sign in, and upload a recording you own or have permission to use. If the speakers mostly take turns, start with Standard. Set the known speaker count or use automatic detection.
For people speaking at the same time, Advanced Overlap is the mode designed for that situation. It automatically detects speakers and requires a successful purchase. Check the mode and estimated credits before processing; it is not included in the free Standard trial.
If you need the full setup walkthrough, use how to separate two voices in one audio file. This guide focuses on what to do once the speaker tracks are available.
2. Find the voice you actually want
Do not assume that Speaker 1 is the person who matters to your edit. Track numbers are not verified identities.
Choose two or three places where you already know what the person says. For an interview, these might be an introduction, a clear answer in the middle, and a closing comment. Compare each place on the original and the candidate track. Confirm that it contains the same person throughout.
In the shared example, listen to Maya's answer after 00:33 and another earlier passage. Check the interviewer track as well if a word seems to be missing. The names in the demo were added for this known script; Seply is not identifying people by name.
3. Download just the selected speaker's WAV
Once you have confirmed the voice, use that track's download control. You do not need to download every speaker. In the demo above, Advanced — Maya downloads the unmodified guest output from the interview test.

Your own result is also a time-aligned WAV. That means it keeps the original timeline: if the person is silent for the first few seconds, the file may start with silence. A pause while someone else speaks can remain a pause rather than being removed automatically.
To create a shorter voice-only clip, trim those pauses in an audio editor after downloading. If you are putting the WAV under a video, keep its original start position until you have checked synchronization.
4. Check the parts where both people speak
Listen around interruptions and quiet words. Some of the other person's voice may remain, and some of the target person's speech may sound less clear. Compare with the original before deciding whether the result is usable.
Downloading one speaker does not preserve every background sound from the recording. Music, room sound, and incidental noises can change during separation. If those sounds matter to your finished edit, review the full mix in your editor rather than assuming that a speaker track contains them unchanged.
For sustained overlap, see how to separate two voices talking at the same time.
Common questions
Can I remove the interviewer and keep the guest?
Yes: choose the guest's track after separation. If you want to keep parts of both sides, import both WAVs into your editor and lower or mute the interviewer where needed. Follow the interview editing workflow for that use case. Some residual interviewer speech can still remain on the guest track.
Can I isolate my own voice from a group recording?
You can try the same workflow and identify your track by listening to known passages. Automatic detection can help with a changing group, but always check the outputs. Similar voices, distant speech, or frequent interruptions can make the result less reliable.
Can I choose the person before uploading?
No. This workflow separates the recording first. You then listen and choose a resulting track. It does not use a name, photograph, or reference voice sample to target a person before separation.
Is keeping one voice free?
Selecting and downloading a completed track does not change the processing rate. New accounts receive 30 trial credits, valid for 30 days, for up to about 3 minutes of Standard Separation. Advanced Overlap requires a successful purchase. The sample on this page uses Advanced; it is available to preview and download without processing your own recording.
Formats and limits
- Accepted input
- Standard: WAV, MP3, M4A, AAC, FLAC, OGG, OPUS, WEBM, WMA, SPX, MP4, AVI, MOV, MKV.
Advanced Overlap: WAV, FLAC, MP3, AAC, M4A, MP4, MOV. MP4 and MOV must contain AAC or PCM audio. - Output
- One time-aligned WAV track per detected speaker
- Upload limit
- Up to 1 GB; Advanced Overlap: up to 90 minutes
- Billing increment
- Standard Separation: 10 credits per started minute, 10-credit minimum.Advanced Overlap: 60 credits per started minute, 60-credit minimum.
The server-calculated duration is rounded to the nearest whole second before billing. Check the exact credit estimate before processing. Advanced Overlap requires a successful purchase; trial credits cover Standard Separation.