To separate speakers in one audio file, upload an authorized recording, confirm the detected duration and credit estimate, choose automatic or manual speaker count, start processing, then listen to every time-aligned WAV track against the original mix. The review step matters: source quality and overlapping voices affect the result.
This workflow is designed for a recording that already contains two or more people in one mixed file—not for splitting music into vocals and instruments.
1. Inspect the source before uploading
Start with the least-processed version you are allowed to use. A high-quality original usually preserves more voice detail than a file that has been repeatedly compressed, noise-gated, or exported through messaging software.
Use this preflight check:
- Confirm that at least two distinct people speak in the recording.
- Make sure the file contains an audible audio track.
- Prefer stable levels without clipping or aggressive background music.
- Note sections with crosstalk, applause, echo, or off-mic speech for later review.
- Confirm you own the recording or have permission to process it.
Seply accepts audio and compatible video files. For some videos, the browser extracts the original audio track without re-encoding before upload; otherwise the compatible container is uploaded as provided.
Formats and limits
- Accepted input
- WAV, MP3, M4A, AAC, FLAC, OGG, OPUS, WEBM, WMA, SPX, MP4, AVI, MOV, MKV
- Output
- One time-aligned WAV track per detected speaker
- Upload limit
- Up to 1 GB
- Billing increment
- 1 credit per 6 seconds, 1-credit minimum
2. Choose automatic or manual speaker count
Use automatic detection when the number of voices is uncertain or when a short participant appears only once. Use a manual count when you know the recording has a fixed cast and want to constrain the result between 2 and 10 speakers.
| Recording pattern | Better starting choice | Why |
|---|---|---|
| Host and one guest | Manual: 2 | The cast is known and stable |
| Group interview with drop-ins | Automatic | Short appearances can be easy to overlook |
| Panel with five named microphones | Manual: 5 | Production notes already define the cast |
| Crowd, audience, or many incidental voices | Automatic, then review carefully | “Speaker” boundaries may not match editorial roles |
A manual count is a constraint, not proof that every resulting track belongs perfectly to one person.
3. Review the estimate before starting
Seply reads the media duration in the browser and shows the required credits before upload. Speaker Separation bills 1 credit per started 6 seconds, with a 1-credit minimum. A full 12-minute recording therefore requires 120 credits. A 30-credit trial covers up to about 3 minutes.
The server-calculated duration is authoritative. Processing will not start if the available credit balance is below the reservation shown in the tool.
4. Let the job finish, then compare every track
Seply uploads the authorized source to private storage, submits the job to specialist processing infrastructure, and preserves the completed source and tracks privately for 30 days. You can leave the processing screen and reopen the job from account history.
When the result is ready, do three passes:
- Identity pass: verify that the same person remains on the same track through the recording.
- Overlap pass: jump to interruptions and simultaneous speech; compare each stem with the original.
- Artifact pass: listen for clipped words, voice leakage, metallic texture, or missing low-level speech.
The tracks remain aligned to the source timeline, so they can be placed together in an editor without manually rebuilding timing.
Shared Seply demo
Compare the mix with each speaker track
19 seconds · two speakers
This is the same public demonstration used on the Seply homepage. It is evidence of one example, not a universal accuracy claim.
5. Export and edit conservatively
Download each separated speaker as WAV. In your audio or video editor, place the tracks at the same start time. Keep the original mix on a muted reference track so you can check edits.
A practical cleanup order is:
- Correct obvious routing or identity mistakes.
- Balance speaker levels.
- Reduce unwanted voice leakage only where it is distracting.
- Apply noise reduction and EQ after separation, using modest settings.
- Compare the finished passage against the original before export.
What to do when the first result is weak
Do not repeatedly apply heavy enhancement and expect lost speech to return. Instead, locate the cause:
| Symptom | Likely source issue | Useful next action |
|---|---|---|
| Both voices appear in one track | Similar timbre or dense overlap | Check the original overlap and try the known speaker count |
| Words disappear | Clipping, masking, or very low level | Return to the earliest source export |
| Music leaks into tracks | Speech competes with a music bed | Use a version without music if available |
| Room sound pulses | Strong reverb or aggressive prior denoising | Reduce downstream gating; edit only affected sections |
For crosstalk-specific decisions, continue with how to handle overlapping speech. If you need text rather than audio stems, compare speaker separation and diarization. For a field workflow, see interview speaker separation.
Private processing
Uploads and results are available only to the signed-in account.
30-day result access
Completed source media and outputs are retained privately for 30 days, or until deleted.
Specialist infrastructure
Seply uses third-party professional processing infrastructure and private object storage to complete jobs.
Upload only recordings you own or have permission to process.