Interview transcription with speaker labels converts an authorized audio or video recording into editable dialogue grouped by detected voice. Review the labels, names, wording, and timestamps before using the result for quotes, research, subtitles, or publication.
Seply's Speech to Text workspace supports language detection, speaker diarization, word-level timing, speaker filters, editing, and downloads in TXT, SRT, VTT, and JSON.
Choose transcription when text is the deliverable
| Need | Best output |
|---|---|
| Find a quote or topic | Searchable transcript |
| Read interviewer and guest separately | Diarized speaker labels |
| Prepare video captions | SRT or VTT |
| Move data into another workflow | JSON |
| Control each voice in an audio mix | Speaker Separation WAV tracks |
Diarization identifies conversational turns in text; it does not isolate the voices into audio stems. If an editor needs separate tracks, use interview speaker separation.
A review-first transcription workflow
1. Upload the closest original
Use the source with the clearest speech and least lossy re-encoding. Seply accepts common audio and video formats up to 1 GB or 10 hours. Confirm that you own the interview or have permission to process it.
2. Select language and diarization
Auto-detect is useful when the language is uncertain. Choose a known language when you want to constrain recognition. Leave diarization on when the speaker boundary matters; turn it off for a single-speaker monologue where labels add no value.
3. Confirm cost before processing
Speech to Text bills 1 credit per started 3 minutes, with a 1-credit minimum. The 30-credit trial covers up to about 90 minutes. Seply displays the calculated requirement before the job starts, and the server-calculated duration is authoritative.
4. Review speaker labels before renaming
Detected labels are working identities such as Speaker 1 and Speaker 2. Listen to multiple known passages before assigning real names. Check short interjections, the opening introduction, and sections with overlap because those are common places for a label to need correction.
5. Correct consequential wording
Prioritize proper names, numbers, technical terms, direct quotes, and any sentence used to support a decision. Search and speaker filtering make this faster, but the source media remains the reference.
6. Export for the next job
- Use TXT for a clean reading copy.
- Use SRT for broad subtitle compatibility.
- Use VTT for web caption workflows.
- Use JSON when another system needs segments, speakers, or timing structure.
The editable result and source media remain private to the account for 30 days, or until deleted earlier. Download required deliverables before expiration.
Formats and limits
- Accepted input
- MP3, WAV, M4A, FLAC, OGG, OPUS, WEBM, AAC, MP4, MOV, MKV, AVI, WMA, SPX
- Output
- Editable transcript, TXT, SRT, VTT, and JSON
- Upload limit
- Up to 1 GB or 10 hours
- Billing increment
- 1 credit per 3 minutes, 1-credit minimum
Where human review is mandatory
For dense crosstalk, the process in how to handle overlapping speech explains what the source may and may not support. To choose between audio stems and labeled text, see speaker separation vs. diarization.
A defensible handoff checklist
Before sharing the transcript:
- Confirm that every named speaker was identified from known context.
- Check dates, figures, names, and direct quotations against the audio.
- Preserve uncertain words as uncertain rather than guessing.
- Match the export format to the downstream tool.
- Remove access when collaborators no longer need the file.
- Keep the original recording according to your organization's consent and retention policy.
Private processing
Uploads and results are available only to the signed-in account.
30-day result access
Completed source media and outputs are retained privately for 30 days, or until deleted.
Specialist infrastructure
Seply uses third-party professional processing infrastructure and private object storage to complete jobs.
Upload only recordings you own or have permission to process.