Skip to content

speech to-text · commercial

Interview Transcription with Speaker Labels

Turn an interview into editable, speaker-labeled text with timestamps, then review names and wording before exporting TXT, SRT, VTT, or JSON.

Written by
Seply Editorial Team
Reviewed

Interview transcription with speaker labels converts an authorized audio or video recording into editable dialogue grouped by detected voice. Review the labels, names, wording, and timestamps before using the result for quotes, research, subtitles, or publication.

Seply's Speech to Text workspace supports language detection, speaker diarization, word-level timing, speaker filters, editing, and downloads in TXT, SRT, VTT, and JSON.

Choose transcription when text is the deliverable

NeedBest output
Find a quote or topicSearchable transcript
Read interviewer and guest separatelyDiarized speaker labels
Prepare video captionsSRT or VTT
Move data into another workflowJSON
Control each voice in an audio mixSpeaker Separation WAV tracks

Diarization identifies conversational turns in text; it does not isolate the voices into audio stems. If an editor needs separate tracks, use interview speaker separation.

A review-first transcription workflow

1. Upload the closest original

Use the source with the clearest speech and least lossy re-encoding. Seply accepts common audio and video formats up to 1 GB or 10 hours. Confirm that you own the interview or have permission to process it.

2. Select language and diarization

Auto-detect is useful when the language is uncertain. Choose a known language when you want to constrain recognition. Leave diarization on when the speaker boundary matters; turn it off for a single-speaker monologue where labels add no value.

3. Confirm cost before processing

Speech to Text bills 1 credit per started 3 minutes, with a 1-credit minimum. The 30-credit trial covers up to about 90 minutes. Seply displays the calculated requirement before the job starts, and the server-calculated duration is authoritative.

4. Review speaker labels before renaming

Detected labels are working identities such as Speaker 1 and Speaker 2. Listen to multiple known passages before assigning real names. Check short interjections, the opening introduction, and sections with overlap because those are common places for a label to need correction.

5. Correct consequential wording

Prioritize proper names, numbers, technical terms, direct quotes, and any sentence used to support a decision. Search and speaker filtering make this faster, but the source media remains the reference.

6. Export for the next job

  • Use TXT for a clean reading copy.
  • Use SRT for broad subtitle compatibility.
  • Use VTT for web caption workflows.
  • Use JSON when another system needs segments, speakers, or timing structure.

The editable result and source media remain private to the account for 30 days, or until deleted earlier. Download required deliverables before expiration.

Formats and limits

Accepted input
MP3, WAV, M4A, FLAC, OGG, OPUS, WEBM, AAC, MP4, MOV, MKV, AVI, WMA, SPX
Output
Editable transcript, TXT, SRT, VTT, and JSON
Upload limit
Up to 1 GB or 10 hours
Billing increment
1 credit per 3 minutes, 1-credit minimum

Where human review is mandatory

For dense crosstalk, the process in how to handle overlapping speech explains what the source may and may not support. To choose between audio stems and labeled text, see speaker separation vs. diarization.

A defensible handoff checklist

Before sharing the transcript:

  • Confirm that every named speaker was identified from known context.
  • Check dates, figures, names, and direct quotations against the audio.
  • Preserve uncertain words as uncertain rather than guessing.
  • Match the export format to the downstream tool.
  • Remove access when collaborators no longer need the file.
  • Keep the original recording according to your organization's consent and retention policy.

Private processing

Uploads and results are available only to the signed-in account.

30-day result access

Completed source media and outputs are retained privately for 30 days, or until deleted.

Specialist infrastructure

Seply uses third-party professional processing infrastructure and private object storage to complete jobs.

Upload only recordings you own or have permission to process.