Skip to main content
Konthora

Knowledge Center

How to Transcribe a Podcast

Providing written text alongside your podcast audio can help audiences who prefer reading or need accessibility options. Learn how to prepare your podcast files, choose timestamp modes, and export your transcript.

Why Transcribe a Podcast?

Creators use speech-to-text tools to generate readable transcripts of their recorded episodes. A written transcript makes it easier to repurpose podcast discussions into articles, pull out notable quotes, or verify what was said.

Transcribing your episodes also helps provide accessible content for users who are deaf, hard of hearing, or those who simply prefer to read audio to text rather than listen to media.

Alternatively, if you are looking to generate spoken narration from a written script rather than converting existing audio into text, see our text-to-speech for podcasts guide.


Preparing Your Podcast File

Before you can transcribe audio using Konthora, you must prepare your media. Konthora currently accepts media up to 100 MB and up to 10 minutes in duration.

Longer podcast episodes must be divided into shorter files before upload. We support audio inputs (MP3, WAV, M4A, AAC) as well as video inputs (MP4, WebM, MOV) if you record video podcasts.

Keep in mind that overall audio transcription accuracy relies heavily on the quality of the recording. Clear speech with minimal background noise yields better results.


Choosing a Timestamp Mode

Konthora allows you to choose from three timestamp modes: sentence, paragraph, and word-level.

Selecting paragraph mode groups large blocks of spoken text together, which is often useful for reading long discussions. Sentence mode provides a closer timing reference for finding specific quotes, and word-level timing offers exact alignment for creating captions.


Choosing an Export Format

After generating the transcript, you can select from four different export formats.

A plain TXT file is best for reading the transcript as a standard document. If you plan to add the text as closed captions to a video podcast, use SRT or VTT. The JSON format is available if you need to extract the raw transcript and timing data for structured data workflows.


Transcribing a Podcast with Konthora

Konthora provides a browser-based workflow with no account required. Remember that uploaded media and generated transcript data follow a temporary 60-minute lifecycle, so you must export your results before the session expires.

1

Prepare a supported podcast audio file within the current limits

Ensure your audio or video file is under 100 MB and under 10 minutes in duration. Split longer episodes into shorter files before proceeding.

2

Upload the file

Select your prepared media file and upload it to the tool. No account is required.

3

Choose sentence, paragraph, or word-level timestamps

Select how you want your text broken down to match your editing or publishing needs.

4

Start transcription

Click the button to process your media. Wait for the browser to finish extracting the text.

5

Export the result as TXT, SRT, VTT, or JSON

Download the final transcript in your preferred format before your 60-minute session expires.

Frequently Asked Questions

Can I transcribe an entire 60-minute podcast episode at once?
No. Konthora currently accepts media up to 10 minutes in duration. Longer podcast episodes must be divided into shorter files before upload.
What file formats do you support?
We support both audio and video transcription. Supported audio inputs include MP3, WAV, M4A, and AAC. Supported video inputs include MP4, WebM, and MOV.
Is my podcast audio stored permanently?
No. Uploaded media and generated transcript data follow a temporary 60-minute lifecycle and are not permanently stored.