Quick Steps: Transcribe Audio in 4 Steps
- 1Go to Konthora’s free Audio-to-Text tool in your web browser.
- 2Upload your audio or video file (MP3, WAV, M4A, AAC, MP4, WebM, or MOV — up to 100 MB / 10 mins).
- 3Select your preferred timestamp grouping: sentence-level, paragraph-level, or word-level.
- 4Click Transcribe Audio and export your result as TXT, SRT, VTT, or JSON format.
Step-by-Step: How to Transcribe Audio on Konthora
Step 1 — Upload Your Audio or Video File
Open the Konthora Audio to Text workspace. Drag and drop your media file directly onto the upload area, or click to select a file from your computer.
Konthora supports MP3, WAV, M4A, AAC audio files and MP4, WebM, MOV video files. Files must be within 100 MB in size and 10 minutes in duration.
Step 2 — Choose Your Language and Timestamp Mode
Verify that your audio language is set to English. Next, choose how you want your timestamps formatted:
- Sentence-level: Groups text into complete sentences with start and end timestamps. Ideal for reading.
- Paragraph-level: Groups text by natural speech pauses. Ideal for meeting notes.
- Word-level: Provides precise start and end timings for every individual word. Ideal for audio indexing.
Step 3 — Start Transcription
Click Transcribe Audio. Konthora uses the Whisper speech recognition model to convert your audio into accurate text with precise timestamps.
Step 4 — Download or Copy Your Transcript
Once transcription is complete, review your transcript in the browser. You can copy the text directly to your clipboard or download it in your format of choice: TXT, SRT, VTT, or JSON.
Supported Audio and Video Formats
Konthora accepts 7 common media formats:
Audio Formats
MP3, WAV, M4A, AAC
Video Formats
MP4, WebM, MOV (audio track extracted)
For all formats, the maximum allowable file size is 100 MB and the maximum audio duration is 10 minutes.
Understanding Timestamp Options
For a detailed breakdown of timestamp modes and format compatibility, see our guide on transcription timestamps.
Sentence Timestamps
Sentence-level timestamps assign a start and end time to each full sentence. This produces clean, natural paragraphs suitable for reading articles or creating show notes.
Paragraph Timestamps
Paragraph timestamps group speech segments by natural pauses in conversation, ideal for long-form speeches and interview transcripts.
Word-Level Timestamps
Word-level timestamps record the exact start and end millisecond of every spoken word, useful for precise audio editing and video synchronization.
Exporting Your Transcript
TXT Format
Plain text format containing speech segments without timestamp headers.
SRT Format
SubRip Subtitle format with sequential numbers, timestamps, and lines.
VTT Format
WebVTT caption format supported natively by HTML5 video players.
JSON Format
Structured JSON containing segment objects with word-level timing data.
Example SRT Output
1 00:00:00,000 --> 00:00:03,500 Welcome to Konthora's audio transcription workspace. 2 00:00:03,800 --> 00:00:07,200 Upload your audio or video file to get accurate timestamps.
Tips for Better Transcription Results
- Record with a dedicated microphone in a quiet room to reduce background noise.
- Maintain a steady speaking distance from your microphone to keep volume levels consistent.
- Avoid overlapping voices — single-speaker speech transcribes with highest accuracy.
- Use high-bitrate MP3 or WAV files instead of heavily compressed voice memos.