Skip to main content
Konthora

Knowledge Center

How to Transcribe a Video

Extracting speech from video files allows you to create readable notes and prepare standalone subtitle tracks. Learn how to process your video files, apply timestamps, and choose an export format.

Why Transcribe a Video?

Creators and editors use speech-to-text to convert the dialogue inside their video recordings into written documents. Once you convert the audio to text, you can review the content without constantly scrubbing through a video player timeline.

Additionally, transcribing a video is the first required step if you plan to create captions or subtitles to improve the accessibility of your content.

Note that video transcription converts existing spoken audio into text. If you instead need to generate voice narration from a written script, see our guide on text-to-speech for YouTube videos.


Preparing Your Video File

Konthora accepts MP4, WebM, and MOV video files (as well as MP3, WAV, M4A, and AAC audio files). When you transcribe audio from a video, Konthora processes the available audio track embedded inside the file.

Before starting, ensure your file fits within the platform limits. The maximum upload size is 100 MB, and the maximum media duration is 10 minutes. Longer videos must be manually divided into shorter files before upload.

For the best audio transcription accuracy, ensure the video's audio track is clear and free from heavy background noise.


Choosing a Timestamp Mode

Applying timestamps makes it easy to match the written text back to the video timeline. Konthora provides sentence, paragraph, and word-level timestamp modes.

Sentence mode is highly recommended if you are creating subtitle tracks, as it breaks the text into manageable chunks. Paragraph mode is better if you just want to read the transcript like a document. Word-level timestamps provide exact frame-level alignment for precise editing workflows.


Choosing an Export Format

Konthora generates a standalone transcript or subtitle file, rather than burning the text directly into the video. You must choose an export format that fits your goal.

If you need a simple readable document, export as TXT. If you want to add captions to video using a player or editing software, you should export as SRT or VTT. Developers and technical users can also export JSON for programmatic access to the data.


Transcribing a Video with Konthora

You can perform this process securely in your browser. Uploaded media and generated transcript data follow a temporary 60-minute lifecycle, so you must download the final output during your active session. No account is required.

1

Prepare a supported video file within the current limits

Ensure your MP4, WebM, or MOV file is under 100 MB and 10 minutes in length. Divide longer videos into shorter files before uploading.

2

Upload the MP4, WebM, or MOV file

Select your video and upload it to the tool. Konthora will process the available audio track.

3

Choose sentence, paragraph, or word-level timestamps

Select a timestamp mode to help you navigate the resulting text or align subtitles.

4

Start transcription

Click the button to process your media and wait for the extraction to finish.

5

Export as TXT, SRT, VTT, or JSON

Download your standalone transcript or subtitle file to your device before your active session expires.

Frequently Asked Questions

Does Konthora edit my video to add captions?
No. Konthora generates a standalone transcript or subtitle file. It does not edit the source video or permanently burn captions into it.
Can I transcribe a 20-minute video?
No. Konthora accepts media up to 10 minutes in duration. Longer videos must be manually divided into shorter files before upload.
Are my uploaded videos stored permanently?
No. Uploaded media and generated transcript data follow a temporary 60-minute lifecycle and are automatically removed.