Skip to main content
Konthora

Knowledge Center

Audio, Video, and Transcript Formats Supported by Konthora

This guide details the media input, generated-audio output, and transcript export formats currently supported by the platform, including limits and practical workflows.

Audio Outputs for Text-to-Speech

When you convert written text into spoken audio, the resulting file can be downloaded in two standard formats. Each generation is currently subject to a verified limit of 2,000 characters per request to ensure optimal performance.

MP3

A compressed audio format that balances sound quality with smaller file sizes. Ideal for sharing online, embedding in web pages, or keeping storage requirements low.

WAV

An uncompressed audio format that preserves maximum acoustic fidelity. Ideal for professional editing workflows, video production, or podcast mastering.


Supported Audio and Video Inputs for Transcription

If you are exploring speech to text solutions, you can upload pre-recorded media files directly for processing. The system automatically extracts the audio track from video files, meaning you do not need to convert video to audio beforehand.

Supported Audio

  • MP3
  • WAV
  • M4A
  • AAC

Supported Video

  • MP4
  • WebM
  • MOV

Processing Limits

Uploaded media files must not exceed 100 MB in size, and the maximum allowed media duration is 10 minutes.


Transcript Export Formats

After an audio or video file is processed, you can download the transcription in one of four formats. Each format serves a specific practical purpose depending on your workflow.

TXT (Plain Text)

A simple, readable document containing only the transcribed words without any timestamps or formatting tags.

SRT (SubRip Subtitle)

The most widely accepted format for timed subtitles. Ideal for video editors (Premiere, Resolve) and social platforms.

VTT (Web Video Text Tracks)

A modern caption format designed for HTML5 web video players, supporting precise alignment and styling.

JSON

A structured data format that developers use to ingest the transcript and raw timestamp data programmatically.


How to Choose the Right Format

Selecting the correct file type depends entirely on what you intend to do with the result:

  • Choose MP3 for compact audio sharing or web embedding.
  • Choose WAV for higher-quality audio editing workflows.
  • Choose TXT for a readable plain transcript to review or publish as an article.
  • Choose SRT for broadly supported timed subtitles in video editors or YouTube.
  • Choose VTT for detailed captions workflows in HTML5 web players.
  • Choose JSON for structured transcript and timestamp data in custom applications.

Timestamp Modes and Export Behaviour

When exporting to a timed format like SRT or VTT, the pacing of the subtitles is determined by how the text is grouped. You can adjust the visual density of the export by selecting different grouping modes:

  • Sentence timestamps: Groups text logically by sentence structure, making captions easy to read.
  • Paragraph timestamps: Creates larger blocks of text on screen, useful for long monologues or transcripts.
  • Word-level timestamps: Aligns every individual word to its exact timecode, primarily used for dynamic social media captions (e.g., TikTok/Reels text effects).

For a deeper explanation of how these modes affect readability, read the guide on transcription timestamps.


How Uploaded Files and Generated Audio Are Handled

Konthora operates without a user-account system, which directly informs how files are processed and stored on the server.

For transcription jobs, uploaded media files and their corresponding generated transcripts are held in temporary storage to allow you to preview and download the results. These files are subject to an automatic 60-minute deletion lifecycle, after which they are permanently removed from the server.

For text-to-speech jobs, the input text is processed entirely in memory, and the generated TTS audiomust be downloaded during your active browser session.

Because there is no long-term file retention, it is important to download your preferred formats immediately after processing. For more details on data handling, refer to the Privacy Policy.

Frequently Asked Questions

Can I upload video files for transcription?
Yes, you can upload MP4, WebM, and MOV files. The system will automatically extract the audio track from the video and process the transcription.
What is the maximum file size for uploads?
The transcription tool accepts files up to 100 MB in size and up to 10 minutes in duration.
Which format should I use for YouTube subtitles?
Both SRT and VTT are widely supported for subtitles and web captions. SRT is the most universally accepted format for video editors and platforms like YouTube.
Are my uploaded files stored permanently?
No. All uploaded media files and generated transcripts are automatically deleted from the server after 60 minutes.
Is there a limit on text-to-speech generation?
Yes, each text-to-speech generation is currently limited to 2,000 characters per request.