Skip to main content
Konthora

Reference

Transcript Timestamps

A timestamp is a start and end time attached to a piece of text, so you can jump from a sentence to the exact moment it was spoken. Choose sentence level for reading, paragraph level for notes, and word level for subtitles.

Short answers

What are timestamps in a transcript?
A start and end time for each piece of text, which lets you jump straight from a line to the moment it was said rather than scrubbing through audio to find it.
What is the difference between sentence, paragraph and word-level timestamps?
Sentence level times each sentence, paragraph level groups related speech into blocks, and word level times every individual word. The choice depends on whether you are reading, summarising, or making subtitles.
Why do word-level timestamps matter for subtitles?
Captions need to appear when the words are spoken. Sentence-level timing holds a whole line on screen for its entire duration, so the viewer reads it while the speaker has already moved on, and captions end up trailing the voice.

What can I do with timestamps in a transcript?

A transcript without timing tells you what was said. A transcript with timing tells you what was said and when, and the second is far more useful for anything that involves moving between the text and the recording.

That means finding a quote. If you need one specific line from a forty-minute interview, sentence timestamps let you skim the text, note the time, and jump straight there. Without them you listen to forty minutes to check one sentence.

It also means subtitling and chaptering, which are timing problems rather than text problems. And it means verifying a transcript against its source, which is the difference between a transcript you can publish and one you have to trust.


Which timestamp mode should I choose?

Sentence level suits reading and quote-finding. One timestamp per sentence is enough granularity to locate anything, and the text stays in natural paragraphs rather than breaking on every word.

Paragraph level suits summarising and note-taking. It groups related speech into blocks that read like notes rather than dialogue, which is usually what you want when a meeting or lecture is being turned into study material.

Word level suits subtitles and anything that needs tight alignment. It is the only mode where a caption can track the voice instead of trailing it, at the cost of more timestamps and a busier transcript.


Why do subtitle timings differ from transcript timings?

Subtitles have a reading-speed constraint that transcripts do not. A viewer reads roughly 160 to 180 words per minute, so a caption has to arrive when the words are spoken rather than when the sentence finished.

Sentence-level cues break that. The whole sentence stays on screen for its entire duration, which means the viewer is reading a line while the speaker has already moved past it. The captions look correct and feel constantly behind.

Word-level timestamps fix it, because each cue reflects what is actually being said at that moment. It is the single biggest quality difference between an automatically generated caption file and one that is genuinely watchable. The subtitle format guide covers the file side of this.


How do I add timestamps to a transcript I already have?

If you already have text without timing, re-running it through a transcription tool is usually faster than inserting timestamps by hand, because the timings come from the audio rather than from a guess.

Hand-inserting is only worth it for a very short excerpt, where the overhead of re-transcribing exceeds the work of typing a handful of timecodes.

If your source was MP3, note that the encoding affects both the text and how cleanly the timings land. The MP3 to text guide covers which settings matter and why.


How do I get timestamps out in a usable format?

Plain text with a leading timestamp per block is the readable option, useful for show notes and written summaries.

SRT and VTT carry machine-readable timing and are what caption players expect. Use SRT for most platforms and VTT for HTML5 players; both are covered in the format reference.

JSON is the option when you are scripting, because it carries per-word timing in a structure you can actually use rather than a flat string you would have to parse.


Frequently Asked Questions

Can I add timestamps to a transcript I already have?
Usually by re-running it through a transcription tool, which is faster than typing timecodes and gives you timings taken from the audio rather than estimated. Hand-inserting only makes sense for a very short excerpt.
Which timestamp mode should I choose?
Sentence level for reading and finding quotes, paragraph level for summarising a lecture or meeting, and word level for subtitles. Word level is the only mode where captions track the speaker instead of trailing them.
Does choosing word-level timestamps slow transcription down?
Slightly, because the model produces a timestamp for every word rather than every sentence. For files at the ten-minute limit the difference is minor, and the subtitle quality gain is usually worth it.
Which export format should I use for timestamps?
TXT for readable notes, SRT for most subtitle platforms, VTT for HTML5 players, and JSON when you are scripting and want per-word timing in a usable structure.
Why do my subtitles fall behind the speech?
Almost always because sentence-level timing was used. A whole sentence stays on screen for its duration, so the viewer reads it while the speaker has moved on. Switch to word-level timestamps and the captions track the voice.