MalcolmCC™ › WebVTT caption generator
WebVTT caption generator
Turn any video into accurate, perfectly timed WebVTT (.vtt) captions
in minutes — with a real timing editor and export you can drop straight into an HTML5
<track>, a video platform, or an LMS.
What is a WebVTT file?
WebVTT (Web Video Text Tracks, the .vtt format) is the web-native caption
format — the one HTML5 video reads through the <track> element, and
the one most modern players, LMSes, and video hosts accept. Each cue is a start and end
timestamp with a line of text, so captions appear exactly when the words are spoken.
The hard part isn't the file format; it's getting the timing right. That's what
MalcolmCC is built for.
Two ways to get perfectly timed .vtt
1. AI transcription
Upload your video and MalcolmCC transcribes the speech and times every cue automatically. You get a complete WebVTT draft in minutes, then fine-tune it in the editor. The spoken language is worked out from the audio, so there's nothing to set before you start.
2. Bring your own script
Have an approved script? MalcolmCC uses your exact words as the caption text and times them to the audio — no transcription guesswork, no paraphrase. This is unique to MalcolmCC, and it's the difference between "close enough" captions and captions that match your approved copy word-for-word. See how script-based captioning works →
Which languages? English, Dutch, French, German, Italian, Polish, Portuguese and Spanish — on both of the paths above, transcription and script alignment alike. MalcolmCC works out which one is spoken and captions in it; you can set it by hand on any file if it gets that wrong. It captions speech in the language it was spoken — it does not translate. Audio in anything outside those eight is refused with a clear message rather than turned into bad captions.
A real timing editor, not just an export button
Reviewing and fixing captions is fast: jump from cue to cue, retime with a drag, and edit the text inline — the editor is built to get you to a finished, accurate track quickly, not to bury you in busywork.
- Nudge any cue on a visual timeline until it lands exactly.
- Silence detection — the timeline shows where the audio goes quiet, so cues land against real pauses to aid in perfect cue placement.
- Remote media — caption a file that already lives on a public URL.
- Eight languages — detected from the audio, with a manual override on any file.
- Real .vtt and .srt export — download a standards-compliant WebVTT or SubRip file, ready to attach.
- Import what you already have — open an existing .vtt or .srt in the editor and fix its timing.
SRT vs VTT? SRT is the older, plainer subtitle format; WebVTT is its web-native successor, with cue styling and positioning and native HTML5 support. If you're captioning for the web, an LMS, or a modern player, WebVTT is the format you want. If you're uploading to YouTube or handing a file to a video editor, SRT is usually what's asked for. MalcolmCC exports both, from the same cues — and imports both, so you can bring captions you already have.
Priced for everyone
Start free — your first 15 minutes of media, every feature, no card. After that it's $6 per hour of media, pay-as-you-go, with $0 owed in months you don't caption. That's the lowest rate in the field. See full pricing →