Text & Writing

Audio to Text (AI Transcription)

Transcribe voice memos, interviews and meetings with AI that runs entirely in your browser — free, no upload, with subtitle export.

  • Free forever
  • No sign-up
  • Runs in your browser
Share X LinkedIn

Your recording stays on this device. OpenAI's Whisper model runs entirely in your browser — no upload, no login, no per-minute fee. The first run downloads the AI model once (~60 MB); after that it is cached.

No recording yet — drop a voice memo above, or .

What this tool does

This tool converts speech into text — voice memos, interviews, lectures, meeting recordings — using OpenAI's Whisper, the same model family behind most commercial transcription services. The difference is where it runs: instead of uploading your recording to someone's server, the model itself is downloaded into your browser (about 60 MB, once), and the transcription happens on your own device.

That architecture is not a gimmick for audio. Recordings are among the most sensitive files people handle — a meeting contains everyone's voice, a memo contains half-formed thoughts you never meant to publish. Commercial transcription sites ask you to upload exactly that, then meter you per minute. Here nothing is uploaded, nothing is stored, and there is nothing to pay per minute — which is only possible because your device does the work.

What you get

  • An editable transcript in a plain text box — fix a name, delete an "um", copy the whole thing with one click.
  • Subtitle files — the transcript keeps timestamps, so you can download ready-made .srt (YouTube, video editors) or .vtt (HTML5 video) files in addition to plain .txt.
  • Any common language — Whisper detects the spoken language automatically, or you pick it manually. A "Translate to English" option turns a German or Spanish recording directly into English text.
  • Live preview — words appear as the model works through the recording, so you see progress instead of staring at a spinner.

Honest expectations

Whisper's base model is a workhorse, not a court stenographer. Clear speech into a decent microphone transcribes remarkably well; mumbled crosstalk in a echoing room does not. Punctuation is good, speaker separation does not exist (everything lands in one continuous text), and unusual names get spelled phonetically. For most voice memos and meetings the result needs a quick read-through, not a rewrite.

Speed depends on your hardware. With WebGPU — every current browser has it — transcription runs several times faster than the recording's length. On browsers without WebGPU the tool falls back to a slower CPU mode and limits recordings to 10 minutes so your tab stays responsive.

How to use it

  1. Drop an audio file onto the box above (MP3, M4A, WAV, OGG — or a video, whose soundtrack gets extracted). No file handy? Load the built-in sample.
  2. Optionally pick the spoken language, or leave auto-detect on. Tick Translate to English if you want an English transcript of a foreign-language recording.
  3. Hit Transcribe. The first run downloads the AI model once; after that it is cached and the tool starts instantly.
  4. Edit the transcript in place, copy it, or download it as .txt, .srt or .vtt.

Pairs well with

Count the words of an interview with the Word Counter, tidy casing with the Case Converter, or go the other direction and pull text out of a screenshot with Image to Text (OCR) — like this tool, it runs entirely on your device, so nothing in the chain ever gets uploaded.

Frequently asked questions

Comet's got your back

Stuck on something? Every tool has a short guide and FAQ — and Comet can point you to the right spot.

Visit help centre