Free tool

Audio to Text

Free transcription for files up to 90 minutes — a whole episode, not a voice note. Speaker labels, text, SRT and VTT, emailed to you.

Looked after by Clack, a typewriter with a sheet of paper in it.

We show the transcript here and email you the text, SRT and VTT, so you still have them tomorrow.

Free. No account. 2 files a day per person, up to 90 minutes each.

Ninety, because across 96 transcripts we measured the median episode was 47.3 minutes and 96% were 90 minutes or under.

How to transcribe a file

  1. 1Drop in your fileAudio or video, up to 90 minutes and 600MB.
  2. 2Say where to send itThe transcript appears here and lands in your inbox as TXT, SRT and VTT.
  3. 3Wait a few minutesYou can close the tab. The email arrives either way.
  4. 4Copy or downloadRead it here, copy it, or take the subtitle files into your editor. Turns are already labelled.

Questions

Is it really free?
Yes. No card, no trial, no watermark on the text. We pay AssemblyAI about 15 cents an hour of audio and we have capped what the tool can spend in a month, which is why there is a limit of 2 files a day per person. If the month’s budget runs out the tool says so and offers to email you when it resets, rather than pretending to be broken.
How long can the file be?
Ninety minutes. We chose that by measuring our own work rather than copying the fifteen-minute cap you see elsewhere: across 96 transcripts we made between February and September 2026 the median was 47.3 minutes and 96% were 90 minutes or under. A fifteen-minute limit would have covered 17% of them.
Does it say who is speaking?
Yes, and that is unusual on a free tier. If more than one person is talking, each turn is labelled Speaker A, Speaker B and so on, in the text file and in the subtitles. One person talking gets no labels at all, because a monologue with "Speaker A:" on every paragraph is just noise. It works from the sound of the voices, so it cannot know anyone’s name — find and replace takes about ten seconds.
What files does it take?
MP3, M4A, WAV, FLAC, OGG, AAC, and video files — MP4, MOV, WebM, MKV. Up to 600MB. Video is fine: we only read the audio track, and you get text back, not a video.
What do I get back?
The transcript on screen to read and copy, and three files: a plain text version in paragraphs, an SRT and a VTT for subtitles. If more than one person is speaking, every one of them is labelled — in the text and in the subtitles. All three files are emailed to you as well, so you still have them after you close the tab.
Who can see my audio, and how long do you keep it?
Your file is uploaded to private storage that is not readable from the web, sent to AssemblyAI to be transcribed, and deleted after 7 days. The text is kept for 30 days so the links in your email keep working, then that goes too. We do not use your audio to train anything.
Does it label who is speaking?
No. The free tool gives you the words. Speaker labels, chapters, clips and captions are what Slice’s paid tools do, and that is the honest difference between the two.
What languages does it handle?
It is set up for English and detects the language of what you send. If what comes back is in another language the accuracy is not something we will promise, and the page tells you what was detected rather than pretending.
Do I have to sign up for emails?
No. We ask for your email address because we send the transcript to it, and nothing else is asked before you upload. Once the transcript is done there is a button on the results screen offering one email a month about new free tools — pressing it is the only way you are ever added, and every email has a working unsubscribe link.
If you are a script, or an agent

There is an HTTP API for exactly this: POST a media URL, get text, SRT and VTT back as JSON. Same limits, same engine, same transcripts — no upload, no email step, and a free key in one request. Without a key you still get one short file a day, so a first try costs nothing at all.

Why ninety minutes, and not fifteen

Most free transcribers stop somewhere between ten and thirty minutes. That number is not chosen around what people actually record; it is chosen around what is cheap to give away. So we measured our own work instead.

Across 96 transcripts made between February and September 2026 — our own customers’ podcast episodes, interviews and recorded calls, measured on 16 September 2026:

  • Median length: 47.3 minutes. Half of everything we transcribe is longer than that.
  • 96% were 90 minutes or under. The longest was 93.6 minutes.
  • Only 17% were 15 minutes or under. A fifteen-minute free tier would have been useless for five files in six.
  • Average 7,775 words, at about 172 words a minute. 94% were audio rather than video.

Ninety minutes is where those two facts meet: it covers the overwhelming majority of real episodes, and it is a length we can afford to give away. These are aggregate figures from our own account — no customer file, name or content is in them.

How this compares

We read the free tiers of six other transcription tools off their own pricing pages and put them in one table — longest free file, whether they label speakers, whether you need an account, and what happens when the free part runs out. It includes where we lose, because a comparison that does not is an advertisement.

Who is speaking

Most transcription tools put speaker labels behind the paywall, and we did too until September 2026. It was the wrong line to draw. Most of what people send this tool is an interview or a two-hander, and a conversation transcribed without turns is not a cheaper document — it is a worse one, because the reader ends up doing the work in their head.

So the free tool diarises. Each turn is labelled in the text and in the subtitle files, the label appears where the speaker changes rather than on every line, and a file with one voice in it comes out with no labels at all. What separates free from paid here is the model, the ninety-minute cap and the 2 files a day — not whether the result is readable.

What it costs us to run

Transcription is about 15 cents an hour of audio, which we pay. That is why there are limits, and why we would rather print them than let you find them: 2 files a day per person, ninety minutes each, and a monthly budget for the tool as a whole. If a month runs out, the tool says so and offers to email you when it resets. It does not quietly get worse or start asking for a card.

What happens to your file

It is uploaded to private storage that is not readable from the web, passed to AssemblyAI to be transcribed, and deleted after 7 days. The text is kept for 30 days so the download links in your email keep working, then that is deleted as well. Nothing is used to train anything.

Where the free tool stops

It gives you the words and who said them. It does not cut clips, burn captions onto video, or find the good bits — those are what Slice’s paid tools do, and pretending otherwise would waste your time. It is also tuned for English: it will transcribe other languages and tell you which one it detected, but we are not going to promise the accuracy.

This tool is looked after by Clack.Listens hard, types fast, never asks you to repeat yourself.

More Free Tools

Shorts by Slice

From the team behind these free tools

One video. A week of posts.

Upload an episode and get back the moments worth posting, cut and captioned.

  • Reads the whole episode, then picks the excerpts worth posting
  • Captions burned in and framed vertical, ready to post
  • Seven days free, no card
Try Shorts by Slice