Guide

How to Turn a Recording into Text You Can Actually Use

What automatic transcription gets right and wrong, how to ask so the transcript is usable rather than merely accurate, and what to do with the text after.

Updated August 6, 2026

Fifty-two minutes, recorded on a Tuesday, untouched since. You know roughly what is in there. You also know that getting at the one part you need means playing the other fifty-one minutes again with a finger on the spacebar, so the file has stayed exactly where it is.

Automatic transcription is software writing down every word that was said, with nobody typing. Whether it is good enough is the real question, and that question has no answer until you say what the text is for. Something you can search is a far lower bar than something you can print, and most people arrive assuming the higher one.

Two honest pictures of the same tool. Give it a clear recording of one or two people and what comes back is readable, searchable, and wrong in a handful of places, most of them names. Give it a phone lying in the middle of a boardroom table while four people talk across each other and you get a rough index of a conversation rather than a record of one. Both are worth having. They are not the same thing.

Good enough for what?

Three jobs, three different bars, and the distance between them is where the worry comes from.

Something to search. You want to know whether the pricing objection came up, and roughly when. A wrong word costs you nothing here, because you are reading the sentences around the hit and jumping back to the audio anyway. Almost any recording with audible speech on it clears this bar, the boardroom phone included.

Something to quote. A sentence is going into an article, a deck, or a customer story, and it has to be what the person actually said. The text gets you to the fifteen seconds where the quote lives. You then play those fifteen seconds and read along. Checking one quote is a fifteen-second job. Checking a whole hour, in case you end up quoting from anywhere in it, is an hour.

Something to publish as the text itself. Court records, compliance archives, a transcript page on your site, captions a broadcaster signs off. Every line gets read by a person before it ships. Automatic transcription turns that into reading rather than typing, which is most of the work and not the last of it. A human transcriber, priced by the audio minute and back in a day or two, still earns the fee when the party who has to accept the document is a court or a regulator.

Accuracy gets sold as a percentage, and the percentage flatters everybody who quotes it. Ninety-seven sounds like a comfortable pass, the sort of mark you stop worrying about. Work it out and it is one wrong word in every thirty-three, which is two or three in a paragraph the length of this one, and they will not be the two or three you would have picked. The words that survive are the ordinary ones. The words at risk are the ones carrying the meaning.

Where it actually goes wrong

The errors are not scattered at random, which is the useful part. They gather in five places.

Words it has no reason to know. Your company, your product, the surname of the person you interviewed, the vocabulary of your field. Whatever is doing the listening has heard an enormous amount of ordinary speech and none of your roadmap, so it puts down the nearest ordinary word and does it with total confidence. On a clean recording this is where most of the damage is, and it is also the cheapest thing to fix, because you already know the list.

Numbers going past at speed. Fifteen and fifty are nearly the same sound in the middle of a sentence, and so is any string of digits somebody rattles off from memory. Every number you plan to repeat in public is worth hearing again.

Two people at once. A transcript is a single line of words on a single clock, so when two voices overlap, what comes back is not two lines. One wins, and the other is dropped or blended into the winner. Polite interruption survives that. Overlapping speech, which broadcasters call crosstalk, does not, and it tends to be exactly where the meeting got interesting. The tell is a sentence that changes direction halfway through, because that is two half-sentences welded into one.

The room. A phone in the middle of a table mostly hears the table. If there is a steady hum under everything, or the voices sound far away, cleaning the audio first and transcribing the clean version is not fussiness. It is the difference between a page you skim and a page you rewrite.

Silence, and music. Nobody warns you about this one. Over a stretch where nobody is speaking, the text can be invented outright: hold music, applause, the four minutes of dead air before a call properly starts. The tells are specific. A phrase repeating three or four times in a row. Confident lines sitting in a stretch you know was quiet. A single word whose timing spans several seconds, when real words take a fraction of one. Any of those means going back to the audio at that point instead of reading on.

An unfamiliar accent does not break a transcript. It raises the rate of exactly the failures above, so the thing to watch for is the combination. A strong accent plus specialist vocabulary plus a bad room is a hard recording. Any one of the three on its own usually is not.

Does it help to tell it what the recording is about?

Not the part that listens, and that is the thing worth understanding. The listening pass gets the audio file and nothing else. Not the names, not the subject, not the language, not the careful sentence you wrote explaining what the meeting was about. Your instructions travel with the request, but they are not in the room when the words are being written down.

So context does not make it hear better. What context does is remove the second round trip. The words come back, and everything after that is work on text: correcting the spellings, cutting the filler, shaping the result into notes or a document. If the names were in your first message, the fixing has already happened by the time you read the transcript. If they were not, you find the errors yourself and then ask.

You do not set the language either. It is detected from the audio, which is fine for a recording in one language and worth a look at any recording that switches halfway.

What to send, and what to say with it

Send the file as it is. Any audio format that plays on your machine works, and video works too: hand over the MP4 and the sound is pulled out for you, so there is no exporting step first. If the recording is on your phone, the share sheet passes it straight to the Chat Octopus app. If part of what you need was on screen rather than in the words, the picture can be read as well without a second upload.

Worth knowing before a client call or an unreleased episode goes anywhere. Uploads and results sit in the conversation for as long as you leave them there, and the scratch copies written while a job is running get swept within seven days. Anything submitted on or after 2 August 2026 also feeds into training the system. Opting out begins with an email to support, who then take you through what it involves, and the policy is where the exact terms live.

Then say more than "transcribe this". A message that earns its length:

Transcribe this customer interview, about forty minutes. Spellings you will need: Novatek, Instrumatic, and the product is called Halo. I want it as a document with times against it, plus a separate list of everything that was said about pricing.

Every clause there is doing a job. The spellings arrive as corrections rather than as mistakes you go hunting for later. The length is a sanity check, because a transcript that stops at eleven minutes has told you something went wrong. The destination picks the file, so a caption file does not turn up where you wanted a Word document. And the last sentence is the thing you were actually after, which is rarely the transcript.

Can it tell you who is talking?

Not from the sound of the voices, and this is the one expectation worth resetting before you start. The words come back with times against them and no names on them. Nothing in the pass that listens measures who is speaking; it hears speech and it writes speech down. Separating a recording by voice is a different job with its own name, diarization, and it is not part of what produces your transcript.

Nothing puts a name next to a line unless you ask for one, and a name you ask for is a guess read out of the words rather than heard in the voices, so you are the one who has to agree with it. On a two-person interview that is easy work, because one voice asks the questions and the other answers them, and the timings let you check a handover in seconds instead of hunting for it. On four people around one microphone it is real work, and no transcript is going to spare you from it.

The fix is at recording time rather than afterwards. Some recorders and most call apps can save each person as their own file, which is what multitrack means, and one file per person turns attribution from a guess into a fact: transcribe each on its own, then merge them in time order with the right name on each. One thing to know if your two voices are in a single stereo file instead: the audio is mixed down to one channel before anything listens to it, so the left and right separation you were counting on is gone by then. Ask for the channels to be split first, then transcribe each side.

How long can the recording be?

There is no time ceiling, and the mechanism behind that is worth knowing because you can see it in the output. Each request carries a size limit, and the prepared audio, meaning the converted copy that actually gets sent rather than your original file, crosses it somewhere around the thirteen-minute mark. Past there the recording gets cut into eight-minute pieces, each one transcribed on its own, then stitched back together with every timestamp shifted onto the original clock. A two-hour session comes back as one transcript, and the time sitting next to a line at 01:42:11 is the real position in your file, not a position inside some fragment. Nothing here runs at the speed of listening.

Under that mark none of this applies: a ten-minute recording goes over in one piece with no joins in it. Above it, the cuts do not overlap, so a word landing exactly on a boundary can come back cut in half. A sentence that looks strangely broken at eight minutes, or sixteen, or twenty-four, is the seam rather than the speaker. Language gets settled piece by piece as well. On a recording that starts in one language and turns into another, each piece is read on its own, so the change lands on a seam instead of where it actually happened.

What comes back

Ask for what the next step needs rather than for "a transcript".

The words in the reply, when all you came for is a paragraph or a number. Nothing to download and nothing to open.

A Word document, when the text is heading into a report, an article draft, or a set of notes for people who missed the meeting.

An SRT, when the recording belongs to a video. That is a caption file: the lines, plus the moment each one should appear and disappear. It is what caption uploads ask for nearly everywhere, and editing software reads it too. VTT is the same idea in a different wrapper. If the actual point is captions showing on the video itself, ask for that instead and the burned-in version arrives alongside the file.

Timing goes down to the individual word, not just the line: each one has a start and a finish of its own. That is what makes "cut from where she starts that sentence" a workable instruction, and it is what caption styles that light up each word as it is spoken are built on. It also means the caption file comes in two shapes: one word per line for that word-by-word treatment, or lines grouped up to eight words for anything a person has to read.

The transcript is the index, not the deliverable

Hardly anyone wants the text. They want the four action items, the three quotes with their times, the summary for the people who were not there, the show notes, the thirty seconds worth clipping. Each of those is one more sentence once the words exist, and it can be as narrow as you like: every question that got asked, everything that got promised, the exact stretch where the price came up.

The other half of it is that the recording stops being something you have to listen to. Weeks later, "did we ever agree a date for the migration?" is a question the same conversation can answer, because the words are still sitting in it.

The recording needs one message, and "transcribe this" is the weakest version of it. Attach the file, spell the names it has no way of knowing, say roughly how long it runs, and say what you want at the end of it. Send it that way and the first version back is close to the version you keep. Then play the two minutes you actually care about and read along while they run. That check is the whole difference between a transcript you quote from and one you hope about, and the other fifty minutes stay unplayed.

Related tools

Your next video is one conversation away.

Free account with credits included. No credit card, no learning curve.