Guide

How to Get the Text Out of an Image Without Retyping It

Why free converters return a wall of characters, what to say when you send the picture instead, which parts of the result to check, and what to ask after.

Updated August 16, 2026

Someone is coming to wipe the board. You take a photo on the way out, and what you have saved is a picture of the words rather than the words. Nothing in that file can be searched, pasted into a task list, or read by the two people who were not in the room.

Getting the words back out has two versions that share a job and share nothing else. The old one has a name, optical character recognition, OCR for short, and it is what the free converter sites run. It sweeps the picture for letter shapes, matches each shape to a character, and prints the characters in the order it met them. It reads shapes. It has no idea it is looking at a whiteboard.

The other version starts from a question. You send the picture with a sentence saying what you want out of it, and the reply comes back written to answer that sentence. Same photo, but "give me the left-hand column as a checklist and put a name against every item that has initials next to it" is a request the shape-matching kind cannot even receive.

Why the converter you already tried gave you a wall of characters

Three things break it, and they break it on exactly the pictures people take.

Arrangement. Words on a board or a receipt are placed, not streamed. A shape-matcher works across the image and emits characters in the order it finds them, so two columns come back interleaved a line at a time, item names and prices weld into one string, and the boxed note in the corner lands in the middle of an unrelated sentence. The words are all there. The thing that made them mean something is gone.

Everything that is not a letter. Arrows, ticks, circles, the line joining two ideas, the star next to the one that matters. On a whiteboard those carry as much of the decision as the writing does, and there is no character to match them to, so they are dropped in silence.

Ambiguity. A shape-matcher sees each mark on its own. In handwriting, a 5 and an S are often the same mark, so are a 1 and a lowercase L, and so are an "rn" and an "m". With nothing but the shape to go on, it picks one and moves on. That is where the run-together nonsense comes from: not one bad guess, but a few hundred of them, with no sentence holding any of them together.

What to say when you send it

The reading is a reply to a request, so the request is where most of the quality is decided. "Extract the text" is close to the least useful thing you can ask for, and it is what nearly everybody types. It leaves every decision open: which order the words arrive in, what becomes of the arrows, whether the boxed note is a heading or an aside, whether a column is even a column. Every one of those gets settled. Say nothing and none of them get settled by you.

Describe the shape you want back instead. Four pictures, four different sentences:

Whiteboard from a planning session. Give me each column as its own list, write the arrows as "X leads to Y", and put a name against any item that has initials next to it.

Read the error in this screenshot exactly as written, including the code at the end, then tell me in one line what it means.

Six receipts. One row each: date, vendor, total, currency, and the tax line if it is printed. Say which ones you are unsure about.

Handwritten meeting notes. Type them out keeping the line breaks, and mark anything you cannot read with [?] instead of guessing.

That last instruction is worth stealing for all four. Ten words, and it changes the work you are left with, because the parts that need your eyes come back marked instead of buried.

There is a second thing those sentences do that is easy to miss. Each one says what the picture is. A receipt, a planning board, an error dialog, a page of notes. Naming it is not politeness. It is the context that decides whether a smudged four-digit figure comes back as a year, a total, or a room number.

Send the file the way the device saved it. A screenshot is a PNG and a camera photo is a JPG, and both go over untouched. The HEIC files an iPhone saves by default are fine to send as they are too, Live Photos and burst frames included: they get turned into a JPG on the way in. Picture and sentence travel in the same message, so drop the photo in and put the sentence under it.

Sending more than one

Eight images is the ceiling on one reading, and the eight are read together rather than one after another. Three overlapping photos of a long whiteboard arrive as a single board, so you can ask for one merged set of notes instead of stitching three replies together yourself. Two shots of a page and its back work the same way.

Nine is where it stops, and nothing behind that number quietly splits a bigger pile into batches for you. Twenty receipts is three groups of eight, eight and four. The grouping is worth two seconds of thought, because anything you want weighed against something else should travel with it. Everything downstream, sorting by month, totaling by vendor, spotting the two that are the same receipt photographed twice, happens on the text and no longer cares how the photos were grouped.

From an iPhone, the share button in Photos pushes up to ten pictures into a message at once. That is two more than one reading takes, so a full ten wants splitting across two messages.

What comes back wrong, and where to look

Anything that reads by understanding rather than by matching fills gaps from context. That is why it beats shape-matching on a smeared word in the middle of a sentence, and it is the thing to watch on a number.

Around ordinary words, context works as a correction. Around a serial number, an invoice number or a total, context has nothing to offer, and the same machinery still returns something that looks like a reading. A guess comes back in the same typeface as a fact.

So the check is narrow rather than general. You are not re-reading the whole result against the picture. You read the digits, the proper nouns, and anything you would be embarrassed to get wrong, and you read only those. On a receipt that is four numbers. On a page of meeting notes it is usually the names.

One test settles most of it before you send anything. Zoom into the photo yourself. If you cannot read the word at full magnification, the pixels for that word were never captured, and nothing downstream invents them back. It takes five seconds and it saves you from arguing with a result that was decided at the moment of the photo.

Reading is not rebuilding. What comes back is the words in the arrangement you asked for, not a reconstructed copy of the page, so if what you need is a document that matches the original down to the margins, this is the wrong route to it. If what you need is that page set again in your own tools, ask for the font and the color values in the same breath as the words, and you leave with all three.

Position is the exception, and only when you ask for a piece of the picture rather than a description of it. "Crop out just the table" and "give me the fine print in the bottom right on its own" both work, because the region gets found in the actual image and cut on those coordinates instead of somebody eyeballing a crop box.

If you are the one taking the picture

Since the result is largely decided at the shutter, a board still standing in front of you deserves the next minute more than the request does.

Stand square to it and shoot from the middle rather than from the end of the table. A photo taken at a sharp angle squeezes the far half into a third of the frame, and small writing over there stops being writing at all.

Move until the glare is somewhere else. Where a window or a flash has blown a strip to white, there is nothing underneath it to recover. That is the one failure with no fix later, and two steps sideways solves it now.

Fill the frame with the board and get all four edges inside it. Words cut by the edge come back as half words or not at all.

Shoot a long board as three overlapping frames rather than one wide one.

On the phone app, which button you press matters. The camera button puts an iOS crop step in front of the shot, and that crop is locked to a square, so a wide board loses its ends whatever you do with it. Take the picture in the Camera app and attach it from your photos instead, because that route has no crop step in it at all. Both send the file at full quality rather than as a shrunken copy, which is what keeps small writing legible.

You do not have to ask for the text at all

The request is a question, so it can be the question you actually have. Almost nobody tries this.

The error screenshot is the clearest case. You did not want the string. You wanted to know what it means and what to do next, so ask for that, and take the exact wording as the by-product you paste into a search if the answer does not land.

The receipts are the same. "Which vendor did I spend the most with last quarter" is a question you can ask a pile of photographs without building the table first.

The board too. "Which of these has nobody's initials against it" beats "type this out", and it is the question you were going to ask yourself the moment the typing was done.

The picture and the answer both stay in the conversation, so the next question needs no second upload. Three weeks on, the thread still holds both, and you can ask what the board said about the migration without going near your camera roll.

Getting it out of the chat

Say where the text is going and it arrives in that shape. A Word document for notes somebody else has to read. A spreadsheet or a CSV for the receipts, which is one file for your accountant instead of a folder of photographs. Plain text if the destination is a wiki or a code repo. Or no file at all, because four action items short enough to copy out of the reply do not need one.

That CSV comes out of your own account rather than from a link anyone could open, and it arrives with no logo on it.

If a strategy board is not something you want inside a training set, the opt-out is an email to support and the policy spells out what it involves. The default, for anything sent from 2 August 2026 onward, is that it trains the system. Retention is a separate and simpler question. Your picture sits in the thread, along with everything made from it, for as long as you leave it there. Working copies from a job in flight clear inside seven days.

You need an account for any of it. That is an email address, no card, and it lands with credits on it already, credits being the thing usage gets counted in here, topped up again every day. Enough to find out whether your handwriting is legible to anything other than you.

When it is not really a picture

Two jobs look like this one and have shorter routes.

A PDF is often not a picture problem at all. Drag your cursor across a line before you screenshot anything. If the text highlights, the words are already inside that file and can be lifted out exactly, with nothing to read and nothing to check, so send the PDF itself instead of a screenshot of it. If the cursor grabs nothing, you have a scan, which is pictures of pages, and everything above applies.

Words that were on screen in a video do not need a screenshot either. Name the moment and the frame gets pulled and read: "read the slide at 14:20". Finding the moment is its own job, and the reading is the easy half of it.

The board gets wiped either way. What decides whether that costs you anything is the thirty seconds between taking the photo and sending it: say what the picture is, say what shape you want back, and say that anything unreadable should be marked rather than invented. Then check the numbers, and only the numbers.

Related tools

Your next video is one conversation away.

Free account with credits included. No credit card, no learning curve.