Audio AI

Text to Speech That Comes With Captions

You paste the script in, and what comes back isn't just an audio file. The captions that match it, word-timed and segment-timed, land in the same reply, already synced to a recording that didn't exist a minute ago.

Try asking

How long a read can be

A script up to about five thousand characters comes back as one read, roughly five minutes of audio. Send something longer and it goes out in parts, each one read separately and then joined into a single file. Keep the direction line, the instruction you put before the script telling it how to read, identical across every part, or the joins land as one narrator changing character mid-sentence.

Picking a voice

There's no dropdown of voices to click through and no sample to preview first. Say what you want instead, warm and gentle, firm and mature, and the read gets matched to that description from the thirty voices on file. Ask for the same script again in a different voice and that's the whole revision: a new sentence, not a new upload.

Two voices in one file

Label two speakers in the script and Chat Octopus assigns each one its own voice, so a two-person dialogue comes back as one piece of audio instead of two takes you line up yourself afterward. A third speaker pushes past what one read can hold. The script goes out in passes of two speakers at a time, and those passes get merged after. Slower, but it works.

Getting the name right

There's no separate field for pronunciation. If a name keeps coming back wrong, respell it in the script the way it should sound, Xiomara as "See-oh-MAR-ah", and the read follows the spelling on the page rather than the one you had in mind.

Getting the file

The audio comes back as a WAV file. Alongside it, two caption files land automatically: SRT, the format most video editors and platforms expect for subtitles, and VTT, the version browsers use for web video, both timed at the word level for captions that pop one at a time and at the segment level for a standard subtitle track. You never ask for them separately.

Putting it under the footage

The audio is already sitting in the thread next to your other files, so laying it under a clip or burning the captions in is the next message, not a trip to another app. Ask for the voiceover under your footage, or the captions burned into the video, and it happens in the same conversation the read came from.

How it works

1

Paste your script, describe the voice, and say how it should be read

2

Get the voiceover back with matching SRT and VTT captions

3

Ask for another read, or take the audio into your video

Frequently asked questions

Yes, and every read includes a fresh batch of daily credits with no card needed, and the audio comes back with no logo of ours on it.

No. Chat Octopus reads from a fixed set of thirty voices rather than a cloned one of yours. Say what you want, steady and clinical, or bright and energetic, and the read gets matched to the closest voice on file. Ask again in a different voice any time; there's no sample to browse first.

Yes, for two. Beyond that, one call can't hold a third voice, so a three-person script needs a second pass that gets stitched to the first once both are done. It costs you time, not quality.

It speaks the language your script is written in. Write the transcript in Spanish and you get Spanish. The direction line cannot override that, so translate the script itself rather than asking for it to be read in another language. Keep any bracket tags in English even when the script is not.

About five thousand characters in a single read, somewhere around five minutes of spoken audio. Longer than that and the script goes out in parts that get joined into one file, so there's no hard ceiling, just a point where it takes more than one pass.

A WAV audio file, plus SRT and VTT caption files timed to it automatically, at both the word level and the segment level. Nothing has to be requested separately or timed by hand afterward.

Nothing reads a name off a pronunciation guide; it reads exactly what's typed. Fix it at the source: respell the word phonetically in the script itself, "Xiomara" as "See-oh-MAR-ah", and the read follows the spelling you sent instead of the one you meant.

Keep reading

Related tools

Your next video is one conversation away.

Free account with credits included. No credit card, no learning curve.