Text to Speech That Comes With Captions
You paste the script in, and what comes back isn't just an audio file. The captions that match it, word-timed and segment-timed, land in the same reply, already synced to a recording that didn't exist a minute ago.
Try asking
How long a read can be
A script up to about five thousand characters comes back as one read, roughly five minutes of audio. Send something longer and it goes out in parts, each one read separately and then joined into a single file. Keep the direction line, the instruction you put before the script telling it how to read, identical across every part, or the joins land as one narrator changing character mid-sentence.
Picking a voice
There's no dropdown of voices to click through and no sample to preview first. Say what you want instead, warm and gentle, firm and mature, and the read gets matched to that description from the thirty voices on file. Ask for the same script again in a different voice and that's the whole revision: a new sentence, not a new upload.
Two voices in one file
Label two speakers in the script and Chat Octopus assigns each one its own voice, so a two-person dialogue comes back as one piece of audio instead of two takes you line up yourself afterward. A third speaker pushes past what one read can hold. The script goes out in passes of two speakers at a time, and those passes get merged after. Slower, but it works.
Getting the name right
There's no separate field for pronunciation. If a name keeps coming back wrong, respell it in the script the way it should sound, Xiomara as "See-oh-MAR-ah", and the read follows the spelling on the page rather than the one you had in mind.
Getting the file
The audio comes back as a WAV file. Alongside it, two caption files land automatically: SRT, the format most video editors and platforms expect for subtitles, and VTT, the version browsers use for web video, both timed at the word level for captions that pop one at a time and at the segment level for a standard subtitle track. You never ask for them separately.
Putting it under the footage
The audio is already sitting in the thread next to your other files, so laying it under a clip or burning the captions in is the next message, not a trip to another app. Ask for the voiceover under your footage, or the captions burned into the video, and it happens in the same conversation the read came from.
How it works
Paste your script, describe the voice, and say how it should be read
Get the voiceover back with matching SRT and VTT captions
Ask for another read, or take the audio into your video