# Chat Octopus vs Canva for Video

Chat Octopus turns a long recording into captioned vertical clips from a plain-language brief, and redoes them in a sentence. Canva keeps graphics and decks.

The video work probably arrived in Canva by accident. The tab was already open, there was a button for it, and the first one came out fine.

Chat Octopus has no canvas. You upload a file, say what should happen to it, and a finished file comes back: an MP4, an MP3, a PNG, a caption file. Nothing to place, nothing to drag, no page to arrange. So anything you lay out by hand is out. With a recording there is nothing to arrange in the first place: the footage already runs in an order, and your job is choosing what to keep.

You post a carousel and you are done with it. A video comes back at you. Someone watches it, asks for one change, and you cut it a second time.

## The four jobs worth moving

**Cutting a long recording down.** Upload the webinar or the interview and ask which parts stand alone. You get the clips, and a line under each one saying where it starts, where it ends, and why that moment. Seeing the reasons teaches you how to ask next time, so a week later "more like the second one from last week" is enough of a brief.

Give it numbers. A vague ask comes back vague. Two to four clips out of a long video, not eight, because weak clips cost a channel more than they earn it. Thirty to sixty seconds each. Start where the point starts, and end on a finished sentence rather than a fade. A moment that needs two minutes of setup before it lands is not a clip, and asking for one anyway is how you get a bad one.

**Captions.** Chat Octopus listens to the recording, times every individual word rather than every line, and draws the words over the picture. Keep them few: seven words on screen is the ceiling, and each group of words should hold for a second, maybe three. Break where the phrase breaks, never mid-phrase, so no line ends on "we tried it and". In a vertical frame, sit the caption strip, the band of text over the picture, higher than looks right, because the platform prints its own buttons across the bottom. You get the captioned MP4 and a standalone SRT file, [the plain-text list of lines and their timings that a platform reads off](/guides/captions-vs-subtitles), with timings that match what is burned into the picture.

**A vertical version.** Turn landscape footage into the tall shape a phone feed wants, or into a square, and turn a tall clip back to landscape when a client asks for one. Hold it to one rule: a vertical version should be cropped in on the subject, not a wide picture shrunk into the middle of a tall black rectangle.

**Voice and audio.** Hand over a script and it comes back read aloud, in a voice you pick. Send a recording with a fridge humming behind it and ask for [the hum taken out](/guides/remove-background-noise-from-audio). Ask for quiet music running under the whole thing, and you get that. Some weeks what you were asked for is [the words themselves, typed out](/guides/transcribe-audio-to-text), and that is a job you can hand over too. One honest gap here, because it bites interview work: nothing labels who is speaking. Two voices on one track come back as a single column of text with no way to tell which line belongs to which person. Recording each person to their own file is the only real fix.

## What happens when the edit changes

Say the change in the same thread. "Lose the first eight seconds of clip two." "Captions in yellow." "Same cut, square." The conversation still holds your source files and the decisions that produced version one, so a change is a sentence instead of a rebuild.

Be clear about what that does not buy you. The video gets made again from the top, and it costs credits again, the same way the first version did. Credits are the unit work is counted in here. A redo is cheap in your hours and not free in your balance.

A picture is the exception. Name the part you mean and that part changes: move the person on the left over, swap the background out, make the jacket red. What comes back is the picture you already had, corrected, rather than a fresh one to choose all over again.

You also do not have to sit through it. Close the tab, and the thread and the files are in your account when you come back.

## What belongs in a design app

Anything composed by hand. A chat cannot lay out a carousel, drag a headline around until the balance looks right, or fan out a page of covers for you to choose between.

So Canva keeps the graphics, the carousels, the ads and covers, the decks and documents and whiteboards and print, anything where the arrangement is the work, and two of you in the same file at once leaving comments on it. It also keeps your brand kit, the colours, fonts and logo you saved once and now apply by picking them. Nothing here does that job, and pretending otherwise would strand you in week two, when the third caption comes back in the wrong grey.

One word crosses over and means something else on the way. A template here is [wording that worked once, kept under a name and rerun](/guides/using-templates), not a layout to start a design from.

Thumbnails work here, up to the point where the text has to be exact. Ask, and four options come back, each with a line explaining what it is going for and who it is aimed at, which is [the hard part of making a thumbnail alone at eleven at night](/guides/make-a-youtube-thumbnail). But the words in a generated picture are painted in rather than set in type, so read the spelling on the result and expect a second pass when a letter doubles. If the title has to be exactly right in exactly your font, that belongs in Canva. Making the picture in one place and the words in another is a reasonable way to work.

## What about your brand colours and fonts

They go in a note: a standing instruction saved to your account that carries across every conversation. "Captions are white on a black bar. My accent is #FF5A1F. I publish vertical." It gets applied when it fits what you are asking for, so treat it as a strong preference, not a locked setting. When a result comes back off-brand, saying so in the thread fixes that result. Write the correction into the note as well, so the next thread starts closer to right.

Writing that note needs the actual values, and if your colours have lived in a brand kit for two years you have probably never typed one. Hand over a graphic you already made and ask what is in it: the codes come back, and the typeface gets named too.

Fonts are the harder half.

[An animated title or a caption strip](/guides/create-motion-graphics) uses a real typeface, though the choice narrows to widely available families it can fetch for itself. Do not count on handing over the licensed font file you bought and having it used. Pick the nearest common face yourself and ask for it by name, rather than letting the choice get made without you.

Colour behaves better. Give an animated title the exact colour code you keep for your brand, the six characters starting with a hash, and it lands on that colour. Give a generated picture the same code and it lands nearby.

## Where it stops

Four things, and none of them get better on a bigger plan.

You describe, you do not drag. Asking to move something over works. What you cannot do is nudge it a hair and watch it land: every adjustment is another ask and another result to judge. Chasing an exact position that way is slow, and it finishes close rather than right, so the fussy last stretch of a layout is better done somewhere you can see it move. Resizing, cropping and rotating are the exception. Those are exact, and they land on the number you name.

Video generated from a description runs four to twelve seconds, and tops out at 720p, a step below the 1080 most phones shoot. Built from your own images, three to ten seconds. That is enough for the shot you drop over someone still talking, or the few seconds of movement before the first word. It is not a supply of footage.

Nobody else can be in the conversation with you. Review still happens where it always did, in a message thread somewhere else, and there are no comments pinned to a frame.

And the waiting does not disappear. A long job takes what it takes. What changes is that the wait costs you nothing but time: no one is making the cuts by hand while the clock runs.

## What it costs on top

You are not cancelling anything, so the number that matters is what gets added.

Free is a real tier rather than a trial. A new account starts with 500 credits, then thirty more arrive each day for as long as you keep making things, with no card and no deadline on the offer.

The daily credits expire after seven days, and your balance stops at 210. They are meant to be spent, not banked.

Paid starts at $19 a month for 2,000 credits, and $49 for 6,000. Pictures, generated clips, voices and music are what cost the most. Every charge appears in the chat as it lands, so a run that got costly is something you watch happen rather than discover on an invoice. The [per-job numbers and a calculator](/pricing) are on the pricing page.

## The test that settles it

Pick the recording first, and if it is a podcast rather than a webinar, [our comparison with Descript](/compare/descript-alternatives) is the closer read, since that one edits by cutting words out of the transcript. Should carving long streams into shorts be the whole job rather than one of six, [the Opus Clip comparison](/compare/opus-clip-alternatives) covers a tool that ranks the clips it finds and does nothing else.

Otherwise: take one recording you have already published, so you know what good looks like on it. Ask for two vertical clips with captions and a matching SRT, then open the result beside the version you cut by hand.

If your notes are "start eight seconds later" rather than "start again", the job is worth moving. Then change one thing and send it back.
