Guide

Motion Graphics: What They Are, and How to Describe the One You Want

The name bars, title cards and counting numbers in other people's videos. What each one is, when a video needs it, and how to describe one precisely enough.

Updated August 5, 2026

Three things you have seen this week and probably cannot name. The strip that slides in under someone's face with their name and job title on it. The number that climbs from zero to 40,000 while a voice says forty thousand. The two seconds of logo that open every upload on a channel you watch.

Those are motion graphics: text and shapes that move, made rather than filmed. Nothing in front of a camera, nothing that existed before somebody decided it should. Someone chose what appears, where it sits and when it arrives, and the frames in between were drawn to match. That is the entire category, and it is a smaller idea than the phrase makes it sound.

The reason your videos do not have them is worth being precise about, because it is probably not the software. Learning to animate is a real obstacle and it takes months. But most people who reach for a shortcut get stopped one step earlier, at saying what they want, and no tool fixes that for you.

Which ones does your video actually need?

Four kinds cover almost everything a normal video wants, and none of them need an art director.

The name strip. A person's name and job title appearing under them the first time they speak, then leaving. Editors call it a lower third, because it lives in the bottom third of the frame. It earns its place exactly once per person, at the moment a viewer would otherwise be wondering who is talking.

The opening. Two or three seconds of something before the video proper starts: a title card if it is words on a background, a logo sting if it is your mark animating and getting out of the way. The near relative is a bumper, the same short thing dropped between sections rather than at the top. Worth making once if you publish every week, because the same two seconds go on the front of everything you post after it.

Words that move. A line of text landing one word at a time, or a figure counting up from zero rather than simply being there. The trade name is kinetic typography. The reason to use it is not decoration: motion buys attention for a phrase you would otherwise have to say twice.

A chart that builds. Bars growing from nothing while you talk over them, one after another. The growth is the whole point. A bar chart already at full height is a picture, and a picture gets read in half a second and then ignored.

Most weekly videos want two of these and get zero.

Why what comes back is never quite what you pictured

Every animation, however small, is a stack of decisions. Where each element sits. What it says, exactly. What color it is against what is behind it. The moment it arrives, how long it holds, the moment it goes. A five-second name strip is roughly twenty of those.

When you type "add a lower third with my name," you have made two.

The other eighteen still get made, just not by you. Whatever you hand the job to picks a position, a size, a typeface, a length, an entrance. Some of that is a fair guess from your material. Some of it is a plain fallback. Hand the job to Chat Octopus, where the description goes into a chat and an MP4 comes back, and a request with no frame shape in it returns 1920 by 1080 landscape, which is the wrong shape for every vertical feed. Then you watch it, dislike it, cannot say why, and ask again slightly differently. That loop is where the afternoon goes, and nothing in the drawing of it was ever the problem.

Four questions every animation answers, with or without you

Answer these and you are describing rather than gesturing.

Where. Each element as a place on the frame, in the words you would use pointing at a screen. "Bottom left, above where my captions sit." "Centered, taking about half the width." Position is also the only way to prevent the most common collision in short video: captions burned into the picture live in the lower third too, so a name strip and a caption line will sit on top of each other unless one of them is told to move.

What it says. The exact words. Names spelled the way that person spells them, titles as they appear on the business card, real figures rather than placeholders. Nothing can proofread a name it was never given.

What color. Name the pair, not the mood. "Cream text on a deep green bar" is a color scheme; "clean and modern" is a feeling. If you have brand hex codes, paste them. There are floors here worth knowing even if you never put them in a request: text needs a contrast ratio of at least 4.5 to 1 against what sits behind it, and 3 to 1 once type is 24 pixels or bigger. Gray on a slightly different gray does not clear that, which matters because "subtle" is usually a request for exactly that.

When. Timing in plain seconds. How long it takes to arrive, how long it holds, whether it leaves at all. Half a second is a brisk entrance. Anything a viewer has to read needs about three seconds of stillness after it lands, and a name strip that holds for eight is normal.

Some of that gets checked for you, and the checking is worth understanding because it is weaker than it sounds. The animation is written down as a plan before any frames exist: every element as a box with its position, every text-and-background pair with its contrast ratio beside it, every arrival and exit as a moment. Then the plan gets walked. Is each box inside the frame, do two things overlap that should not, does the piece that slides away actually clear the edge instead of stopping just short of it, does each pair clear those floors.

The thing doing the checking is the thing that wrote the plan, so it is arithmetic marking its own homework: it catches most collisions and not all of them. And afterwards, once the render has run and every frame has actually been drawn, a handful of stills gets pulled back out of the finished video and read against the plan, about four of them, so a defect that only happens between the sampled moments goes out with the file.

None of it can catch that you wanted the strip on the right, or that the tagline on your title card is last year's. Watch the first version yourself.

Turning "put my name on it" into something specific

Here is the request most people send:

can you put my name on this interview

And the same request with the four questions answered, which takes under a minute to write:

Name strip for an interview. "Dana Ruiz" on the first line, "Head of Research" underneath. Slides in from the left over half a second, holds eight seconds, slides back out to the left. Cream text on a deep green bar, bottom left, sitting above my captions. Vertical, 1080 by 1920.

None of that is jargon, and none of it is extra. Every line in it is a decision that gets made either way.

The same move works on the others. "Animate my signup number" becomes "the figure 40,000 counting up from zero over one second, then holding, black background, one orange accent, nothing else on screen." "Make me an intro" becomes "three seconds: my logo scales up from small with a slight bounce, the tagline fades in underneath a beat later, both hold, off-white background, silent."

When there is a voice over the top, the timing question has a better answer than a number. Ask for the voiceover in the same conversation and the narration comes back with a timing for every word in it, so reveals can be pinned to the words themselves. "The number lands on the word forty" becomes a thing you can ask for rather than a thing you nudge one frame at a time.

Your logo, your colors, your font

Upload the logo, the product photo, the clip, the music bed, then point at them in the same message: "use the logo I attached, bottom right, fading in after the title lands." Your own assets are most of the distance between a sequence that looks like yours and one that looks bought.

Fonts have a hard boundary, worth knowing before you plan a look around one. Name a well-known font and you get it, fetched by name from a large free type library. Ask for something outside that library and the render fails on it, and the usual recovery is a swap to whichever available font sits closest, which you will meet as a finished video rather than as a question. Uploading your brand's licensed font file is not a path to rely on either. So pick the substitute yourself: name the closest well-known face in the request and the decision stays yours.

What comes back, and the one thing it cannot be

An MP4, no watermark, at whatever frame size you asked for. Say "vertical for Reels" and it is 1080 by 1920. Square is 1080 by 1080. Say nothing and it is 1920 by 1080 landscape at thirty frames a second, so name the shape unless the video really is going somewhere wide. Ask for sixty frames a second when the motion is fast, which is the only time the extra smoothness shows.

Sound can live inside the file: narration, a music bed under it, captions burned into the picture. Ask for the caption file on its own instead and you get an SRT, because burned-in captions and a separate subtitle file suit different destinations.

The limit that surprises people: the background is part of the video. There is no transparent version to drop over your own footage in another editor. So when the name strip needs to sit over a shot of the person talking, put the shot inside the animation rather than the other way around. Upload the clip, ask for the strip over it, and one MP4 comes back with both. If the footage still needs cutting, that happens in the same conversation and nothing gets exported anywhere in between.

When it comes back wrong

It often will on the first pass, and the fix is a sentence. "Hold the title one second longer." "Move the strip up, it is hitting the captions." "Count from zero, do not fade the number in." "Slow the whole thing by about a fifth."

A follow-up edits the sequence you already have rather than regenerating one, so the parts you did not mention usually survive the round. That is the difference from a generator that starts over each time and hands you a different animation, where half of every note you write goes on defending the bits that were already right.

The honest cost sits in the render. There is no partial re-render: any change redraws the whole piece from the first frame, and you watch the progress bar climb while it happens. Nothing caps how long a sequence can be, but long and busy is where it strains, and a render that stalls gets called off rather than shipping a half-finished file. The practical answer is to work in scenes. Build a two-minute explainer as six short pieces and stitch them at the end, and then a note on scene three costs you scene three instead of two minutes. Sections are easiest to draw when the script already has them, which is what the explainer template starts you with: it writes the narration first and derives every visual from it, so the four questions get asked line by line rather than about a two-minute blur.

When you should hire someone instead

When the motion is the product. A title sequence, a brand film, anything where the animation is the reason people are watching. That work wants hands on individual frames, fixing where a thing sits at one moment and where it sits at the next, which is what animators mean by keyframes. Describing is the wrong instrument for it.

When the thing has to be filmed. Motion graphics are drawn: type, shapes, charts, your own assets. A coastline at sunset or a person on camera saying the words is footage, so generate the shot or film it, then cut the two together.

When you need a drawing at an angle. A simple drawing, the flat kind made of shapes and outlines, is reliable straight on or from directly above. Ask for a three-quarter view, the angle you would get photographing a toy car from its corner, and the result tends to come apart: a wing hidden behind the body, one eye missing, a wheel that never got drawn. Ask for side-on or top-down and it holds up.

When somebody else has to take it over. You get a video and a conversation that remembers the video. There is no layered project file to hand to a motion designer next quarter.

Start with the one you keep noticing

Pick the animation you envy most in someone else's video and write down four things about it: where it sits, what it says, what color it is, and when it moves. The first version comes back in the conversation you asked in, and the second version is one sentence after that.

A free account comes with credits, no card needed. Put the description you just wrote into a chat and see how close the first pass lands.

Related tools

Your next video is one conversation away.

Free account with credits included. No credit card, no learning curve.