What B-Roll Is, and How to Get Some When You Only Shot Yourself
A-roll is the take carrying the words. B-roll is what the picture cuts to while those words keep running. Four ways to get it from one recording.
Updated August 16, 2026
In every interview you have ever watched, the person answering is not on screen for the whole answer. Their voice keeps going. The picture leaves their face and shows the workshop they are describing, or their hands, or the street outside, or a photograph from the year they are talking about, and then it comes back to them.
The face is the A-roll: the take that carries the words. Everything the picture goes to while those words keep running is B-roll.
You have watched that trade a thousand times and never had to name it, because it is built not to be noticed.
Now play your own recording back. One camera, one distance, one framing, for eleven minutes. Every shot is the same shot, because the only thing you shot was A-roll, and there is nothing to cut to. That is not a problem you can rearrange your way out of, which is why the last hour has felt like moving furniture.
What is B-roll actually doing up there?
Four jobs. Which one you are doing decides what you go out and get.
It shows the thing. You say "the old warehouse" and the picture shows an old warehouse. The viewer stops holding your description in their head and just looks at it.
It covers a cut. This is the one that matters to you today. Cut out the forty seconds where you lost your thread, join what is left, and your head jumps to a new position at the join. That jolt is a jump cut, and it is the single loudest signal that a video was edited by someone in a hurry. Put two seconds of something else over the join and it vanishes, because the two head positions never appear next to each other.
It sets a place before anyone speaks. The building, the room, the street, so your first sentence lands somewhere instead of nowhere.
It gives the eyes somewhere else to be. Attention on one unmoving face runs out sooner than the person on camera ever believes.
A shot doing none of those four is decoration, and decoration is what makes a cut look homemade. There is a name for the thing it would be decorating. The spine is the line of speaking takes, trimmed and joined end to end, that holds a video up, and B-roll rides on top of the spine without ever carrying it. It goes on where it earns its place and nowhere else.
Where does it go, and how much do you need?
Get the transcript of your own recording and mark every sentence where you name something the viewer cannot see. "For example." "When I was there." "It looks like this." Those sentences already asked for a picture and never got one. Each one is a cutaway, meaning a shot you cut away to and come back from without the sentence ever stopping. That is your shot list.
Then mark the joins. Every place you cut something out is a jump waiting to be covered.
Two to four seconds each is usually the entire job. A cutaway that outlasts the sentence it belongs to stops supporting the video and starts being a different one. And if a moment feels long while you are watching it, it is long, which is a thing the person who shot it is always the last to accept. One cutaway per idea. Not one per sentence.
Keep your voice running underneath all of them. That is the part people get wrong on the first attempt: they drop the new shot into the sequence, which replaces your picture and your sound for those seconds, and the video lurches. What you want is the shot over the picture with the audio left alone, so ask for exactly that.
Is B-roll even the fix?
Check this before you spend an evening or a penny on shots, because sometimes the recording is not visually boring. It is just slow, and cutaways over a slow video produce a slow video with pictures on it.
The tell is whether you can say what any given thirty seconds is for. If you cannot, no shot of a coastline is going to rescue it. Cutting is the fix: pull the false starts, the long pauses, the sentence you said twice, the ninety seconds of preamble before the actual point. It always takes out more than you expected to lose, and it does more for the video than any shot you could lay over it. Cut first, then look at what still needs covering, and watch the shot list get shorter.
The other honest case is that you are covering something you should re-say rather than hide. A jump cut over a fumbled explanation is still a fumbled explanation. Thirty seconds re-recorded on the same camera in the same seat beats every workaround for it.
Where do you get B-roll with one recording and no second camera?
Cheapest first.
Inside the file you already have. If you recorded in 4K and are publishing at 1080p, a second angle is already sitting in your own take. Drop the recording into a Chat Octopus chat and ask for that stretch cropped to a tighter framing on your face, which is what an editor means by a punch-in. Still you, so not strictly B-roll, but it breaks the sameness and nothing about it is charged by the second, because no new footage is being made. The limit is arithmetic: shoot 1080p and publish 1080p and there are no spare pixels to crop into, so the punch-in comes out soft.
Things you already own. The screen recording of the thing you are talking about. The clips sitting on your phone from that week. Photographs, which can be made to move instead of sitting there: hand over one to four stills and ask for a slow push in, and back comes a moving shot of three to ten seconds. Photos of real people work on that route, with one geographic exception. In the UK, Switzerland and the EEA, a starting picture with a person in it is refused.
Ten minutes and your phone. Your hands doing the thing. The desk. The object you keep mentioning. The walk to the door. This is what B-roll was for decades before anything could generate it, and it is still the best-looking option you have, because it is genuinely your world and it matches your lighting.
Shots that do not exist. The coastline, the warehouse interior you cannot get into, the year 1974, the product you do not own yet. This is where a searched library used to be the only answer, and where describing the shot now beats going and looking for it.
What can a generated shot actually be?
Described in words, a shot runs four to twelve seconds and comes back as an MP4 at 720p, or at 480p if you are only checking whether the idea works. You choose the shape of the frame: wide for YouTube, tall for a phone, square for a feed, and four more, boxier or more extreme. A cutaway for a vertical video arrives vertical rather than cropped to fit afterwards. Sound comes with it unless you say you want it silent.
Generated from your own pictures instead, a shot runs three to ten seconds. Up to four starting pictures, each under seven megabytes and fourteen for the set. Two shapes only on this route, wide or tall, so anything square has to come from a described shot. There is no resolution choice either, and no way to switch the sound off, so plan on muting it when it goes in.
Twelve seconds is the ceiling on any single shot and no phrasing gets you past it. A twenty-second shot is two shots stitched together, which for B-roll almost never comes up, since your cutaways are two to four seconds.
Then there is cost, because "just generate it" sounds like it should be free.
Generating is billed by the second of finished footage. A four-second described shot costs fifty-eight credits, an eight-second one a hundred and sixteen. Credits are what usage gets counted in here, and a new account starts with five hundred, then picks up thirty more each day you are working, banking up to two hundred and ten. That is about eight short cutaways before the daily top-up is what sets your pace. Everything around them, the trimming and the joining and the laying-in, is not billed by the second at all.
One gate to know before you plan around it. Making new picture needs an account signed in with Apple or Google, and that covers three requests rather than two: the described shot, the shot built from your photographs, and asking for something already inside a filmed shot to be changed. Sign up with an email and a password from late June 2026 onwards and those three come back with a no. Accounts opened before that were let through and still are, whatever they signed up with. Cropping, cutting, trimming and laying shots in sit outside the gate entirely.
How do you ask for all this?
The recording goes into a chat, and the work happens in one place rather than four.
Ask where things are first. "Find every place I describe something the viewer cannot see, and give me the timestamps." That is the same question you would ask to find one moment in hours of footage, pointed at your own script, and what comes back is the shot list you were going to write by hand.
Then ask for the shots. Describe each one the way you would brief a camera operator: subject, what it does, where it is, how it is lit, what the camera does. "A slow push in on rain hitting a workshop window, grey afternoon light, no music" gets you something usable. "Cool B-roll of rain" does not, and the gap between those two is the same gap that decides whether an animation comes back looking like the one in your head.
Then ask for them to go in. Name the timestamp, name the length, and say the audio stays. Nothing gets re-uploaded between those three steps, because your recording, your shots and your finished cut are all in the same conversation.
One thing not to ask for yet. A rough cut is the first pass at a video, where everything usable is found and put in the right order before anything is polished, and you can ask for one by name. It sorts a pile of footage into speaking takes, supporting shots, wides that set a place, and the broken ones, and it hands back a reason for every clip it threw out. It is built for the pile. Give it one short clip and it stops and asks for the rest of the material, because there is nothing there to sort. With one recording, ask for the specific thing you want done to it.
What goes wrong the first time?
The generated shot looks softer than your footage. It is softer. A described shot tops out at 720p, and dropping it into a 1080p cut scales it up to fill the frame. On a two-second cutaway nobody notices. On a six-second hero shot they will, so keep generated material short and keep your own footage on screen for the moments that need to look sharp.
Black bars around a cutaway, or the whole video in the wrong shape. One disagreement, showing up twice. When shots are simply joined one after another, anything that does not match the frame is fitted inside it and the space left over is filled with black, so a wide cutaway in a tall video gets a bar top and bottom. And the frame everything else is matched to comes from the first shot in the line, so a generated opener sitting in front of your camera footage hands its shape to the whole file. Ask for cutaways in the shape you publish in, keep your own footage first, and if a mismatched shot has to go in anyway, ask for it cropped to your frame before it is joined, which loses the edges instead of adding black.
A cutaway brings its own soundtrack. Generated shots come with sound of their own: wind, footsteps, the hum of a room you have never been in. On a described shot you can ask for silence up front. On a shot made from your photos you cannot, so say you want it muted under your voice when it goes in.
The first one to make
The first cutaway is also the easiest one to judge. Find the place in your recording where you say "for example", and put a picture of the example there. Watch it back. If the video got better in those two seconds, you now know exactly what the rest of your list is worth, and you can ask for the shots and the edit in the same conversation rather than going looking for a library to subscribe to.