Guide

Finding One Moment in Two Hours of Video

Ask footage where something happens and get a time back. How exact those times are, how much video a single pass can read, and where it stops helping.

Updated August 6, 2026

Two hours of recorded call. One decision inside it, made somewhere after the first hour, and you need the wording rather than your memory of the wording. You know exactly what you are looking for and nothing about where it is. That is the worst amount of knowing: too little to jump to, too much to sit through the whole thing again.

So you scrub. Drag the marker along the bar, land past it, drag back, overshoot, sit through ninety seconds of someone's screen share, drag again.

You can ask instead. Put the file in a Chat Octopus conversation, type the question you have been asking yourself, and the answer arrives with a time on it. At 1:47:12 someone says "so we're saying the fourteenth, final answer," and nobody objects after it.

"AI watched your video" is a claim that has let most people down at least once, and it is not what happens here. Nothing watches. The picture gets sampled, meaning still frames pulled out at known times and read one at a time. The audio becomes words, each carrying the moment it was spoken, which is a transcript with a clock attached to it. Every answer is assembled from those two, which is why every answer can carry a time, and why any of them can be checked.

The questions that work have an answer at the end of them rather than a summary:

  • Where does anyone commit to a number?
  • Does the product ever actually make it into the picture, or does it sit off the left edge the whole time?
  • Which stretches of this have nobody talking?
  • What is on the whiteboard behind him, and can you read the third line?

How exact is the time it gives you?

Approximately right, and how approximately depends on which part of the file the answer came out of.

A run-through of the picture reads about one still a second, which is also why it only reaches about half an hour of footage at a time. So a time that came from something seen, a gesture, a product finally coming into the picture, a slide change, is good to about a second either way. To find the moment, that is plenty. To place a cut, it is not: video runs at twenty-four or thirty still pictures a second, each one a frame, so a second is thirty frames, and a cut thirty frames late is a cut anyone can see.

Two things in the file are known exactly, and both make better anchors than the picture.

The words. Each one carries its own start and end, so a boundary put on a word is on that word, not near it.

The shot changes. A shot is one unbroken run of camera between two cuts. Where one ends and the next begins is found by measuring how much the picture changes from frame to frame, so those edges are exact rather than described.

The working method is two moves instead of one. Take the rough time from the picture, then snap it to whichever exact anchor is nearest: the start of a word, a gap in the speech, the shot change. "Find where she picks up the product, then start the clip at the beginning of the sentence she is saying over it" is one message, and what comes back is a boundary you can publish.

If the picture time itself has to be tighter, name a range to look in. A hunt across a whole file samples twice a second. A hunt inside a range you name samples five times a second.

How do you know it is not making it up?

Because you can look at exactly what it looked at.

Every claim about the picture came from a still at a stated time, so ask for that still. "Show me the frame at 1:47:12." It comes back as an image at the video's own resolution, and either the thing is in it or it is not. The check takes five seconds and is worth doing on the first two or three answers, until you have a feel for what this is reliable at.

Resolution matters more than it sounds. Ask what is in a file and what you get is a map of its shots, one still described per shot, and those stills are read at 512 pixels wide: enough to tell a wide shot from a close-up, nowhere near enough to read the small print on a slide. A frame you ask for by time is not shrunk like that. So when the job is reading words off the screen, ask for the frame at that second and have it read the image, rather than hoping the map caught it.

Spoken claims check the same way. Ask for the transcript around that time and read the sentence yourself.

One failure is worth knowing in advance. Over music, room tone or wind, the pass that turns audio into words sometimes returns words nobody said. It is a known behavior of the technology rather than a rare glitch, and there is a check for it: implausibly few words a minute, one phrase repeating, a single "word" that supposedly lasted several seconds. When the check trips, the words are thrown out and the clip gets read as pictures instead. If a clip you know was silent ever comes back with a confident transcript on it, that is what happened, and none of it is real.

How much video can one pass read?

A single pass over the picture gets a budget of about 1,800 still frames. All three ceilings people run into are that one number, spent at different densities.

  • A run-through of what happens samples once a second, so it reaches about thirty minutes.
  • A hunt across the whole file for one thing samples twice a second, so it reaches about fifteen minutes.
  • A hunt inside a range you name samples five times a second, so the range itself has to come in under about six minutes.

Past the budget the pass does not run at all. Nothing skims instead, quietly and worse. The work becomes naming a stretch: "tell me what happens from 1:10:00 to 1:35:00."

There is a size ceiling as well, and it bites sooner than people expect. The passes that read the file as video, the run-through and the hunt and the close look at a single take, take files up to 200 MB. A forty-five-minute recording at 1080p is usually well past that. Name a stretch and that stretch gets cut out of the file and read on its own, which is the same move again. For long footage, naming the section is not a way to go faster. It is how you get an answer at all.

Asking for one frame is not bounded by any of that. A still at a time you name gets pulled straight out of the file where it sits, so "show me the frame at 1:47:12" behaves the same on a recording well past 200 MB as on a small one. The map of the shots is built the same way.

The words carry no equivalent ceiling. A long recording is split into eight-minute pieces, transcribed, and stitched back together with every time returned to the original clock, so a two-hour meeting transcribes end to end.

That gap decides the order to work in. Get the transcript first. It costs one request, it covers the entire file, and it tells you which ten minutes are worth pointing the picture at. Then aim.

What if it is forty clips instead of one long file?

Same principle, different pain. This is the hour before the edit rather than the edit itself: opening files one at a time to remember which one had the good take.

Upload the batch and ask what is in them. What runs first is deliberately cheap. Every file with an audio track gets transcribed and measured for how much of it has anyone speaking, and that measurement decides how the file gets read from there. Half speech or more and it gets read as words. Almost none and it gets read as pictures. Anything in between gets read both ways, stretch by stretch. Nothing burns a picture pass on a talking head, and nothing goes hunting through words on a silent drone shot, which is the only reason forty files is a sensible thing to ask about at once.

That cheap pass answers more than people expect it to. What fraction of each clip has anyone talking. Where the gaps of two seconds or more sit, which is your dead air, the pause before someone finally answers, the eleven minutes of setup before the event starts.

Ask for a rough cut instead and the sorting gets a purpose to serve. Worth asking for by that name: Rough Cut From Footage is a saved instruction you can open and read, and the sorting scheme comes with it. Clips separate into the main speaking takes that carry the story, called A-roll, the supporting shots you lay over them, called B-roll, the wide shots that establish where you are, and the broken ones. The broken pile comes back with a reason attached to each: soft focus, audio that cannot be saved, framing that missed the subject. That is the pile worth reading, because a reason is a thing you can argue with. It also names what the shoot never got, which is what you actually want to know before the next one.

Expect a question before any of that arrives. What the cut is for, roughly how long it should run, who is watching it. Those three are not worth guessing at, so they come back at you instead. Any edit stops for that kind of question when the call is yours to make, and nothing gets cut until you say go.

What will it get wrong?

Anything shorter than the gap between samples can fall through it. At one still a second, a half-second glance at the lens, a single bumped frame, a graphic that flashes for twenty frames: each can sit between one sampled moment and the next and never be seen. Naming a range makes it look five times a second instead of once, which is better rather than certain. That is what sampling is, not a setting waiting to be turned up.

Nothing sweeps the file for on-screen text. No pass walks an hour of video gathering every slide title, caption and name strip into a list. Reading words off the picture works when you aim it at a time. Asking for all of them across an hour is asking for something that does not exist. For anything spoken aloud, the transcript is both more accurate and genuinely complete.

It cannot count scenes. That map of the shots tops out at about twenty stills for the whole file, whether the file runs four minutes or forty. It is a map, and a good one for "roughly what is in here, in what order". It is not a census, so "how many scenes does this have" has no honest answer coming out of it.

Keep and filler are opinions. A run-through marks stretches worth keeping and stretches that are dead time, each with a reason. The reasons are the part to read. Overrule them freely.

Watch the cut before you publish it. That is the rule the editing side already works to, not a line of small print, because the things sampling misses are exactly the things a person catches in the first few seconds of watching.

Turning the time into the cut

The footage is already in the conversation, so the next request is the edit itself. Nothing gets re-uploaded and nothing goes out to another tool first.

The move worth knowing, if you cut for a living: once you have chosen a take, ask what is wrong with it. A close pass runs over that take plus about two and a half seconds either side, five looks a second, and comes back with times for the things you would otherwise find yourself. Someone glancing at the lens. Crew or gear in frame. A bump or a shake. The picture hunting for focus. Framing that breaks.

The trim itself is worth understanding, because it changes what you ask for. A defect inside a take you are keeping is not removed, it is cut around: everything up to the bump, then everything after it, both from the same clip. The join gets a fade of about thirty milliseconds so it does not click. You never describe the defect. You describe where the clip stops and where it starts again.

Boundaries that land on speech are word-exact, because the word timings are. "Cut from the start of that sentence to the end of the next one" is a boundary rather than an estimate. Boundaries on the picture are the approximate ones, which is why they get snapped to a shot change.

You will see the plan before anything is cut. Which parts are being kept, in what order, roughly how long the result runs. Editing is subjective enough that agreeing in words beforehand is faster than disagreeing with a finished cut. And if the finished clip is heading somewhere people watch on mute, captions are the next sentence, not the next project.

The length of the file decides your first move, and it is the only setup decision worth thinking about. Under half an hour, point the question straight at the picture. Over it, take the transcript first and let it tell you which ten minutes deserve a closer look. Either way, what you type is a question with one answer in it.

The best file to try this on is one you already know the answer inside. Upload it and ask something you can check, then go and look at the frame it names. A new account is an email address, and the credits this costs are already sitting on it.

Related tools

Your next video is one conversation away.

Free account with credits included. No credit card, no learning curve.