Video creation
Make AI video clips match: a six-shot continuity walkthrough
To make separate AI video clips match, define what must stay consistent, plan the ending of each shot and the beginning of the next, and approve reference frames before generating motion. Review the joins in an editor. Repeating a prompt can help communicate intent, but it does not guarantee matching objects, lighting, or action.
In this article
Six attractive clips can still make one confusing video
A pencil changes color between shots. An empty cup becomes full before anyone puts anything inside it. A hand moves toward the left edge of the screen, then enters the next shot from the opposite direction. Each clip might look convincing on its own. Together, they ask the viewer to reconstruct a scene that no longer makes sense.
That is a continuity problem: the details that connect one moment to another have changed without an explanation. Adding cinematic, realistic, or professional to every prompt does not specify those connections. I would solve the joins before requesting more elaborate camera movement.
This walkthrough builds a fictional, 24-second desk-reset video. The desk, objects, shot descriptions, and revision results below are illustrative. I have not generated or tested these clips, and the sequence is not evidence that any model can reproduce every action reliably. It is a production plan you can adapt and check against your own outputs.
The intended viewer should understand a simple transformation: four loose pencils go into a cup, a notebook moves into the cleared space, and the desk is ready for work. There is no product claim, before-and-after performance claim, or complicated story to hide behind. If that small sequence is unclear, the next useful step is fixing the sequence, not buying another effect.
You need a place to write the shot plan, an approved image or video tool, and an editor that can trim and assemble clips. A phone recording can replace a difficult generated action. The finished piece does not need to be entirely generated to be useful.
Choose the change the viewer must actually see
My starting question is not which model makes the prettiest footage. It is what must be visibly different at the end. For this example, the answer is specific: the four pencils are upright inside the same cup, and the same closed notebook sits in the space they previously occupied.
That sentence gives the video a testable purpose. A montage of attractive desks would not pass. Neither would a clip where the pencils disappear while a caption says organized. The audience needs enough visual information to connect the starting arrangement, the actions, and the final arrangement.
Decide the format before making reference images. I would use a vertical composition for this example, leaving comfortable space around the cup and notebook. The exact export dimensions belong in the editor and the generation tool's supported settings. Typing an aspect ratio into a prompt is not a substitute for selecting it in the interface.
Keep the creative brief separate from the generation prompt. The brief can include the audience, the intended lesson, the total edit length, and the final caption. A single clip prompt only needs the information required for its particular shot. Sending the entire production plan with every generation makes it harder to see which instruction caused a problem.
Also decide what is outside the brief. We do not need a visible face, a logo, a readable notebook title, or spoken dialogue. Removing those requirements is an editorial choice for this exercise, not a claim that models cannot create them. It leaves more attention for the action the video is supposed to explain.
Purpose: show a small desk reset, not advertise a product. Final edit: 24 seconds, vertical. Beginning: four loose yellow pencils; empty white cup; closed red notebook at the back. Ending: all four pencils in the cup; notebook in the cleared center. Caption added in editing: Clear a little space to begin. Audience takeaway: understand the two organizing actions without sound.
Write down the details you refuse to let drift
A continuity card is a short record of what should stay the same. For our desk, that means a white cylindrical cup without a handle, four yellow pencils, a closed red notebook, a pale gray tabletop, and light falling from the viewer's left. It also records where those objects start.
Make the descriptions observable. An elegant workspace leaves many possibilities open. A white cup without a handle is something you can inspect. Four pencils is a count. A closed notebook is a state. Those details help a reviewer decide whether a beautiful output still belongs in this particular sequence.
The card should distinguish permanent identity from changing state. Pencil color stays fixed; pencil location changes. The notebook stays closed but moves forward. A rule that says preserve everything conflicts with a shot whose purpose is rearranging the desk. Specify the permitted change so the instruction is not fighting itself.
I would keep the camera on the same side of the desk throughout this first exercise. The cup stays on screen right even when the framing gets tighter. This is a simplifying choice, not an absolute filmmaking rule. Once the action reads clearly, you can experiment with another angle and judge whether the viewer can still follow it.
Use the card as a review document, not a magic incantation. It belongs beside your reference images and shot list. A model may still alter an object despite a precise instruction. The value is that you can identify the alteration and choose whether to regenerate, edit, replace, or abandon that shot.
Fixed identity: 4 yellow pencils; 1 white handleless cup; 1 closed red notebook. Fixed setting: pale gray tabletop; light from screen left. Fixed orientation: cup on screen right; notebook starts at the back. Allowed changes: pencils move into cup; notebook slides into cleared center. Hand, where visible: right hand, plain gray sleeve, enters from screen bottom. Reject: changed object count, new cup handle, open notebook, reversed layout.
Build a 24-second edit before generating a single clip
I would divide this sequence by what the audience needs to understand, rather than giving every shot the same duration. The first two shots establish the arrangement and the empty destination. The middle two show the pencil action. The fifth moves the notebook, and the last lets the viewer register the result.
The example timings total 24 seconds: three plus three plus five plus five plus four plus four. These are intended durations on the editing timeline, not promises about a generator's available clip lengths. Select a supported generation duration, then trim suitable material to fit. If the movement cannot fit naturally, change the edit rather than accelerating it until it looks implausible.
The difficult join is between shots three and four. Shot three must end with the pencils gathered in the hand, and shot four must start with those pencils still held. If the first ends with the pencils already inside the cup, the next shot repeats the action. If it ends with only two gathered, the next shot cannot honestly begin with four in the hand.
Treat the opening and closing overview as a pair. Their framing should make the change easy to compare. The final shot is not a new desk, a new cup, or a prettier replacement scene. It is the same arrangement after the planned actions.
The six-shot structure is only a worked example. A three-shot version might tell the story better if the intermediate actions are distracting or unreliable. The purpose of a shot list is to expose decisions before generation, not to create an obligation to keep every planned shot.
01 | 0-3s | Overview: four loose pencils, empty cup, notebook at back. End unchanged. 02 | 3-6s | Cup detail: empty interior visible. Pencils remain on the desk outside the crop. 03 | 6-11s | Hand gathers four pencils. End: pencils held together above the tabletop. 04 | 11-16s | Hand lowers those pencils into cup and releases. End: four pencils in cup. 05 | 16-20s | Hand slides closed notebook from back to cleared center. Cup stays on right. 06 | 20-24s | Matching overview of finished desk. Hold the arrangement steady.
Approve the still frames before asking them to move
For a sequence built around recognizable objects, I would first assemble a small set of approved stills: the opening overview, the closer pencil view, the cup interaction, and the finished overview. They might be photographs you own, permitted reference assets, or generated concepts you have reviewed. They should all agree with the continuity card.
Runway documents an image-reference feature for creating new images from supplied visual references. That is a way to develop related stills; it is not itself a guarantee that separate video clips will match. Its image-to-video guide separately explains that the starting image supplies visual information while the text guides movement. These are two different stages of the process.
Before proceeding, count the pencils in each image and check the cup's shape. Look for a notebook that quietly acquired a spiral binding or a tabletop that changed texture. Correct those differences while you are still choosing images. Animating an unsuitable starting frame does not remove the underlying mismatch.
Do not ask a reference image to reveal information it does not contain. An overhead view may hide the side of the cup. A close crop may omit the notebook entirely. If the next shot needs a different view, prepare and inspect that view instead of assuming the tool will reconstruct every hidden detail correctly.
This is also the point to consider a simpler production method. If you can photograph or record the actual desk in a few minutes, that may be a better starting point than regenerating multiple almost-matching versions. AI is an option within the production, not a requirement for every frame.
Sources: Runway: creating images with Gen-4 References; Runway: image-to-video prompting guide
Give the subject and the camera separate jobs
A useful motion instruction answers two different questions: what moves inside the scene, and what happens to the view? For shot five, the hand slides the notebook toward the center while the camera stays in place. That distinction is clearer than asking for a dynamic desk transformation.
Google DeepMind's Veo guide discusses framing, camera movement, lighting, action, and other scene details as distinct prompt elements. I would use those categories to find missing decisions, not to add every possible effect. In this example, keeping the view simple makes the object movement easier to judge.
Keep a complete scene description for text-to-video work. For an image-to-video shot, start from the approved image and make the requested change explicit. You do not need to reintroduce a different desk design merely because the next shot has a new prompt. Match the wording to the input mode you are actually using.
The sample below describes shot five only. It does not request a montage, an opening title, an ending card, and several camera moves at once. Those decisions belong elsewhere in the project. If an output fails, a small instruction gives you fewer competing causes to investigate.
A prompt is still a request, not a mechanical constraint. Review whether the hand touches the notebook naturally, whether the notebook stays closed, and whether the pencils remain in the cup. The most relevant question is not whether the clip looks cinematic. It is whether it performs the planned action while preserving the rest of the scene.
Prompt to try
Animate the supplied approved desk image as one continuous shot. The right hand in the gray sleeve slides the closed red notebook from the back of the desk toward the clear center, then withdraws toward the bottom of the frame. The camera stays fixed. The white cup on screen right remains in place with all four yellow pencils inside it. Lighting stays consistent with the starting image. End with the notebook settled in the center.
Sources: Google DeepMind: Veo prompt elements
Check the join, not just the best frame in each clip
Put the end of shot three beside the beginning of shot four. Ignore the attractive background and inspect the hand, pencils, and cup. Are four pencils held together in both images? Is the hand approaching the same side of the cup? Does the light come from the same direction? This comparison is more useful than judging each clip's thumbnail.
For a continuation from the same viewpoint, Runway documents using a completed clip's final frame as the image input for another generation. That can provide a shared starting point. It does not prove the subsequent motion will remain correct, and a deliberately different camera angle needs a separately checked composition.
There is another catch: a bad ending is a bad reference. If the last frame has a bent cup or an extra pencil, carrying it forward preserves the mistake. Select only a suitable frame that also represents the correct moment in the action. Moving back a few frames changes the handoff state, so update the next shot's plan accordingly.
Try the join at ordinary playback speed as well as frame by frame. A small shift may be invisible during a purposeful cut, while a frozen pause can make otherwise matching shots feel disconnected. The editing decision depends on what the audience can follow, not just whether two still images look similar.
Do not hide a missing action behind an unrelated cutaway. An empty cup followed by a full cup can be a deliberate time jump, but it is no longer a demonstration of placing the pencils. Either show enough of the action or make the jump explicit in the structure. Keep the promise of the video aligned with what it actually shows.
Rewrite the failed instruction, not the entire project
Suppose shot four changes the cup into a mug. That is an object-identity failure, not a reason to add more adjectives about realism. First inspect its input image. If the handle is already present, replace the frame. If the frame is correct, simplify the action or revise the instruction while keeping other settings unchanged where practical.
Now suppose the cup stays correct but the hand deposits only three pencils. The problem is different. You need to inspect the transfer, not recolor the scene. A closer view may make the action easier to evaluate, but it can also expose more detail that must remain correct. Choose the revision because it addresses the observed failure.
I would keep a short attempt log with the shot number, input frame, requested change, observed issue, and decision. This is not a scientific benchmark: a single successful retry does not prove a prompt is universally better. It is a record that helps you avoid repeating unsuccessful changes without learning anything.
Set a stopping rule before a difficult action consumes the whole project. For this exercise, you could allow two targeted revisions after the first attempt, then switch to a recorded hand movement or simplify the sequence. That limit is a planning choice, not a statement about expected generation success or cost.
Preserve the original versions until the edit is settled. The newest generation may fix the hand but spoil the lighting. Comparing named versions is easier than trying to remember which preview had the correct cup. Keep a selected version for each shot and a short explanation of why it was chosen.
04-A | Cup gained a handle | Input frame also has handle | Replace input; reject clip. 04-B | Cup correct; only 3 pencils transferred | Input shows 4 | Simplify transfer instruction; inspect action. 04-C | Transfer still unclear | Revision limit reached | Record the action or revise the edit. Decision: do not label any version approved until the handoff and object count pass review.
Make the silent rough cut work before adding polish
Place the selected shots in order without music, captions, or transitions. Can someone explain what changed? If the answer depends on a voiceover describing an action that is not visible, decide whether the picture sequence needs repair or whether the video's purpose has changed.
Trim around complete, readable actions. A generated clip may begin with a pause or end with an unwanted movement. You do not have to use its full duration. However, trimming cannot create a transfer that never occurred. It can remove a distraction, but it should not conceal a missing step while the narration claims that step was demonstrated.
Use the same opening and closing framing where possible, then compare the object arrangement. The notebook's move should explain the cleared space. The cup should still be on the same side of the desk. The final shot needs enough time for a viewer to understand the result; the four seconds in this plan are a starting choice to assess during editing.
Add the final caption as editable text in the editor. For this exercise, Clear a little space to begin is sufficient. There is no reason to generate a changing price, a product label, or a supposed customer quote inside the footage. Keep important wording separate so it can be checked and revised without regenerating the scene.
If you add narration or sound, listen across the cuts. A sudden background change can make connected visuals feel like unrelated recordings. Use audio you are permitted to use, and avoid implying that an illustrative generated scene is a recording of an actual event. The caption, narration, and images should all describe the same modest idea.
Judge the finished sequence at the size people will watch
A desktop editing preview can make a small detail look obvious. On a phone, the viewer may not distinguish the four pencils from the background or see whether the notebook moved. Watch the exported video at its intended display size, with the surrounding interface in mind, before deciding the explanation is clear.
I would review this example three times for different purposes. First, watch without sound and describe the sequence. Second, focus on continuity: cup shape, pencil count, notebook state, hand position, and light direction. Third, watch with the intended audio and text to check that they reinforce rather than contradict the pictures.
Separate must-fix problems from optional polish. A pencil appearing before it is transferred is a must-fix problem for this demonstration. A less dramatic camera angle is not. A softer background might be acceptable if the action is clear. A perfect-looking final desk is not acceptable if it replaces the objects from the opening.
Ask a reviewer a neutral question: what happened in this video? Avoid explaining the intended story first, because that gives away the information you are checking. If they say the notebook appears from nowhere, investigate that join even if the clip itself is visually impressive.
Finally, inspect the exported file, not only the editor preview. Check the full sequence, text placement, sound, and ending. Save the approved export together with the shot plan and selected inputs. If someone later asks for a shorter version, those records make it possible to remove a shot deliberately instead of breaking the story by accident.
Use the method on another idea without copying the same six shots
The reusable part is the relationship between states, not the desk props or the number of cuts. For a packing tutorial, the states might be an empty bag, selected items beside it, and those items packed inside. For a simple explanation of a digital feature, real screen recordings may be more appropriate than generated interfaces.
Start each new project by identifying what the viewer needs to recognize across shots. It might be a person, an object, a location, or a stage of a process. Then identify which details can change. A story about redecorating a room should allow furniture placement to change while keeping the room's identity understandable.
The planning prompt below asks for those decisions before any clip prompts. It is meant for a text assistant. It should return a production plan for you to inspect, not a claim that the requested footage is possible in your chosen tool. Replace the placeholders and add the real constraints of your project.
If the assistant proposes several actions inside one short shot, ask which action is essential. If it assumes a reference feature your tool does not offer, revise the method. If it invents an attractive result that the source material cannot support, remove it. The plan should become more faithful to the job as you revise it, not simply more elaborate.
My completion rule would be straightforward: every retained shot communicates its intended moment, and every join preserves or deliberately explains what changes. That gives you a coherent video to evaluate. More adjectives, more generations, and more shots are useful only when they help meet that rule.
Prompt to try
Help plan a short video about [visible change] for [audience]. Target edit length: [seconds]. Available material: [owned or permitted reference images, recordings, and tools]. First separate fixed visual details from allowed changes. Then propose only the shots needed to explain the change. For each shot, give its purpose, approximate edit duration, starting state, one main action, ending state, camera position, and continuity check against the next shot. Make the durations add up. Flag difficult interactions and propose a recorded or simplified alternative. Do not invent test results, tool capabilities, or product claims. Wait for my approval before writing individual generation prompts.