Multimodal references for a more directed video brief

Reference to Video AI Generator

Use the Reference to Video AI Generator to combine images, video clips, and audio in one production brief. Choose a compatible model, then use @ mentions to tell it which reference should guide the subject, motion, camera, or sound.

Reference types and upload limits vary by the selected model.

Images, video clips, and audio arranged as references for an AI video brief

Give Every Reference a Clear Role

Reference-to-video generation works best when each upload has one job. Decide what must remain recognizable, what should move, and what should set the pace before you generate.

Images anchor the look

Use a reference image for the subject, product, character, palette, or composition that should remain visually recognizable.

Clips describe movement

Point to a video reference when motion, camera behavior, pacing, or a transition is easier to show than explain.

Audio sets timing

Use an audio reference to communicate rhythm, ambience, or another sound cue supported by the selected model.

The prompt connects them

Mention uploaded files with @ and state how each reference should influence the generated clip.

From a Reference Set to a Testable Video Brief

Select the source material that matters, assign each reference a role, and review one generation at a time.

Start with the few assets that define the shot. A focused set makes it easier to write precise instructions and diagnose the result.

Reference-Led Video Production in Veo 3

Combine visual, motion, and audio references in one clear path from upload to finished video.

Mixed-media references

Bring images, video clips, and audio together to guide one new scene.

Explicit asset mentions

Use @image, @video, and @audio tokens to connect prompt directions to uploaded files.

Model-aware settings

Available durations, resolutions, reference types, and limits update with the selected model.

Frame-based alternative

When a shot needs fixed endpoints, choose first and last frames instead of multimodal references.

Visible task progress

Follow generation status and reopen completed results from My Creations.

Downloadable result

Preview the completed clip and download the returned video when it is ready.

Reference-to-Video AI Questions

What to know before generating a new video from mixed reference assets.

What is a reference-to-video AI generator?

A reference-to-video AI generator creates a new clip from a prompt plus source assets such as images, video clips, or audio. Each reference can guide a different part of the result, including identity, style, motion, camera language, rhythm, or sound. In Veo 3, you organize those inputs into one directed brief and review the output against the role assigned to every asset.


How is reference to video different from image to video?

Image to video usually begins with one image or a frame pair that anchors the shot. Reference to video can assign separate roles to multiple supported assets, including images, motion clips, and audio cues.


Can I upload images, videos, and audio together?

Yes. Add the images, video clips, and audio cues your scene needs, then give each upload a distinct purpose. A focused set is easier to direct than a large undifferentiated mood board. The upload panel shows accepted types and limits as you prepare the brief, and the available combination depends on the model and reference mode you choose.


How do I tell the generator which file to use?

Upload the assets, type @ in the prompt, and select the matching reference token. Then state the role of that asset, such as subject from @image1, camera movement from @video1, or rhythm from @audio1. Place the instruction close to the matching token and avoid assigning two conflicting references to the same decision unless you explain how they should be combined.


When should I use first and last frames instead?

Choose first and last frames when the shot needs fixed visual endpoints. Choose multimodal references when separate images, motion clips, or audio cues should influence the scene.


How do I choose the right model for my references?

Start with the assets that matter most, then choose a model that supports those reference types and the duration, resolution, and sound controls your project needs. Review the available upload panel after selecting the model, remove references that have no clear job, and check the credit estimate before generating. The best choice is the workflow that can express your brief without unnecessary inputs.


How should I prepare a mixed-media video brief?

Write the intended outcome in one sentence, then list the references and the decision each one controls. Put the visible action in chronological order, describe the camera path, and connect audio cues to specific moments. Finish with the ending state or edit point. If the brief cannot be summarized clearly, split it into separate shots or remove references that do not change the result.


What is the difference between a style reference and a subject reference?

A subject reference defines what must remain recognizable, such as a person, product, character, or object. A style reference guides palette, lighting, texture, composition, or visual treatment. State the distinction in the prompt so an artistic cue does not accidentally replace an identity cue. When one image serves both roles, name the specific attributes to preserve and the qualities that may change.


How can a video clip guide motion without copying the whole scene?

Mention the clip and describe the motion feature you want to borrow: camera speed, subject gesture, transition timing, blocking, or rhythm. Separately define the new subject and environment. This tells the generator that the reference demonstrates movement rather than supplies every visual detail. Review whether the resulting camera path and action timing follow that intended role before refining surface style.


How do I use audio references for timing and atmosphere?

Assign the audio a clear function, such as dialogue cadence, musical rhythm, room ambience, or an effect synchronized to a visible action. Describe when the cue enters and what happens on screen at that moment. Review the completed clip with sound on and off: first confirm that the visuals tell the scene, then check timing, clarity, balance, and whether the audio supports the intended pace.


What should I review after a reference-led generation?

Use the original brief as a scorecard. Check whether each source performed its assigned role, then inspect subject continuity, motion logic, camera movement, background stability, small details, in-frame text, and audio timing. Separate must-fix issues from optional polish, keep the strongest result in My Creations, and revise one high-impact instruction or reference before generating the next version.


How can a team keep reference-to-video revisions organized?

Save the prompt, model, output settings, uploaded asset names, and a short role note for every reference. Label each result by the one change being tested and collect feedback against the same criteria. This creates a traceable production brief, helps collaborators compare versions fairly, and prevents useful identity, motion, or sound decisions from being lost when the next generation is prepared.


Turn Your References into One Directed Shot

Open the generator, keep the references that matter, and tell the model exactly how each one should shape the new scene.