Reference Generation
- →Assemble a valid reference-to-video request within every documented count, duration, size, and format limit
- →Apply the reference-mode prompt structure - Subject labels and subject_definitions - from the official reference prompt guide
- →Price a reference job before submitting it, including input-video billing
Frame control, from last lesson, pins pixels at a clip's endpoints. Reference generation solves the opposite problem: keeping identity stable while everything else changes. You hand the model material — images of a character, clips of a motion, audio of a voice — and prompt scenes that none of that material depicts. The documented combination is simple: text + any mix of reference images (role=reference_image), reference videos (role=reference_video), and reference audio (role=reference_audio). This is r2va, the third mode of the content array, and it is how a character survives shot two, shot five, and shot twelve of a real production.
Remember the wall: the moment any reference role appears, first_frame and last_frame are forbidden in the same request — mixing them is a 400. Reference mode never promises you a specific frame; it promises you a consistent subject.
Unlock “Reference Generation”
The 4-lesson Foundation module is free and always will be — you have already read the part most guides get wrong. The paid lessons go deeper, with labs, knowledge checks, and every technical claim verified against live documentation.
Already bought it? Sign in with the email you used at checkout.