CoreLesson 435 min

Reference Generation

By the end of this lesson you can
  • Assemble a valid reference-to-video request within every documented count, duration, size, and format limit
  • Apply the reference-mode prompt structure - Subject labels and subject_definitions - from the official reference prompt guide
  • Price a reference job before submitting it, including input-video billing

Frame control, from last lesson, pins pixels at a clip's endpoints. Reference generation solves the opposite problem: keeping identity stable while everything else changes. You hand the model material — images of a character, clips of a motion, audio of a voice — and prompt scenes that none of that material depicts. The documented combination is simple: text + any mix of reference images (role=reference_image), reference videos (role=reference_video), and reference audio (role=reference_audio). This is r2va, the third mode of the content array, and it is how a character survives shot two, shot five, and shot twelve of a real production.

Remember the wall: the moment any reference role appears, first_frame and last_frame are forbidden in the same request — mixing them is a 400. Reference mode never promises you a specific frame; it promises you a consistent subject.

Core — locked

Unlock “Reference Generation”

The 4-lesson Foundation module is free and always will be — you have already read the part most guides get wrong. The paid lessons go deeper, with labs, knowledge checks, and every technical claim verified against live documentation.

MiniMax H3 — Complete Course
The whole course, one purchase
$129one-time

Already bought it? Sign in with the email you used at checkout.