Quick answer
Use image-to-video when one authorized image should be the literal starting frame. Use reference-to-video when one or more authorized images should guide recurring elements such as a person, product, or location while the generated shot can begin elsewhere.
xAI documents up to seven reference images per generation, but reference-to-video is currently capped at 720p. It does not accept a starting image and reference images in the same request.
For API voice reference, plan around the documented preset voice_id workflow—not uploaded voice cloning—and verify eligibility. Current xAI documentation restricts that capability to trusted partners in the United States.
Choose the anchor before writing the prompt
| Need | Mode | Input | Current resolution boundary |
|---|---|---|---|
| Start from an exact first frame | Image-to-video | Prompt + one starting image | Up to 1080p |
| Preserve a character, product, or location across a new shot | Reference-to-video | Prompt + reference images | Up to 720p |
| Generate without visual source media | Text-to-video | Prompt only | Up to 1080p |
| Carry a permitted preset voice into a referenced scene | Reference-to-video | Prompt + permitted preset voice ID, optionally with reference images | Up to 720p |
Read the model record for duration, editing, extension, polling, and error boundaries. If you are still choosing a product surface, use the API versus SuperGrok guide.
Assign one job to each reference
xAI’s release describes each reference as locking one thing in place. Treat that as workflow guidance, not a mathematical guarantee.
A useful seven-slot budget is:
- primary character or subject identity;
- wardrobe or secondary character;
- hero product front or three-quarter view;
- product detail that must survive motion;
- location or set geometry;
- lighting, palette, or material reference;
- optional secondary scene or brand constraint.
You rarely need all seven. Start with the minimum set that isolates the requirement, then add a reference only when a named failure persists. Conflicting faces, camera angles, lighting, product states, or locations can make the conditioning ambiguous.
Prepare reference images
For each image:
- confirm copyright, likeness, trademark, location, and contractual permission;
- use a clear subject with minimal obstruction and sufficient detail;
- remove accidental personal data, logos, people, or confidential material;
- keep identity references consistent in age, styling, and facial detail;
- include product angles needed for the planned camera movement;
- separate the subject from the scene when you need independent control;
- record source, owner, consent, permitted use, retention, and deletion date.
Do not use reference generation as identity verification, forensic evidence, or permission to imitate a real person. Synthetic similarity can still be wrong, harmful, or deceptive.
Write a shot brief, not a wish list
A reference prompt should explain what changes and what must remain stable:
Keep the referenced presenter and studio layout. Medium shot, slow camera push, presenter turns toward the product on the desk, warm practical lighting, six seconds, no on-screen text.
Review five fields before submission:
- anchors: which character, product, or location references must persist;
- action: one primary motion with a clear start and end;
- camera: framing, movement, lens feel, and orientation;
- environment: light, atmosphere, time, and background behavior;
- avoidances: unauthorized marks, extra people, unstable text, or unwanted transformations.
Avoid packing multiple scene changes into a short clip. Generate separate shots and edit them together when continuity matters.
Voice reference requires a separate gate
The July 31 consumer announcement demonstrates a character image paired with a voice reference. The current API documentation is narrower: it describes reference_audios containing built-in voice_id values, a maximum of three voices per request, and prompt tokens such as <AUDIO_0>.
Before using voice:
- confirm that the API team is an eligible trusted U.S. partner;
- select only a provider-documented preset voice;
- map each voice index to the named speaker in the prompt;
- disclose synthetic speech when context or policy requires it;
- block impersonation, fraud, harassment, and misleading endorsement;
- keep a text-only or non-dialogue fallback.
Do not promise that the API accepts a user’s uploaded recording. Do not assume consumer voice access grants API voice access.
Evaluate consistency with evidence
Create a small shot grid that varies one dimension at a time:
- close, medium, and wide framing;
- static, pan, push, and subject motion;
- neutral, indoor, and outdoor lighting;
- front, profile, and partial occlusion;
- quiet, dialogue, and ambient-audio scenarios;
- one, two, and several reference images.
Score each output for identity, product geometry, location, action, camera, motion artifacts, audio, lip behavior, unwanted marks, policy fit, and reviewer acceptability. Record the prompt, ordered references, mode, model ID, duration, aspect ratio, resolution, output, and rejection reason.
An attractive single sample is not evidence of repeatable consistency. Use the acceptance rate across the planned shot set and include human review time in the cost.
Frequently asked questions
How many images can Grok Imagine Video 1.5 use as references?
xAI’s July 31, 2026 release documents up to seven reference images per generation. Use each authorized image for a clear element such as a character, product, or location rather than assuming that more references always improve the result.
What is the difference between a starting image and a reference image?
A starting image defines the first frame in image-to-video. Reference images guide recurring elements such as identity, product appearance, or location in reference-to-video. xAI does not allow a starting image and reference images in the same request.
Can reference-to-video generate 1080p?
No, not under xAI’s current developer documentation. Reference-to-video is capped at 720p, while text-to-video and image-to-video on grok-imagine-video-1.5 can support 1080p.
Can I upload a voice recording to the xAI API?
The public API documentation does not describe uploaded voice clips. It documents preset voice IDs, up to three voices per request, for trusted partners in the United States. Consumer voice-reference behavior and eligibility are separate.
Do reference images guarantee identity consistency?
No. References provide conditioning, not a guarantee. Test face, product, wardrobe, location, motion, text, audio, and continuity across a representative shot set, then reject or regenerate outputs that fail the brief.
What rights do I need for reference media?
Use images, likenesses, voices, trademarks, products, and locations you are authorized to process. Obtain appropriate consent, avoid deceptive impersonation, review current provider terms, and complete human rights and brand review before publishing.
Bottom line
Pick one request mode, give every reference one named job, accept the 720p reference cap, and treat identity or voice consistency as something to test—not a promise. Rights, consent, disclosure, and final review remain your responsibility.
Official sources
Source check: August 2, 2026. Verify live reference limits, voice access, region, model behavior, terms, and acceptable-use rules before production work.