Rewrite the inputs below into a structured MiniMax H3 image-to-video / reference-to-video prompt.

The user idea may come from ANY source — another video model's prompt, free-form natural language, a loose idea, or a structured non-H3 format. Treat it as raw material: extract the motion/change intent and preserve concrete details, then restructure into the H3 format below. Translate loose camera phrasing into the official motion-type + amplitude + speed vocabulary. Infer concrete diegetic sound and non-diegetic music where the input is silent. CRITICAL for image tasks: do NOT re-describe appearances already visible in the reference image — the alignment directive / subject_definitions already declare what is referenced; describe only the CHANGE. Keep dialogue, lyrics, and visible on-screen text in their original language.

{structure_guidance}

If an alignment directive is given below, it MUST be the first line of the final prompt, followed by one blank line before the first section. Keep it verbatim.

Alignment directive (if non-empty):
{align_directive}

User idea (raw, any format — the motion / change to depict):
{idea}

Reference-material caption (features that MUST be preserved — lock them item by item; describe only the CHANGE on top of them):
{caption}

Specs:
- duration: {duration} seconds
- aspect ratio: {aspect_ratio}
- category advice (apply as default styling where relevant; ignore if empty): {category_advice}

Output language: {output_language}
(`en` = write the prompt in English; `zh` = write the prompt in Chinese. Dialogue, lyrics, and visible on-screen text ALWAYS stay in their original language regardless.)

Use EXACTLY these section headers, in this order, one per line. Copy them VERBATIM (they are pre-localized to the output language and carry the canonical English API field name in parentheses — do not translate, rename, or duplicate them):

{section_headers}

If (and only if) there are banned elements, append a final section using this header line:
{negatives_header}

Few-shot examples of the desired prompt style:
{examples}

Now output ONLY the final structured H3 prompt following the system-prompt rules: alignment directive (if any) verbatim at the top, then the section headers above in order. For image-bearing base modes describe only the change and do NOT re-describe the reference image's visible appearance; lock identity by naming features item by item. No preamble, no explanation.
