Rewrite the input below into a structured MiniMax H3 text-to-video (T2VA) prompt.

The input may come from ANY source — another video model's prompt, free-form natural language, a loose idea, or a structured non-H3 format. Treat it as raw material, NOT a final prompt: extract every visual, action, camera, lighting, style, and on-screen-text intent it contains, and preserve all concrete details. Where the input is silent on a dimension H3 requires (camera motion type/amplitude/speed, diegetic sound, non-diegetic music), infer a concrete, fitting value rather than copying vague wording. Translate loose camera phrasing (e.g. "镜头跟随", "slow motion", "快速穿梭") into the official motion-type + amplitude + speed vocabulary. Keep dialogue, lyrics, and visible on-screen text in their original language.

Input (raw, any format):
{idea}

Specs:
- duration: {duration} seconds
- aspect ratio: {aspect_ratio}
- category advice (apply as default styling where relevant; ignore if empty): {category_advice}

Output language: {output_language}
(`en` = write the prompt in English; `zh` = write the prompt in Chinese. Dialogue, lyrics, and visible on-screen text ALWAYS stay in their original language regardless.)

Structure: {structure_guidance}

Use EXACTLY these section headers, in this order, one per line. Copy them VERBATIM (they are pre-localized to the output language and carry the canonical English API field name in parentheses — do not translate, rename, or duplicate them):

{section_headers}

If (and only if) there are banned elements, append a final section using this header line:
{negatives_header}

Few-shot examples of the desired prompt style:
{examples}

Now output ONLY the final structured H3 T2VA prompt following the system-prompt rules. No preamble, no explanation.
