h3-prompt-writing
minimax-ai/minimax-h3
Write MiniMax H3 video generation prompts for T2VA, I2VA, FL2VA, L2VA, and Ref2VA modes.
What is h3-prompt-writing?
Converts multimodal requests into structured H3 prompt formats for MiniMax video generation. Use this skill when composing integrated descriptions, soundscapes, music tracks, and reference labels for text-to-video, image-to-video, frame-list-to-video, last-frame-to-video, and full-reference video workflows.
- Rewrite user requests into H3 prompt structures with exact field names and section order
- Compose integrated_multimodal_description, overall_soundscape, and non_diegetic_music sections
- Align keyframes and define reference labels consistently across all sections
- Support five input modes: T2VA (text), I2VA (first frame), FL2VA (frame path), L2VA (last frame), and Ref2VA (full reference)
- Preserve dialogue, lyrics, and visible text in original language while describing shots by composition, subjects, environment, actions, camera, and sound
How to install h3-prompt-writing
npx skills add https://github.com/minimax-ai/minimax-h3 --skill h3-prompt-writing- Access to references/base-en.txt and references/ref-en.txt guide files
- Understanding of the five H3 input modes (T2VA, I2VA, FL2VA, L2VA, Ref2VA)
How to use h3-prompt-writing
- 1.Identify the input mode from the user's request (text, first frame, frame list, last frame, or full reference)
- 2.Read the appropriate reference guide (base-en.txt for T2VA/I2VA/FL2VA/L2VA; ref-en.txt for Ref2VA)
- 3.Rewrite the request into the required sections in the exact order specified
- 4.Preserve reference labels (e.g., <Picture 1>, <Video 1>, <Audio 1>) consistently across all sections
- 5.Verify total duration matches the requested video length and timing notation is accurate
Use cases
- Generate a 6-second video prompt from a text description with synchronized audio and music
- Convert a reference image into a video by describing the opening frame and inferring the timeline forward
- Create a video that flows between a supplied first and last keyframe with continuous motion description
- Rewrite a multi-reference video request (images, videos, audio) into a six-section Ref2VA structure with consistent labels
- Align timing notation and duration to match requested video length (4–15 seconds)
- Video generation engineers working with MiniMax H3 API
- Prompt engineers optimizing multimodal video requests
- AI agents handling video synthesis workflows
- Developers building video generation pipelines with reference content
h3-prompt-writing FAQ
T2VA builds the full audiovisual timeline from text alone. Ref2VA rewrites use supplied reference images, videos, or audio and requires six sections (subject_definitions, summary, retention_analysis, detailed_description, overall_soundscape, non_diegetic_music) with consistent reference labels.
I2VA starts from the supplied first frame and develops forward. FL2VA describes the continuous path between first and last frames. L2VA infers a plausible opening and converges to the supplied last frame. Always state how the keyframe(s) connect to the timeline.
Write rewrite sections in English, but preserve dialogue, lyrics, and visible scene text in their original language.
Match the total duration of the description to the requested video length (4–15 seconds) and use exact timing notation from the reference guides.
Use concrete visual and audio details (composition, subjects, environment, actions, camera, sound) instead of abstract words like 'cinematic' or 'beautiful'. Keep reference labels consistent and avoid unresolved labels.
Full instructions (SKILL.md)
Source of truth, from minimax-ai/minimax-h3.
name: h3-prompt-writing description: Write MiniMax H3 video generation prompts for T2VA, I2VA, FL2VA, L2VA, and Ref2VA. Use when rewriting multimodal requests into H3 prompt structures, composing integrated_multimodal_description, overall_soundscape, and non_diegetic_music, aligning keyframes, or defining reference labels for images, videos, and audio. compatibility: Portable to any agent that can read local files — no external API calls, MiniMax Hub tools, or proprietary runtime required. The agents/openai.yaml file only adds optional ChatGPT/Codex UI metadata; it does not restrict the skill to OpenAI agents.
H3 Prompt Writing
Workflow
- Identify the input mode: T2VA, I2VA, FL2VA, L2VA, or full-reference Ref2VA.
- For base text/keyframe modes, read
references/base-en.txtand follow its final prompt structure. - For full-reference mode, read
references/ref-en.txtand follow its six-section rewrite format. - Preserve the exact field names, section order, labels, and timing notation from the selected guide.
Base Modes
- T2VA: build the full audiovisual timeline from text.
- I2VA: start from the first frame and develop forward from it.
- FL2VA: describe the continuous path between the first and last frames.
- L2VA: infer a plausible opening and converge to the supplied last frame.
Use integrated_multimodal_description, overall_soundscape, and non_diegetic_music in the order shown in references/base-en.txt.
Full-Reference Mode
Ref2VA rewrites use subject_definitions, summary, retention_analysis, detailed_description, overall_soundscape, and non_diegetic_music in that order. Reference labels stay consistent across all sections.
Read references/ref-en.txt for label rules, retention analysis, and complete examples.
Output Rules
- Write rewrite sections in English; preserve dialogue, lyrics, and visible scene text in their original language.
- Describe each shot by composition, subjects, environment, actions, camera, sound, and the exact point where referenced content appears.
- Avoid plot summaries, unresolved reference labels, and timing that does not match the requested duration.
Tips for Better Results
- Always match the total duration of the description to the requested video length (4–15 seconds).
- Keep reference labels consistent (e.g.
<Picture 1>,<Video 1>,<Audio 1>) across every section. - Prefer concrete visual and audio details over abstract words like "cinematic" or "beautiful".
- When using keyframes (I2VA / FL2VA / L2VA), clearly state how the first and/or last frame connects to the timeline.
Related skills
More from minimax-ai/minimax-h3 and the wider catalog.

handdrawn-live-video-generator
Create 15-second surreal hand-drawn animation blended with live-action video using MiniMax H3.

minimalist-product-ad-generator
Generate minimalist product ad videos for e-commerce from images and brief requirements.

music-video-subtitle-generator
Design beat-reactive music video prompts with lyric typography and multi-shot stitching for MiniMax Hub.

paper-collage-explainer-generator
Turn narration, story beats, or concepts into premium halftone paper-collage stop-motion animations with tactile sound effects.

papercraft-stop-motion-explainer
Create production-ready papercraft stop-motion explainers with character design, diorama sets, storyboards, and video prompts.

android-native-dev
Android native development guide covering Material Design 3, Kotlin/Compose, project setup, and build configuration.