FastVideo-Minimax-FastH3-Preview-v0.2

A few-step (4-step) distillation preview of MiniMax-H3, the 33B dual-modality (video + audio) diffusion transformer β€” distilled with data-free DMD2 by the FastVideo team.

The base model samples with 50 denoising steps; this student walks a 4-step grid on the release's shift-12 rectified-flow schedule (12.5Γ— fewer transformer evaluations), generating synchronized video and audio in one pipeline call.

Preview status (v0.2): step 2900 of a 4000-step run β€” the same run that produced v0.1 (step 1400), carried 1500 steps further. Sample quality is still maturing, most visibly on high-motion detail.

What's new since v0.1

  • 1500 more distillation steps (1400 β†’ 2900) on the same data-free DMD2 run: sharper still detail and steadier audio/video sync.
  • Corrected sampling contract in the card. v0.1 documented num_inference_steps=4, which makes the scheduler build its own 4-point sigma grid β€” 3 forwards on native spacing, not the 4 trained jump points. Sample with the explicit trained ladder instead (see Usage). The same off-operating-point mismatch affected this run's in-training validation renders, so judge the student by fresh samples on the ladder below, not by earlier validation clips.
  • Fixed the repo id in the usage snippet (v0.1's card had it doubled).

What's in the repo

Diffusers-format (modular pipeline) layout. Only the transformer/ weights differ from the base release β€” the distilled student, in bf16. All other components (Qwen3-VL text encoder, video/audio VAEs, tokenizer, processor, schedulers) are unmodified copies of the base release, included so the repo is self-contained.

The student was trained with block-sparse video attention (VSA, 64-token tiles, 90% sparsity) and carries its trained sparse-gate parameters (attn.to_gate_compress); it can be run dense (default) or with VSA for additional inference speedup.

Usage (FastVideo)

Sample on the trained ladder β€” [999, 749, 500, 250] on the shared 1000-step grid, one forward per entry, each scheduler applying its own shift:

from fastvideo import VideoGenerator

gen = VideoGenerator.from_pretrained(
    "FastVideo/FastVideo-Minimax-FastH3-Preview-v0.2",
    num_gpus=1,
    dmd_denoising_steps=[999, 749, 500, 250],
)
video = gen.generate_video(
    prompt="<your H3-format multimodal prompt>",
    guidance_scale=1.0,   # the base model is guidance-distilled
)

The ladder can also be set without touching code:

export FASTVIDEO_DMD_DENOISING_STEPS=999,749,500,250

To run the student under the sparse attention it was trained with, select the VSA-H3 backend and match both knobs of the training contract β€” sparsity alone, at the default 256-token tile, is a different operating point:

export FASTVIDEO_ATTENTION_BACKEND=VIDEO_SPARSE_ATTN_H3
# generator args: VSA_sparsity=0.9, VSA_tile_size=64

Prompts follow the MiniMax-H3 multimodal prompt format (integrated_multimodal_description: ... overall_soundscape: ...); see the base model card for the prompting guide.

Training summary

  • Method: data-free DMD2 (distribution matching distillation) β€” student / frozen teacher / trained fake-score critic, backward-simulation rollout (the student walks its own 4-step sampling grid during training), x0-space critic regression, shifted score-time sampling matched to the dual video/audio noise clocks (shifts 12 / 3).
  • Student grid: [999, 749, 500, 250] β€” 4 steps on the release sampler's shift-12 schedule.
  • Attention: student trained with VSA block-sparse attention (64-token tiles, 90% video-tile sparsity); teacher and critic dense.
  • Data: text prompts only (data-free) β€” ~258k prompts (VidProM-H3 + synthetic t2va prompt set); no video data used.
  • Resolution: 768Γ—1344, 124 frames (5s) with synchronized audio.
  • Optimization: global batch 64, lr 1e-6 (student and critic), fp32 master weights, bf16 compute, 2900 of 4000 steps.
  • Hardware: 32Γ— NVIDIA GB200.

Limitations

  • Preview checkpoint β€” quality below the base model's 50-step sampling, especially on fine motion and audio detail; improves with training.
  • Inherits all content limitations and usage restrictions of the base model.
  • The 4-step ladder is what the student was trained for; other step counts and other timestep grids are off-distribution.

License

Distributed under the MiniMax H3 Community License (see LICENSE), inherited from the base model. Review the license (including its territory and acceptable-use terms) before use or redistribution.

Notes

  • The transformer_ref component (reference-conditioning variant) is not packaged here; its entry in modular_model_index.json points at the base MiniMaxAI/MiniMax-H3 repo and is fetched from there if used. This preview distills the text-to-video+audio path only.
Downloads last month
-
Safetensors
Model size
35B params
Tensor type
BF16
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for FastVideo/FastVideo-Minimax-FastH3-Preview-v0.2

Finetuned
(82)
this model
Adapters
1 model

Space using FastVideo/FastVideo-Minimax-FastH3-Preview-v0.2 1