Text-to-Video
Diffusers
Safetensors
MiniMax H3
video
audio
text-to-audio-video
distillation
dmd2
few-step
fastvideo
fasth3
Instructions to use FastVideo/FastVideo-FastH3-8-Step-V2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use FastVideo/FastVideo-FastH3-8-Step-V2 with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("FastVideo/FastVideo-FastH3-8-Step-V2", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
Model card: match the 4-step Preview card layout; add acknowledgements
Browse filesDocumentation only. README.md rewritten to the FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree card structure (logo, summary, run, scope, acknowledgements). No weights or provenance files changed.
README.md
CHANGED
|
@@ -12,9 +12,7 @@ tags:
|
|
| 12 |
- text-to-audio-video
|
| 13 |
- distillation
|
| 14 |
- dmd2
|
| 15 |
-
- vsa
|
| 16 |
- few-step
|
| 17 |
-
- 8-step
|
| 18 |
- minimax-h3
|
| 19 |
- fastvideo
|
| 20 |
- fasth3
|
|
@@ -26,44 +24,25 @@ tags:
|
|
| 26 |
|
| 27 |
# FastVideo-FastH3-8-Step-V2
|
| 28 |
|
| 29 |
-
FastH3 8-Step V2
|
| 30 |
-
[FastVideo](https://github.com/hao-ai-lab/FastVideo)
|
| 31 |
-
|
| 32 |
-
|
| 33 |
-
Video Sparse Attention (VSA).
|
| 34 |
|
| 35 |
-
[
|
| 36 |
-
[FastH3 blog](https://haoailab.com/blogs/fasth3-preview/) ·
|
| 37 |
[FastH3 collection](https://huggingface.co/collections/FastVideo/fastvideo-fasth3)
|
| 38 |
|
| 39 |
-
|
| 40 |
-
|
| 41 |
-
|
| 42 |
-
## Model
|
| 43 |
-
|
| 44 |
-
| Setting | Value |
|
| 45 |
-
|---|---|
|
| 46 |
-
| Task | Text to video with stereo audio (T2AV) |
|
| 47 |
-
| Method | Data-free [DMD2](https://arxiv.org/abs/2405.14867), carried backward simulation |
|
| 48 |
-
| Student attention | [VSA-H3](https://arxiv.org/abs/2505.13389), 80% sparsity, 64-token tiles |
|
| 49 |
-
| Denoising | 8 transformer forwards |
|
| 50 |
-
| Video / audio shifts | 10 / 3 |
|
| 51 |
-
| Guidance scale | 1.0 |
|
| 52 |
-
| Checkpoint | Training step 1300, bf16 full weights |
|
| 53 |
-
|
| 54 |
-
Training used prompt conditioning and student-generated latents, not target
|
| 55 |
-
video latents. The corpus used mixed resolutions and aspect ratios. This is
|
| 56 |
-
an eight-step DMD2 checkpoint, not a PDD model.
|
| 57 |
-
|
| 58 |
-
**This checkpoint requires FastVideo's VSA-H3 backend and Video Sparse
|
| 59 |
-
Attention kernel.** The four-step models' default 90% sparsity and 12/3 shifts
|
| 60 |
-
are not this model's sampling recipe. No matching eight-step LoRA is included.
|
| 61 |
|
| 62 |
## Run with FastVideo
|
| 63 |
|
| 64 |
-
|
| 65 |
-
|
| 66 |
-
|
|
|
|
|
|
|
| 67 |
|
| 68 |
```bash
|
| 69 |
git clone https://github.com/hao-ai-lab/FastVideo.git
|
|
@@ -75,58 +54,24 @@ UV_TORCH_BACKEND=cu130 uv pip install \
|
|
| 75 |
-e ".[fasth3]"
|
| 76 |
```
|
| 77 |
|
| 78 |
-
Inference support for this checkpoint (video shift 10 and the trained
|
| 79 |
-
eight-rung ladder) is in FastVideo
|
| 80 |
-
[PR #1852](https://github.com/hao-ai-lab/FastVideo/pull/1852). Until it is
|
| 81 |
-
merged, check out that branch; afterwards, `main` works unchanged.
|
| 82 |
-
|
| 83 |
```bash
|
| 84 |
python examples/inference/basic/basic_fasth3_8step.py \
|
| 85 |
-
--prompt
|
| 86 |
-
--
|
| 87 |
-
--
|
| 88 |
-
--height 768 --width 1344 --num-frames 124 \
|
| 89 |
-
--output outputs/fasth3-8step
|
| 90 |
```
|
| 91 |
|
| 92 |
-
|
| 93 |
-
|
| 94 |
-
|
| 95 |
-
|
| 96 |
-
preview example's CLI, so every other flag works unchanged; pass
|
| 97 |
-
`--model-path` to use a local snapshot.
|
| 98 |
-
|
| 99 |
-
To download ahead of time:
|
| 100 |
|
| 101 |
-
|
| 102 |
-
hf download FastVideo/FastVideo-FastH3-8-Step-V2 \
|
| 103 |
-
--local-dir ./FastH3-8-Step-V2
|
| 104 |
-
```
|
| 105 |
|
| 106 |
-
|
| 107 |
-
|
| 108 |
-
|
| 109 |
-
|
| 110 |
-
|
| 111 |
-
## Scope and limitations
|
| 112 |
-
|
| 113 |
-
- T2AV only. FL2VA and Ref2VA were not distilled or validated for this release.
|
| 114 |
-
- Difficult motion, small details and some audio can remain below Base H3.
|
| 115 |
-
- Sampling outside the trained schedule or attention policy is an ablation.
|
| 116 |
-
- The four-step blog's latency and evaluation numbers do not establish this
|
| 117 |
-
eight-step model's performance. On 4x GB200 (SP4, eager strict profile,
|
| 118 |
-
VSA 0.8 tile 64, 832x480, 124 frames + 32 kHz audio) one measured request
|
| 119 |
-
took 6.3 s end to end including saving (FastVideo PR #1852 smoke). Quality
|
| 120 |
-
has not been formally compared against Base H3 or the four-step preview.
|
| 121 |
-
|
| 122 |
-
This model inherits the [MiniMax H3 Community License](LICENSE), including
|
| 123 |
-
its territorial, use and commercial restrictions. Review it before use or
|
| 124 |
-
redistribution. The bundled Qwen3-VL encoder is covered by
|
| 125 |
-
[Apache 2.0](LICENSE-Qwen3-VL). See [NOTICE](NOTICE) for attribution and modifications.
|
| 126 |
-
|
| 127 |
-
Training and checkpoint provenance are preserved in
|
| 128 |
-
[provenance.json](provenance.json), [checkpoint_metadata.json](checkpoint_metadata.json)
|
| 129 |
-
and [checkpoint_content.json](checkpoint_content.json).
|
| 130 |
|
| 131 |
## Acknowledgements
|
| 132 |
|
|
|
|
| 12 |
- text-to-audio-video
|
| 13 |
- distillation
|
| 14 |
- dmd2
|
|
|
|
| 15 |
- few-step
|
|
|
|
| 16 |
- minimax-h3
|
| 17 |
- fastvideo
|
| 18 |
- fasth3
|
|
|
|
| 24 |
|
| 25 |
# FastVideo-FastH3-8-Step-V2
|
| 26 |
|
| 27 |
+
The FastH3 8-Step V2 checkpoint from
|
| 28 |
+
[FastVideo](https://github.com/hao-ai-lab/FastVideo). It generates synchronized
|
| 29 |
+
video and audio from text with eight transformer forwards. This step-1300 model
|
| 30 |
+
was trained with data-free DMD2 and VSA-H3 at 80% sparsity.
|
|
|
|
| 31 |
|
| 32 |
+
[Blog](https://haoailab.com/blogs/fasth3-preview/) ·
|
|
|
|
| 33 |
[FastH3 collection](https://huggingface.co/collections/FastVideo/fastvideo-fasth3)
|
| 34 |
|
| 35 |
+
> This checkpoint requires FastVideo's VSA-H3 attention backend. Its video
|
| 36 |
+
> scheduler shift is 10, not the base model's 12; use the example below, which
|
| 37 |
+
> reads the trained schedule from the checkpoint.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 38 |
|
| 39 |
## Run with FastVideo
|
| 40 |
|
| 41 |
+
Install [uv](https://docs.astral.sh/uv/getting-started/installation/), then use
|
| 42 |
+
the CUDA 13 / Blackwell path below. It selects FastVideo's published CUDA
|
| 43 |
+
kernel wheel instead of compiling the kernel locally. See the
|
| 44 |
+
[installation guide](https://hao-ai-lab.github.io/FastVideo/getting_started/installation/)
|
| 45 |
+
for other platforms.
|
| 46 |
|
| 47 |
```bash
|
| 48 |
git clone https://github.com/hao-ai-lab/FastVideo.git
|
|
|
|
| 54 |
-e ".[fasth3]"
|
| 55 |
```
|
| 56 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 57 |
```bash
|
| 58 |
python examples/inference/basic/basic_fasth3_8step.py \
|
| 59 |
+
--prompt "your prompt" \
|
| 60 |
+
--no-warmup \
|
| 61 |
+
--repeats 1
|
|
|
|
|
|
|
| 62 |
```
|
| 63 |
|
| 64 |
+
The tested defaults use four B200 GPUs and the trained eight-forward schedule.
|
| 65 |
+
On other multi-GPU CUDA systems, follow the installation guide and add
|
| 66 |
+
`--no-replicated-dit --vsa-kernel triton --no-fa4`. The GPU count must divide
|
| 67 |
+
H3's 56 attention heads.
|
|
|
|
|
|
|
|
|
|
|
|
|
| 68 |
|
| 69 |
+
## Scope
|
|
|
|
|
|
|
|
|
|
| 70 |
|
| 71 |
+
This checkpoint supports text-to-audio-video generation. FL2VA and Ref2VA were
|
| 72 |
+
not distilled. Difficult motion, fine detail, and some audio may remain below
|
| 73 |
+
the base MiniMax H3 model. This checkpoint inherits the
|
| 74 |
+
[MiniMax H3 Community License](LICENSE).
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 75 |
|
| 76 |
## Acknowledgements
|
| 77 |
|