wlsaidhi commited on
Commit
71a17d7
·
verified ·
1 Parent(s): 76878a0

Model card: match the 4-step Preview card layout; add acknowledgements

Browse files

Documentation only. README.md rewritten to the FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree card structure (logo, summary, run, scope, acknowledgements). No weights or provenance files changed.

Files changed (1) hide show
  1. README.md +25 -80
README.md CHANGED
@@ -12,9 +12,7 @@ tags:
12
  - text-to-audio-video
13
  - distillation
14
  - dmd2
15
- - vsa
16
  - few-step
17
- - 8-step
18
  - minimax-h3
19
  - fastvideo
20
  - fasth3
@@ -26,44 +24,25 @@ tags:
26
 
27
  # FastVideo-FastH3-8-Step-V2
28
 
29
- FastH3 8-Step V2 is a text-to-audio-video checkpoint from
30
- [FastVideo](https://github.com/hao-ai-lab/FastVideo), distilled from
31
- [MiniMax H3](https://huggingface.co/MiniMaxAI/MiniMax-H3). It generates video
32
- and stereo audio in eight transformer forwards using data-free DMD2 and
33
- Video Sparse Attention (VSA).
34
 
35
- [FastVideo](https://github.com/hao-ai-lab/FastVideo) ·
36
- [FastH3 blog](https://haoailab.com/blogs/fasth3-preview/) ·
37
  [FastH3 collection](https://huggingface.co/collections/FastVideo/fastvideo-fasth3)
38
 
39
- Powered by MiniMax H3.
40
-
41
-
42
- ## Model
43
-
44
- | Setting | Value |
45
- |---|---|
46
- | Task | Text to video with stereo audio (T2AV) |
47
- | Method | Data-free [DMD2](https://arxiv.org/abs/2405.14867), carried backward simulation |
48
- | Student attention | [VSA-H3](https://arxiv.org/abs/2505.13389), 80% sparsity, 64-token tiles |
49
- | Denoising | 8 transformer forwards |
50
- | Video / audio shifts | 10 / 3 |
51
- | Guidance scale | 1.0 |
52
- | Checkpoint | Training step 1300, bf16 full weights |
53
-
54
- Training used prompt conditioning and student-generated latents, not target
55
- video latents. The corpus used mixed resolutions and aspect ratios. This is
56
- an eight-step DMD2 checkpoint, not a PDD model.
57
-
58
- **This checkpoint requires FastVideo's VSA-H3 backend and Video Sparse
59
- Attention kernel.** The four-step models' default 90% sparsity and 12/3 shifts
60
- are not this model's sampling recipe. No matching eight-step LoRA is included.
61
 
62
  ## Run with FastVideo
63
 
64
- Start with the [installation guide](https://haoailab.com/FastVideo/getting_started/installation/)
65
- or [agent-guided installation](https://github.com/hao-ai-lab/FastVideo#install-with-an-ai-coding-agent).
66
- For NVIDIA B200 with CUDA 13, the public FastVideo environment is installed with:
 
 
67
 
68
  ```bash
69
  git clone https://github.com/hao-ai-lab/FastVideo.git
@@ -75,58 +54,24 @@ UV_TORCH_BACKEND=cu130 uv pip install \
75
  -e ".[fasth3]"
76
  ```
77
 
78
- Inference support for this checkpoint (video shift 10 and the trained
79
- eight-rung ladder) is in FastVideo
80
- [PR #1852](https://github.com/hao-ai-lab/FastVideo/pull/1852). Until it is
81
- merged, check out that branch; afterwards, `main` works unchanged.
82
-
83
  ```bash
84
  python examples/inference/basic/basic_fasth3_8step.py \
85
- --prompt 'A slow cinematic drone shot glides over a coastal town at golden hour; gulls call over the harbor as a church bell rings twice.' \
86
- --num-gpus 4 --vsa-kernel sm100a \
87
- --profile strict --no-inference-torch-compile --no-compile-vae \
88
- --height 768 --width 1344 --num-frames 124 \
89
- --output outputs/fasth3-8step
90
  ```
91
 
92
- `basic_fasth3_8step.py` pins this checkpoint and its trained recipe (nine
93
- sigma-grid points = eight forwards, VSA 0.8, 64-token tiles) and reads the
94
- shifts and rung ladder from `scheduler/`, `audio_scheduler/` and
95
- [fastvideo_inference.json](fastvideo_inference.json). It shares the FastH3
96
- preview example's CLI, so every other flag works unchanged; pass
97
- `--model-path` to use a local snapshot.
98
-
99
- To download ahead of time:
100
 
101
- ```bash
102
- hf download FastVideo/FastVideo-FastH3-8-Step-V2 \
103
- --local-dir ./FastH3-8-Step-V2
104
- ```
105
 
106
- The complete package is approximately 148 GB. It includes the distilled
107
- transformer, text encoder, tokenizer, processor, video VAE and audio VAE.
108
- It is stored in a modular Diffusers layout, but the supported runtime is
109
- FastVideo with VSA-H3, not an unmodified dense Diffusers pipeline.
110
-
111
- ## Scope and limitations
112
-
113
- - T2AV only. FL2VA and Ref2VA were not distilled or validated for this release.
114
- - Difficult motion, small details and some audio can remain below Base H3.
115
- - Sampling outside the trained schedule or attention policy is an ablation.
116
- - The four-step blog's latency and evaluation numbers do not establish this
117
- eight-step model's performance. On 4x GB200 (SP4, eager strict profile,
118
- VSA 0.8 tile 64, 832x480, 124 frames + 32 kHz audio) one measured request
119
- took 6.3 s end to end including saving (FastVideo PR #1852 smoke). Quality
120
- has not been formally compared against Base H3 or the four-step preview.
121
-
122
- This model inherits the [MiniMax H3 Community License](LICENSE), including
123
- its territorial, use and commercial restrictions. Review it before use or
124
- redistribution. The bundled Qwen3-VL encoder is covered by
125
- [Apache 2.0](LICENSE-Qwen3-VL). See [NOTICE](NOTICE) for attribution and modifications.
126
-
127
- Training and checkpoint provenance are preserved in
128
- [provenance.json](provenance.json), [checkpoint_metadata.json](checkpoint_metadata.json)
129
- and [checkpoint_content.json](checkpoint_content.json).
130
 
131
  ## Acknowledgements
132
 
 
12
  - text-to-audio-video
13
  - distillation
14
  - dmd2
 
15
  - few-step
 
16
  - minimax-h3
17
  - fastvideo
18
  - fasth3
 
24
 
25
  # FastVideo-FastH3-8-Step-V2
26
 
27
+ The FastH3 8-Step V2 checkpoint from
28
+ [FastVideo](https://github.com/hao-ai-lab/FastVideo). It generates synchronized
29
+ video and audio from text with eight transformer forwards. This step-1300 model
30
+ was trained with data-free DMD2 and VSA-H3 at 80% sparsity.
 
31
 
32
+ [Blog](https://haoailab.com/blogs/fasth3-preview/) ·
 
33
  [FastH3 collection](https://huggingface.co/collections/FastVideo/fastvideo-fasth3)
34
 
35
+ > This checkpoint requires FastVideo's VSA-H3 attention backend. Its video
36
+ > scheduler shift is 10, not the base model's 12; use the example below, which
37
+ > reads the trained schedule from the checkpoint.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
38
 
39
  ## Run with FastVideo
40
 
41
+ Install [uv](https://docs.astral.sh/uv/getting-started/installation/), then use
42
+ the CUDA 13 / Blackwell path below. It selects FastVideo's published CUDA
43
+ kernel wheel instead of compiling the kernel locally. See the
44
+ [installation guide](https://hao-ai-lab.github.io/FastVideo/getting_started/installation/)
45
+ for other platforms.
46
 
47
  ```bash
48
  git clone https://github.com/hao-ai-lab/FastVideo.git
 
54
  -e ".[fasth3]"
55
  ```
56
 
 
 
 
 
 
57
  ```bash
58
  python examples/inference/basic/basic_fasth3_8step.py \
59
+ --prompt "your prompt" \
60
+ --no-warmup \
61
+ --repeats 1
 
 
62
  ```
63
 
64
+ The tested defaults use four B200 GPUs and the trained eight-forward schedule.
65
+ On other multi-GPU CUDA systems, follow the installation guide and add
66
+ `--no-replicated-dit --vsa-kernel triton --no-fa4`. The GPU count must divide
67
+ H3's 56 attention heads.
 
 
 
 
68
 
69
+ ## Scope
 
 
 
70
 
71
+ This checkpoint supports text-to-audio-video generation. FL2VA and Ref2VA were
72
+ not distilled. Difficult motion, fine detail, and some audio may remain below
73
+ the base MiniMax H3 model. This checkpoint inherits the
74
+ [MiniMax H3 Community License](LICENSE).
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
75
 
76
  ## Acknowledgements
77