Buckets:

HuggingFaceDocBuilder's picture
|
download
raw
3.64 kB

CacheDiT

CacheDiT speeds up DiT pipelines by reusing transformer block outputs across denoising steps. It does not need training and supports most Diffusers DiT pipelines, including Flux, Qwen-Image, Wan, and HunyuanVideo. Diffusers also has built-in caching which doesn't require an extra dependency.

Install CacheDiT from PyPI.

pip install -U cache-dit

Call cache_dit.supported_pipelines() to list the pipeline families CacheDiT supports.

import cache_dit

cache_dit.supported_pipelines()

Enable caching

Call cache_dit.enable_cache on a pipeline to cache it with the default settings, then run the pipeline as usual.

import torch
import cache_dit
from diffusers import FluxPipeline

pipeline = FluxPipeline.from_pretrained(
    "black-forest-labs/FLUX.1-dev", dtype=torch.bfloat16
).to("cuda")
cache_dit.enable_cache(pipeline)

image = pipeline(
    "A cat holding a sign that says hello world", num_inference_steps=28
).images[0]

CacheDiT also works with torch.compile. Compile the transformer after you call enable_cache. See the compile docs for settings that avoid recompilation with dynamic input shapes.

pipeline.transformer = torch.compile(pipeline.transformer)

Call cache_dit.summary after inference to log how many steps were cached and the residual differences between steps.

stats = cache_dit.summary(pipeline)

Call cache_dit.disable_cache to restore the original pipeline.

cache_dit.disable_cache(pipeline)

Configure the cache

DBCache (Dual Block Cache) computes the first n blocks (Fn) at every step. When their output barely changes from the previous step, it reuses the cached output for the remaining blocks, and it can recompute the last n blocks (Bn) to correct it. The TaylorSeer calibrator predicts the cached output from earlier steps instead of reusing it as is.

enable_cache defaults to DBCache with the first 8 blocks always computed (F8B0) and 8 uncached warmup steps. To trade speed for quality, raise Fn_compute_blocks or lower residual_diff_threshold (default 0.08). For the best quality at high cache rates, add the TaylorSeer calibrator.

from cache_dit import DBCacheConfig, TaylorSeerCalibratorConfig

cache_dit.enable_cache(
    pipeline,
    cache_config=DBCacheConfig(
        max_warmup_steps=8,
        Fn_compute_blocks=8,
        Bn_compute_blocks=0,
        residual_diff_threshold=0.12,
    ),
    calibrator_config=TaylorSeerCalibratorConfig(taylorseer_order=1),
)

For supported pipelines, CacheDiT already knows whether CFG runs as a separate forward pass. For other models, set enable_separate_cfg=True in DBCacheConfig if the model runs the conditional and unconditional passes separately, or False if it fuses them or doesn't use CFG.

See the DBCache design docs for how the Fn and Bn blocks work, and the cache benchmarks for speed and quality numbers.

Next steps

Xet Storage Details

Size:
3.64 kB
·
Xet hash:
c4d22b66b04e0ce1dc61bacc3729300c49514acabaca1780a2639b9055cd17d7

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.