Skip to content

Get started ​

Install Vflash to inspect your GPU and choose a supported profile. For complete video generation on one RTX 4090 48 GB, optional paired Base16 keyframe generation on two matching 48 GB 4090s, or reference-image generation on two RTX 3080 20 GB GPUs, continue with the complete pipeline and official-weight preparation. The native bundle workflow is described below.

Before you begin

The Python pipeline and vflash generate support text, one to three images, or one short reference video → a five-second MP4 on SM89, plus reference-image generation on a cooperating SM86 pair. The denoise command and HTTP service accept compiled conditioning → video/audio latents, including the supported SM86 profiles. Model downloads and compilation are explicit steps; no weights are bundled with the package.

Version 0.4.0 also supports official Base16 I2VA/L2VA/FL2VA complete requests at integer durations from five through ten seconds. A prepared Base16 keyframe pipeline accepts a first frame, a last frame, or both without reloading weights. One ten-second, 736 × 992 SM89 L2VA case has bounded completion, media-integrity, endpoint and latency evidence; broader L2VA quality and other duration/canvas combinations remain unqualified. Turbo and reference-video output remain five seconds.

Install the CLI ​

Use Python 3.11 or newer. The base installation is lightweight and does not download model weights or PyTorch.

bash
git clone --branch v0.5.2 --depth 1 https://github.com/Hansimov/vflash.git
cd vflash
python -m venv .venv
source .venv/bin/activate
python -m pip install -e .

Check your GPU ​

bash
vflash doctor
vflash profiles

doctor lists the NVIDIA GPUs visible through nvidia-smi. profiles lists the configurations supported by this release.

Choose the GPU index shown by doctor and check a profile:

bash
# RTX 4090 with 48 GB
vflash plan ref2va-turbo4-exact-sm89 --gpu 0

# Two RTX 3080 GPUs with 20 GB each
vflash plan ref2va-turbo4-exact-sm86 --gpu 0 --peer-gpu 1

plan checks the hardware match without loading model weights. For the 3080 profile, we recommend at least 64 GiB of available system memory per worker for the tested workload, plus headroom for other processes. Larger inputs need separate capacity checks. These profiles target the stated memory capacities; the usual 10/12 GB 3080 and 24 GB 4090 are outside this release's supported configurations.

See profiles and hardware for the full support table.

Run a bundle ​

Use Docker for a pinned GPU environment. For a Python source installation, install the GPU dependencies in your environment:

bash
python -m pip install -e '.[gpu]'

The source runtime uses PyTorch 2.11 and Triton 3.6 on Linux with a compatible NVIDIA driver. The Docker build pins the CUDA 13.0 PyTorch build and its dependencies.

Prepare a matching artifact, schedule, auxiliary tensor file, and conditioning bundle. Replace the example paths with your files:

bash
vflash denoise ref2va-turbo4-exact-sm89 \
  --gpu 0 \
  --artifact /path/to/artifact \
  --schedule-overlay /path/to/schedule \
  --auxiliary-tensor /path/to/auxiliary.safetensors \
  --bundle /path/to/example-bundle \
  --output-latents ./result.safetensors

For a 3080, select ref2va-turbo4-exact-sm86 and use assets compiled for that profile. Renaming a profile or changing its GPU suffix does not convert the assets.

The command writes the video and audio latent tensors to result.safetensors and prints a JSON summary. The tensors are inputs for a compatible decoder; the file is not a playable video.

Use two GPUs for one request ​

To use two RTX 3080 20 GB devices for one request, keep the same SM86 Turbo4 assets:

bash
vflash plan ref2va-turbo4-exact-sm86 --gpu 0 --peer-gpu 1 --strategy sequence-head

Add the same --gpu, --peer-gpu, and --strategy options to vflash denoise. Both GPUs belong to one request; this does not start two independent workers.

StrategyWhat is dividedDefault when a peer is selected
tensorQKV, attention output, FFN, and LoRA projection weightsNo
sequence-headToken rows for GEMMs, then attention heads for complete-sequence attentionYes

Both strategies use BF16 weights, exact attention and the complete four-step schedule. Selecting a peer defaults to sequence-head. See the architecture guide for the implementation.

Version 0.4.0 also accepts two matching RTX 4090 48 GB devices for a prepared SM89 Base16 I2VA/L2VA/FL2VA profile. Pass the same peer options, but use sequence-head; tensor remains rejected. Both devices stream blocks from one prepared SM89 artifact. This opt-in path prioritizes one request's latency and should not replace two independent workers when two requests are ready.

Parallel execution changes GEMM shapes or reduction order, so results need not be bitwise identical to single-GPU execution. Full latent and decoded-media smoke checks passed on one workload; cross-case instruction-quality qualification remains pending. See performance measurement for scope and memory accounting.

Reuse a loaded model ​

For repeated requests, use a Python session or the HTTP service. A Python session loads during construction; the HTTP service loads on its first job. Both then reuse the model. Every new vflash denoise process loads a fresh model.

If a check fails ​

SymptomWhat to check
No GPU is listedRun nvidia-smi and check driver access. In Docker, confirm that the selected GPU is visible in the container.
The profile does not match the GPUSelect the profile for the device and memory capacity shown by doctor.
Runtime assets are missing or incompatibleCheck all four inputs and their model, adapter, schedule, and hardware versions. See runtime assets.
The 3080 process runs out of host memoryFree system memory or reduce the number of workers. Start with 64 GiB or more of available RAM per worker for the tested workload; check capacity for larger inputs.

Vflash · Native MiniMax H3 inference