Skip to content

0.6.12Stable release

Native H3 inference.
Built for your GPU.

A MiniMax H3 engine measured on SM86 and SM89. Run Turbo or Base16 keyframe generation through Python, containers or HTTP.

SM86 / SM89Turbo4 / Turbo8 / Base16Apache 2.0

Choose a denoising configurationRTX / CUDA
48GBVRAM capacity
Execution
One GPU · resident/streamed
Available profiles
Turbo4 / Turbo8 / Base16

Serve text, references and keyframes; the complete pipeline reuses one GPU across stages.

Inspect hardware without loading weights
vflash plan ref2va-turbo4-exact-sm89 \ --gpu 0

The current public interface

Ref/T2 Turbo retains five seconds; Base16 and optional 544p keyframes accept five through ten. Single-SM89 Base16 defaults to Sol. Original v0.1 keyframes permit explicit Sol but default to dense. Media evidence is scoped by mode, not a blanket quality guarantee. Check the prerequisites

Start with what you need

Understand it. Build on it.

Native PyTorch and Triton execution, with owned LoRA, memory scheduling and two-GPU collective layouts.

Explore the architecture

Vflash · Native MiniMax H3 inference