Vflash

Vflash is the clean, reusable publication layer for measured video-inference work. It packages a pinned MiniMax H3 / LightX2V runtime, Docker deployment, public-safe topology reporting, benchmark contracts and quality evidence without carrying a product UI, private prompts, machine inventory or raw experiment history.

What is measured today

Question Measured answer Contract
Can one 20 GB RTX 3080 run 768p H3? Yes: 17.9 GiB peak in the controlled Ref Turbo4 run 1344×768, 124 frames, 4 NFE
Does TP2 help on two RTX 3080 cards? 422.366 → 269.621 s, 1.567× PCIe PHB, no NVLink
How does one RTX 4090 compare? 210.669 s, 2.005× vs one 3080 and 1.280× vs TP2 same FA3 generation contract
Does SageAttention2 help? 1.195× on 3080; 1.118× on 4090 same 1344×768 Turbo4 contract
Can one 3080 deliver a 15 s, 8 fps preview? 197.542 s warm, 277.830 s cold native 124 frames conformed to 120 frames at 910×512
What dominates Base40? Denoising is 93.4% of a measured 15 s request 928×512, 362 frames, one 4090

These are local measurements for declared revisions, not upstream marketing claims or universal hardware guarantees. The full catalog includes accepted, experimental, rejected and incomplete routes so negative results remain visible.

Current runtime choices

  • Ref Turbo4 is the low-latency reference path.
  • Base40 is the balanced full-path baseline; Base50 is the slower quality anchor.
  • Base24 is a measured fast quality tradeoff, not a lossless replacement.
  • Base40 request-scoped cache is qualified as an explicit experimental option.
  • Global INT8, Base16, FP8-SGL and ConvRot are not production defaults.
  • Two 4090 replicas improve candidate throughput; TP2 is reserved for a single request that truly needs multi-GPU parallelism.

Evidence levels

Label Meaning
upstream claim Reported by a framework or model author; not reproduced here
measured Reproduced under one declared local contract
qualified Passed additional inputs and quality checks
production Enabled by default in a deployed profile
experimental Runnable, but quality or generalization is incomplete
rejected Measured and intentionally not selected

Start here

Model and output licensing

Vflash code is Apache-2.0, but MiniMax H3 is not. Its current Community License restricts the territory in which the model and its Outputs may be displayed. No H3 weights or generated videos are distributed by this repository. Read Model publishing before serving, quantizing or publishing.