Skip to content

Benchmark results ​

Each result belongs to its stated software version, hardware and workload. These published summaries preserve their original scope for reproduction; they are not a ranking of the current release. See profiles and hardware for current support and performance and quality for measurement guidance.

Native 544p-trained keyframes · bounded screen ​

Core ce23c10, with application experiment source 3648270f, completed a same-device comparison on one RTX 4090 48 GB. Each arm used BF16 block-ring execution, fresh prompt/image encoding, 768×320, five seconds at 24 fps, the same reference/prompt/seed per scene, and continuous decoded delivery. The candidate uses LightX544 eight evaluations / Torch Flash; the control uses official Base16 / Sol. This compares those combinations, not only step count.

SceneBase16/Sol requestNative544/8 requestBase16/Sol samplingNative544/8 sampling
Human group113.886 s61.953 s89.009 s36.935 s
Group action107.090 s56.165 s87.044 s36.152 s

Requests include fresh conditioning, sampling, video/audio decoding and MP4 delivery, but exclude process initialization and external queue/network time. Each arm ran these scenes serially; initialization was 76.873 / 76.087 seconds (control/candidate), charged additionally to its first result. Full first-result times were 190.759 / 138.040 seconds. This is one observation per scene and arm, not an interleaved drift-controlled benchmark or throughput test. The observed request reduction is 45.6–47.6%; do not transfer it to larger canvases or SM86.

All four complete AV files decoded. Twenty whole-frame observations per video and twelve native face crops per human video retained subjects and action without the previously observed two-pass color speckles. Expressions and motion trajectories changed; background small faces remained soft. Exact dialogue, audio semantics and all instruction details were not qualified. This supports an optional small-canvas candidate, not universal equal quality or a face repair.

The same candidate pipeline also completed true L2VA at 512×512 / five seconds in 60.630 seconds without reinitialization (sampling 39.783 seconds). Its input was a last-frame condition, not an I2VA reversal or endpoint paste. It is execution evidence, not a matched L2VA speed ratio. One-time native compilation took 278.153 seconds with explicit metadata-only verification; that cost is excluded from request times. Peak allocated GPU memory across these requests was about 8.3–8.4 GiB, not full device residency or host-RAM usage.

A separate true FL2VA request at 640×352 / ten seconds completed 240 frames with both input anchors: 178.171 seconds including 75.407 seconds initialization, or 102.763 seconds for the request (72.219 seconds sampling). It reused the same compiled artifact without recompilation. Whole-video review retained the transition from a grounded animated subject to flight, with evolving appearance between the distinct endpoint designs. This is one longer execution/visual observation, not a matched quality or speed comparison.

Exact H3 text-encoder prefix ​

Pinned Diffusers H3 encoding reads Qwen3-VL hidden_states[50], already bypassing the vocabulary head. Vflash now retains 51 decoder layers before offloading instead of executing all 64. Keeping only 50 would expose the final normalized state and change conditioning. The same idea is described in the author's pruned-model card; this implementation does not use that model's low-rank or quantized weights.

Frozen first-frame and last-frame requests each used 928 × 512, eight seconds and BF16 conditioning. Each A/B/A2 sequence used the same physical device; A2 restores all 64 layers. The encoder screen did not run denoising. Baselines were 4910216 on SM89 and 53d6687 on SM86, with pinned Torch 2.11.0, Diffusers 0.40.0 and Transformers 5.9.0.

Encoding device / limitInputA capture51-layer captureA2 capture
RTX 4090 48 GB / 350 WFirst frame16.778 s11.006 s13.444 s
RTX 4090 48 GB / 350 WLast frame11.953 s9.862 s11.963 s
RTX 3080 20 GB / 200 WFirst frame18.834 s12.695 s15.622 s
RTX 3080 20 GB / 200 WLast frame13.666 s11.184 s13.557 s

Each candidate and A2 preserved all 14 conditioning tensors bit for bit on its architecture. The first A call includes first-use overhead; relative to A2 the local capture reduction is about 18–19%, not a complete-video or denoising speedup. The SM86 screen used the encoding device of a reserved pair; it was not parallel text encoding. Another GPU workload shared the host during the SM86 run, so these are bounded stage measurements, not isolated throughput. Removing the tail releases references to 6,338,778,368 BF16 parameters (11.81 GiB of tensor storage), not a measured process-RSS reduction; original checkpoint loading is unchanged. Private experiment ownership and receipts remain in the application repository at 73590e21.

A complete eight-second first-frame A/B on the same 350 W SM89 preserved all 192 delivered RGB frames (273,678,336 channel values) and all 512,000 decoded stereo PCM16 samples exactly. The final 64-layer control also preserved every pixel and decoded PCM16 sample. The preloaded generate calls took 303.798 / 303.251 / 305.163 seconds (A/B/A2); encoding took 18.309 / 16.443 / 18.421 seconds. Denoising remained unchanged and dominant. This small total difference is not a formal latency claim or evidence that the stage percentage applies to a whole video. Preloading, queueing and network delivery are outside these call times. Preloading itself took 79.601 / 74.667 / 72.101 seconds, without establishing a cold-start speedup. The implementation does not claim to repair pre-existing semantic or visual defects.

4090 FFN fusion · promoted in 0.2.2 ​

The implementation promoted in 0.2.2 was measured against the 0.1.0a6 native runtime on the same RTX 4090 48 GB at its default 450 W limit. The fixed workload used BF16 Ref4 v0.1 (rank 128, alpha 8, scale 0.0625), 928 × 512, 124 model frames, 20,828 packed tokens and four evaluations through all 50 layers. Both implementations used block streaming, the same conditioning bundle, cuBLAS GEMMs, Torch Flash attention, PyTorch 2.11.0+cu130 and Triton 3.6.0. Only FFN adapter merge plus SiLU changed. Encoding, VAE decoding, MP4 export and model initialization are outside this native request boundary.

Warm native requestRepeatsMedianRangePeak allocated GPU memory
Separate merge and activation343.033 s42.994–43.082 s9,281,030,144 B
Combined kernel342.569 s42.539–42.600 s8,683,849,728 B

The median reduction is 0.464 s / 1.08%; the GPU allocation reduction is 597,180,416 B. Interleaved requests bracketed the candidate with original-implementation runs; anchor drift was 0.204%, temperature reached 75°C and thermal counters did not grow. All final FP32 video/audio latents matched exactly. This is a target-hardware measurement of one fixed native workload, not a broad speed or quality guarantee.

Complete-pipeline integration was then checked with a different packed length, 18,922 tokens, on the 0.2.1 pipeline using the same kernel. A three-reference original/fused/original sequence preserved conditioning, final latents and every decoded video frame, and cancellation released owned storage. In that case denoising allocation fell from 8,577,681,920 to 8,035,150,336 B; the maximum observed across pipeline stages stayed at 9,945,532,928 B. The runs had different first-use caches and only one sample per arm, so their total times do not establish an end-to-end speed ratio. Host RSS peaked at 114.38 GiB. The 0.2.2 release notes describe the supported dispatch and remaining audio limits.

Two 3080s · a3 ​

The comparison below was measured on the a3 runtime, before a4 introduced segmented pinned storage. Its original timings and memory figures are retained as historical evidence, not a benchmark of every later release. The a5 loader and lifecycle changes do not update this ranking.

A target-hardware measurement on one frozen Ref2VA Turbo4 request (928 × 512, 124 model frames, four evaluations, 18,175 tokens, BF16 weights and exact attention) compared one RTX 3080 20 GB with the same primary device plus a second 3080. Both ran at their default 320 W limits, on PCIe 3.0 x16 host-bridge links without peer access. The runtime used PyTorch 2.11.0+cu130 and Triton 3.6 with eight CPU threads.

Warm executionRepeatsMedian conditioning-to-latent timeRange
One GPU4 anchors85.694 s85.539–85.805 s
Two GPUs, sequence-head349.671 s49.599–49.973 s

This is a 1.725× speedup on the measured workload. Three interleaved A/B/A2 comparisons had at most 0.227% anchor drift, no sampled thermal throttling or thermal-counter growth, and successful CUDA probes before and after every request. The same process retained both single/parallel device rings and shared their host weights to avoid repeated loading during the comparison. Memory figures below come from a separate standalone session.

That a3 standalone sequence/head session initialized in 41.758 s with an already warm filesystem cache; its first request took 51.186 s and its next request 49.441 s. Denoising allocation peaked at 4.745 / 4.614 GiB across the pair. Active pinned host allocation was 58.009 GiB, with peak process RSS of 59.90 GiB. Initialization, storage cache, GPU allocation, reserved memory and process RAM are distinct costs.

Standard weight tensor completed three warm requests in a separate standalone session: median 58.615 s, range 58.526–58.709 s, approximately 1.462× against the preceding same-primary single-GPU anchors. Initialization took 47.899 s and the first request 60.079 s. Denoising allocation peaked at 4.825 / 4.785 GiB, pinned host allocation at 58.792 GiB, and process RSS at 60.70 GiB. This is a separate TP screen without a new interleaved A/B/A2 comparison. One sample reported a transient software thermal flag without thermal-counter growth or a corresponding clock reduction. All three results are retained; this screen has a narrower evidence level than the isolated comparison above.

These measurements stop at the exported latent file. Prompt/reference encoding, VAE decoding, MP4 encoding, queueing and network transfer are outside the boundary. Two cooperating GPUs improve one-request latency; independent single-GPU workers provide a different throughput tradeoff.

Both parallel strategies completed full trajectories and a paired decoded-video/audio smoke. They are not bitwise equal to single-GPU execution. The pipelined head exchange itself matched the unpipelined sequence/head trajectory bitwise, but the partitioning changes GEMM shapes and standard tensor parallelism adds reduction boundaries. One decoded example is not cross-case quality qualification; no same-quality or general prompt-to-video speed claim is made.

4090 host allocation · a4 ​

One Ref2VA check on a single RTX 4090 48 GB / SM89 used three reference images, 928 × 512, and 124 model frames at 24 fps. It compared the 0.1.0a3 allocator with the allocator shipped in 0.1.0a4, using PyTorch 2.11.0+cu130. It kept the fixed BF16 Ref Turbo4 v0.1 configuration: 4 NFE, video/audio shifts 12/3, LoRA strength 1 and alpha 8. Active pinned weight storage fell from about 58.009 GiB to 40 GiB. With identical inputs, final video and audio latents matched the previous allocator bit for bit. Warm native denoising took 41.607 s before and 41.641 s after: the measured benefit is lower host memory use with comparable inference time. The 40 GiB figure covers pinned weight storage; the process still needs additional RAM.

A separate same-GPU, four-step capacity check compared resident weights with block streaming. Final latents matched bit for bit, while block streaming took longer. That check supports the capacity option, not a claim that streaming is faster.

Vflash · Native MiniMax H3 inference