Measured performance

This page is generated from the anonymized public benchmark catalog. Values are measurements for the declared revisions and workloads, not universal claims.

1344×768 Ref Turbo4 hardware comparison

Path Generation Cold total vs 3080 FA3 Peak VRAM Peak temp Scope
单 4090 · Sage2 188.474 s 274.324 s 2.241× 18118 MiB 69 °C 单变量 Sage2;生成阶段
单 4090 · FA3 210.669 s 287.95 s 2.005× 18118 MiB 69 °C 硬件主对照
双 3080 · TP2 · FA3 269.621 s 323.036 s 1.567× 12234 MiB 79 °C 每卡峰值取较高者;无 NVLink
单 3080 · Sage2 353.502 s 432.242 s 1.195× 17918 MiB 82 °C 单变量 Sage2
单 3080 · FA3 422.366 s 499.874 s 17918 MiB 77 °C 硬件基线

15-second Base path

Path End to end Denoise Speedup Status
Base40 · Torch RoPE 1427.813 s 1335.917 s 0.953× measured
Base40 · 严格 Triton RoPE 1360.634 s 1270.942 s measured
Base40 · warm20 cache 1150.903 s 1057.396 s 1.182× measured
Base40 · AdaLN-BF16 混合 INT8 988.805 s 900.106 s 1.376× experimental
Base24 · BF16 853.396 s 763.469 s 1.594× measured
Base16 · BF16 实验 597.222 s 510.426 s 2.278× experimental

Bottleneck ledger

Profile E2E Denoise Denoise share Conditioning VAE Encode
Base40 BF16 1360.634 s 1270.942 s 93.4% 40.509 s 38.999 s 8.556 s
Base24 BF16 853.396 s 763.469 s 89.5% 40.428 s 37.747 s 10.076 s
Turbo4 BF16+Sage 208.609 s 132.667 s 63.6% 31.957 s 37.24 s 6.101 s
Turbo4 INT8+Sage 178.833 s 96.226 s 53.8% 31.294 s 37.396 s 13.256 s

Actual workload matrix

# Profile Canvas Timeline Measured time Basis Evidence family
1 Ref Turbo4 · 单3080 FA3 1344×768 124帧 / 原生约5.17s@24 422.366s 生成;499.874s 含初始化 生成阶段 / 冷合计 HW-1344;仅性能资格
2 Ref Turbo4 · TP2双3080 FA3 1344×768 124帧 / 原生约5.17s@24 269.621s 生成;323.036s 含初始化 生成阶段 / 冷合计 HW-1344;仅性能资格
3 Ref Turbo4 · 单4090 FA3 1344×768 124帧 / 原生约5.17s@24 210.669s 生成;287.950s 含初始化 生成阶段 / 冷合计 HW-1344;仅性能资格
4 Ref Turbo4 · 单3080 Sage2 1344×768 124帧 / 原生约5.17s@24 353.502s 生成;432.242s 含初始化 生成阶段 / 冷合计 HW-1344;仅性能资格
5 Ref Turbo4 · 单4090 Sage2 1344×768 124帧 / 原生约5.17s@24 188.474s 生成;274.324s 含初始化 生成阶段 / 冷合计 HW-1344;仅性能资格
6 Ref Turbo4 · 单3080 928×512→910×512 124原生帧→15s/120帧@8 194.778s 生成;197.542s 热态交付;277.830s 冷交付 生成 / conform / 初始化 08-28历史输入;低fps预览
7 Turbo4 BF16 FA3 · 单4090 928×512 362帧 / 15.08s@24 252.424s 第2轮热态生成 4-update矩阵
8 Turbo4 BF16 Sage2 · 单4090 928×512 362帧 / 15.08s@24 208.609s 第2轮热态生成 4-update矩阵
9 Turbo4 BF16 Sage2 · 单3080 928×512 362帧 / 15.08s@24 461.508s 首轮生成,不含初始化 4-update矩阵
10 Turbo4 INT8 Sage2 · 单3080 928×512 362帧 / 15.08s@24 378.991s 首轮生成,不含初始化 4-update矩阵;有损
11 Base40 BF16 严格 · 单4090 928×512 362帧 / 15.08s@24 1360.634s(22m40.6s) 生成E2E;去噪1270.942s STRICT-LONG
12 Base40 cache · 单4090 928×512 362帧 / 15.08s@24 1150.903s(19m10.9s) 生成E2E;去噪1057.396s STRICT-LONG;实验档
13 Base40 hybrid INT8 · 单4090 928×512 362帧 / 15.08s@24 988.805s(16m28.8s) 生成E2E;去噪900.106s STRICT-LONG;实验
14 Base24 BF16 · 单4090 928×512 362帧 / 15.08s@24 853.396s(14m13.4s) 生成E2E;去噪763.469s STRICT-LONG;生产快速档
15 Base16 BF16 · 单4090 928×512 362帧 / 15.08s@24 597.222s(9m57.2s) 生成E2E;去噪510.426s STRICT-LONG;实验
16 Base40 BF16 · 双4090两seed均值 约928×512 124原生帧→15s@8 545.020s 单候选均值 TEMPORAL;历史批次
17 Base50 BF16 · 双4090两seed均值 约928×512 124原生帧→15s@8 686.885s 单候选均值 TEMPORAL;历史批次
18 Base40 BF16 · 完整时轴 约928×512 约362帧 / 15s@24 2126.672s 历史完整时轴 TEMPORAL;只与第16行比较
19 Base50 BF16 · 完整时轴 约928×512 约362帧 / 15s@24 2336.157s 历史完整时轴 TEMPORAL;只与第17行比较
20 Base50 BF16 · 5秒 928×512 124帧 / 5.17s@24 533.180s 历史生成E2E CANVAS-AB
21 Base50 BF16 · 5秒 512×288 124帧 / 5.17s@24 354.940s 历史生成E2E CANVAS-AB
22 Ref Turbo4 · 双4090独立worker 512×288 124帧 / 5.17s@24 两候选批次87.196s;单条86.325/78.760s 任务生成墙钟 并发验收;非单候选TP

Optimization matrix

# Method Workload Before → after Speed Quality/status Evidence
1 双 3080 TP2 1344×768×124,Turbo4,FA3 422.366→269.621 s 生成 1.567× 精确并行;单4090仍快1.280× A
2 单 4090 替代单 3080 同上 422.366→210.669 s 生成 2.005× 跨架构只比较速度,不做像素真值 A
3 SageAttention2 · 3080 1344×768×124,Turbo4 422.366→353.502 s 生成 1.195× 性能预览;画布超Ref资格范围 A
4 SageAttention2 · 4090 1344×768×124,Turbo4 210.669→188.474 s 生成 1.118× 性能预览 A
5 SageAttention2 · Base40 928×512×362,单4090 1851.519→1443.667 s E2E 1.283× 四轮盲评全平;生产 A
6 SageAttention2 · Base50 928×512×362,单4090 2336.157→1779.808 s E2E 1.313× 四轮盲评全平;生产 A
7 严格 Triton RoPE 15秒 Base40,单4090 1427.813→1360.634 s E2E 1.049× SSIM 0.994831,音频逐样本相同;生产 A
8 Pattern3 warmup20 cache 15秒 Base40,单4090 1360.634→1150.903 s E2E 1.182× SSIM 0.990041;显式实验档 A
9 Base50→Base40 928×512×124,热态 600.833→490.876 s E2E 1.224× 改变轨迹;Base40生产 A
10 Base40→Base24 928×512×362,单4090 1360.634→853.396 s E2E 1.594× SSIM 0.845058;盲评平,运动敏感场景有风险;生产快速档 A
11 Base40→Base16 928×512×362,单4090 1360.634→597.222 s E2E 2.278× 质量余量不足;实验 A
12 AdaLN-BF16 混合 INT8 15秒 Base40,单4090 1360.634→988.805 s E2E 1.376× SSIM 0.857694;运动+0.516%,音频+0.607dB,盲评全平;待扩样 A
13 全局 INT8 + Sage2 15秒 Base40,单4090 1851.519→1060.874 s E2E 1.745× 运动-21.99%,音频-4.244dB;拒绝生产 A
14 INT8 + Sage2 · 4090 928×512×362,Turbo4 热态 252.424→178.833 s E2E 1.412× 快速预览信号;量化质量风险 A
15 INT8 + Sage2 · 3080 928×512×362,Turbo4 461.508→378.991 s E2E 1.218× 20G可跑;相对BF16 Sage B
16 15秒低帧率预览 · Base40 362→124原生帧,交付8fps 2126.672→545.020 s 3.902× 轨迹改变;仅粗预览 A
17 15秒低帧率预览 · Base50 362→124原生帧,交付8fps 2336.157→686.885 s 3.401× 轨迹改变;仅粗预览 A
18 降低画布 5秒 Base50,928×512→512×288 533.180→354.940 s 1.502× 小物体/材质/空间关系风险 A
19 双4090独立 worker 512×288 Turbo4,两候选 86.325+78.760→87.196 s 批次 观测并发比1.893× 吞吐收益,不降单候选延迟 B
20 Ref图 2048→match 合同 SGLang INT8+Sage,单3080 309.458→95.888 s wall;224.006→54.270 s去噪 3.227× / 4.128× 候选盲评4:0胜;先前多算了token A
21 SGLang exact FA→Sage 928×512×124,4 updates,单3080 114.036→96.681 s wall;74.047→56.827 s去噪 1.180× / 1.303× SSIM 0.936698,运动+0.916%,音频-0.226dB,盲评平 A
22 精确 AdaLN 跨请求缓存 928×512×124 Base50,单4090 624.100冷→600.833热均值 约1.039× E2E 数学精确;冷热含其他复用因素 B
23 ConvRot INT8 真实H3一步,4090/3080 相对INT8慢1.38% / 2.12% 负收益 已否决 A
24 FP8-SGL 真实H3稳态,单4090 7.315→9.320 s/step 慢27.4% 显存不降;已否决 A
25 BF16常驻24 blocks 928×512 Base50,单4090 512.364→512.364 / 533.182→536.452 s 无收益或更慢 峰值约42.5GB;已否决 A
26 请求内固定前处理缓存 token/refiner/RoPE 后续步骤仅快0.283% E2E无改善 已否决 A
27 移除音频 token 512×288×124 Base40 470.464→470.035 s 0.09%噪声级 改变视频结果;已否决 A
28 Cache warmup4 本机Base40 actual rollout E2E约1.462× 更快但不过门 SSIM 0.8431;已否决 A
29 Taylor/Scaling/分段 cache Base40 observe-only oracle direct reuse rel-L1中位20.70%,最优分组仍20.56–20.61% 未进入E2E 预测基底失败且history过大;关闭 C
30 W4A4 / SVDQuant 六种真实H3 linear shape 完整linear微基准1.877×;估算去噪上限1.234× 无模型E2E residual/scale/provider未过门;候选 C

Prompt-treatment A/B

All four variants used one paired concurrent group with the same model, seed, canvas, frame rate and NFE. These are contended throughput measurements, not isolated latency.

Variant Treatment Generation Request wall Peak VRAM File size
raw original Chinese brief 154.948304 s 155.028677 s 14293.724 MiB 14.75 MB
skill local H3 prompt-writing skill 165.818291 s 165.902783 s 14340.082 MiB 14.48 MB
qwen Qwen3.8-27B constrained rewrite 151.199552 s 151.259305 s 14563.726 MiB 15.56 MB
skill-qwen skill draft then Qwen3.8-27B constrained polish 154.86052 s 154.941888 s 14587.016 MiB 14.66 MB
Candidate Raw wins Candidate wins Ties Raw median Candidate median
skill 0 0 3 9 9
qwen 0 0 3 9 9
skill-qwen 0 0 3 8 8

Conclusion: No quality improvement was demonstrated on this already-specific brief.

Media publication: withheld: MiniMax H3 Community License territorial display restriction requires additional authorization for globally accessible hosting

Interpretation limits

  • Measurements apply only to the declared revisions, inputs and contracts.
  • Cross-architecture videos are not pixel-equivalent quality references.
  • Throughput from independent workers is not single-request latency speedup.
  • Experimental and rejected rows remain visible to prevent cherry-picking.
  • No model weights, GPU identifiers, user jobs, private paths or generated media are included.

The complete machine-readable source is benchmarks/h3-consumer-gpu-2026-08-31.json.