Python integration
Use H3Pipeline to generate complete videos from text or reference images on SM89. Use a NativeEngineSession when your application already prepares compatible conditioning bundles and needs to reuse loaded weights. A session owns one fixed profile and GPU group, and processes requests serially.
For process isolation and an HTTP queue, use the Docker service instead.
Create a session
In 0.5.0 both NativeEngineSession and H3Pipeline default to attention_backend="auto": single-SM89 official Base16 selects approximate sol-sm89; other supported profiles and pairs select torch-flash. Set attention_backend="torch-flash" explicitly for dense execution. Sol native sessions use block-ring residency (default resolves to it); an explicit resident request is incompatible with Sol. Install the pinned dependency with python -m vflash.install_sol after GPU/pipeline dependencies, or use the standard Docker build. Missing Sol fails without fallback. Prepared assets are unchanged; the result's attention policy identifies actual approximate execution.
Since 0.6.6, the original LightX v0.1 four-step keyframe pair also accepts explicit sol-sm89 on single SM89. Its auto remains dense: three matched complete requests were slower with Sol. Do not infer support for other Turbo profiles.
Install .[gpu] in the supported environment, prepare the four runtime inputs, and replace the paths below. Select devices before your process initializes CUDA.
from pathlib import Path
from vflash.catalog import ProfileCatalog
from vflash.hardware import discover_nvidia_devices
from vflash.native.runner import NativeEngineSession
from vflash.planner import resolve_plan
# Physical GPU indices are the ones reported by `vflash doctor`.
devices = {device.index: device for device in discover_nvidia_devices()}
plan = resolve_plan(
ProfileCatalog.bundled(),
profile_id="ref2va-turbo4-exact-sm89",
device=devices[0],
)
outputs = Path("./outputs")
outputs.mkdir(parents=True, exist_ok=True)
with NativeEngineSession(
plan,
artifact=Path("/path/to/artifact"),
schedule_overlay=Path("/path/to/schedule"),
auxiliary_tensor=Path("/path/to/auxiliary.safetensors"),
) as session:
for name in ("example-a", "example-b"):
result = session.generate(
Path("/path/to/bundles") / name,
outputs / f"{name}.safetensors",
)
print(result["session"])Each output contains video and audio latent tensors for a compatible decoder. It is not an MP4. The returned session fields separate initialization and request timing.
Select a 3080 pair
Replace the plan above and use the matching SM86 assets:
plan = resolve_plan(
ProfileCatalog.bundled(),
profile_id="ref2va-turbo4-exact-sm86",
device=devices[0],
peer_device=devices[1],
strategy="sequence-head",
)Both devices belong to the same request. tensor is another supported strategy; omitting the peer selects single-device execution. See supported hardware and memory before adding concurrent workers.
Receive completed-step progress
Pass a callback when your application needs denoising progress:
def on_step(completed: int, total: int) -> None:
print(f"Denoising: {completed}/{total}")
# Within the active session:
result = session.generate(
Path("/path/to/bundles/example-a"),
Path("./outputs/example-a.safetensors"),
progress_callback=on_step,
)The callback runs after that evaluation finishes on every selected GPU. It is denoising progress, not a percentage of total video-generation time. Keep callbacks short. Enabling them adds synchronization at each step; omit the callback if you do not need notifications.
Leave more VRAM for activations
A 4090 session can pass weight_residency="block-ring" at construction to stream weights from host RAM. Its default is resident weights. The 3080 profile always streams blocks; it does not support full BF16 residency.
Measure both memory and latency with your own inputs. Streaming reduces device weight storage but needs substantial host RAM. The selected strategy remains fixed for the session.
Close the session
Use with, as above, or call session.close() when its owner stops. Closing waits for the selected devices, closes communication resources and releases owned weights. Keeping a reference to the closed Python object does not retain those weights.
The process still owns its CUDA context and allocator caches; process exit releases them. A closed session rejects new requests. After failed inference, close the session rather than trying to resume it. If device completion or cleanup fails, stop its worker process instead of reusing that CUDA context.
Use a separate spawned process for each independently owned GPU group. Do not call one session concurrently, initialize CUDA before device selection, or attempt to change its profile between requests. Changing models or compiled adapters requires a new session with matching assets. The experimental FP32 DiT attention LoRA context is a separate, reversible low-level option for an exclusively owned single-SM89 Base16 session; it does not change the compiled assets or defaults.