Skip to content

Latest commit

Β 

History

154 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

LongLive logo

🎬 LongLive

Long video generation research from NVIDIA. This repository hosts three generations of the project, each in its own directory with its own code, documentation and model weights.

Paper Paper Paper

Demo LongLive-Plug Demo LongLive 2.0 Demo LongLive 1.0

What's in this repository

Directory What it is Use it when you want to Venue
LongLive-Plug/ Once-for-all distillation for video generation Distill a capability once on a base model and reuse it across downstream models, without retraining arXiv
LongLive2.0/ An NVFP4 parallel infrastructure for long video generation Train or serve long-video models fast, with NVFP4 quantization and sequence parallelism arXiv
LongLive1.0/ Real-time interactive long video generation Type prompts and watch a long video appear in real time, steered as you go ICLR 2026

Each directory is self-contained: clone the repository, cd into the one you need, and follow its own README.

git clone --single-branch --branch main --depth 1 https://github.com/NVlabs/LongLive.git
cd LongLive/LongLive-Plug   # or LongLive2.0, or LongLive1.0

Demo videos

LongLive-Plug LongLive 2.0 LongLive 1.0
LongLive-Plug overview video β€” watch on YouTube LongLive 2.0 overview video β€” watch on YouTube LongLive 1.0 overview video β€” watch on YouTube
Once-for-all distillation: train a capability once, reuse it downstream NVFP4 parallel infrastructure for training and inference Real-time interactive long video generation

Click a thumbnail to watch on YouTube.

News

  • πŸ”₯ [2026.09.28] We release LongLive-Plug training and inference code with four recipes covering CFG and few-step distillation. β†’ LongLive-Plug/
  • πŸ”₯ [2026.07.08] LongLive 2.0 supports FP8 inference. Please refer to here.
  • πŸ”₯ [2026.06.01] We released LongLive-RAG, a general retrieval-augmented framework for long video gen.
  • πŸ”₯ [2026.05.30] LongLive 2.0 now supports I2V AR teacher-forcing training and I2V DMD distillation for Wan2.2-TI2V-5B.
  • ⚑ [2026.05.25] We optimized the NVFP4 inference path with fused Triton RoPE/adaLN kernels, reduced KV-cache synchronization overhead, in-place quantized KV-cache updates, faster FP4 KV dequantization, pinned VAE transfers, and safer LoRA-before-quantization setup, improving overall throughput by 18.6%.
  • πŸ”₯ [2026.05.13] We release LongLive 2.0, infra with NVFP4, parallelism and multi-shot for AR training, DMD distillation, and inference (⚑45.7 FPS).
  • πŸ”₯ [2026.04.12] LongLive supports kv cache compression with TriAttention, with 50% KV reduction and no quality drop. Check it here
  • πŸŽ‰ [2026.01.27] LongLive is accepted by ICLR-2026.
  • πŸ”₯ [2026.01.11] LongLive supports adapting LongLive's original RoPE into KV-cache relative RoPE and generates infinite long videos!
  • πŸ”₯ [2025.11.03] We implement LongLive on linear attention model SANA-Video! Now SANA-Video can generate 60s interactive videos in real-time.
  • πŸ”₯ [2025.09.29] We release Paper, this GitHub repo LongLive with all training and inference code, the model weight LongLive-1.3B, and demo page Website.

Models

LongLive-Plug model collection

LongLive-Plug Directory Supported Models
LongLive-Plug-MiniMax-H3-few-step LongLive-Plug/ H3-World, Code World Model, SolarWM-H3, Fun ControlNet-Union, LineartAnime, ...
LongLive-Plug-MiniMax-H3-cfg LongLive-Plug/ MiniMax-H3 (base model), SolarWM-H3
LongLive-Plug-Wan2.1-T2V-14B-few-step LongLive-Plug/ FantasyWorld, DreamZero, Fun Control, Wan-Move, MagicTryOn, ...
LongLive-Plug-Wan2.1-T2V-14B-cfg LongLive-Plug/ FantasyWorld, DreamZero, Fun Control, Wan-Move, MagicTryOn, ...
LongLive-Plug-Wan2.2-TI2V-5B-few-step LongLive-Plug/ Matrix-Game 3.0, SCOPE, Fast-WAM LIBERO, Kiwi-Edit, Ovi, ...
LongLive-Plug-Wan2.2-TI2V-5B-cfg LongLive-Plug/ Matrix-Game 3.0, SCOPE, Fast-WAM LIBERO, Kiwi-Edit, Ovi, ...

Model examples are from the paper appendix, Complete Transfer Coverage and Additional Cases; each row lists up to five examples. Wan coverage includes both few-step and CFG transfer. MiniMax-H3 few-step and CFG adapters are used separately.

Model Directory FPS ↑ Params VBench ↑ Multi-shot
LongLive-2.0-5B LongLive2.0/ 24.8 5B 85.06 βœ…
LongLive-2.0-5B-NVFP4-4Step LongLive2.0/ 29.7 5B 84.51 βœ…
LongLive-2.0-5B-NVFP4-2Step LongLive2.0/ 45.7 5B 83.14 βœ…
LongLive-1.3B LongLive1.0/ 20.7 1.3B 84.87

LongLive-Plug β€” Once-for-All Distillation for Video Generation

Separates reusable capabilities from task-specific customization: train functional LoRAs on a base model, then attach them to compatible downstream models while keeping those models' task-specific weights. Supports Wan2.1-14B, Wan2.2-TI2V-5B and MiniMax-H3.

LongLive-Plug method overview

β†’ Code and documentation in LongLive-Plug/

LongLive 2.0 β€” An NVFP4 Parallel Infrastructure for Long Video Generation

Training and inference infrastructure built around NVFP4 quantization and sequence parallelism.

  • For training, it supports
    • Balanced sequence parallel for T2V/I2V AR training (teacher-forcing).
    • T2V/I2V AR training on multi-shot (or single-shot) videos.
    • NVFP4 (or BF16) for both AR training and few-step distillation.
  • For inference, it supports
    • NVFP4 inference (W4A4) and NVFP4 KV Cache.
    • TorchAO FP8 PTQ inference (W8A8) from the BF16 checkpoint.
    • Multi-shot attention sink.
    • Sequence parallel inference.
    • Async decoding.

LongLive 2.0 framework overview

β†’ Code and documentation in LongLive2.0/

LongLive 1.0 β€” Real-time Interactive Long Video Generation

Accepts sequential user prompts and generates the corresponding video in real time, so a person can steer a long video while it is being produced. The key ideas are the attention sink, KV-recache, and streaming long tuning.

LongLive 1.0 overview

β†’ Code and documentation in LongLive1.0/

Awesome work using LongLive

  • DreamForge-World 0.1: Adapts the LongLive AR video stack with a residual action pathway for low-compute real-time controllable world modeling.
  • DreamX-World 1.0: Follows LongLive by adapting the model on long sequences with long rollouts and local temporal windows for stable long-horizon AR world generation.
  • SANA-Video: Combines SANA-Video with LongLive to build LongSANA, a real-time minute-long video generation variant with constant-memory KV cache.
  • Daydream Scope: Wraps LongLive as a streaming AR video diffusion pipeline for interactive text-to-video and video-to-video workflows.
  • MemFlow: Builds on the LongLive codebase and adds adaptive memory retrieval for more consistent long narrative video generation.
  • ShotStream: Builds on LongLive's distillation procedure for real-time streaming multi-shot AR video generation.
  • Stream-T1: Builds on LongLive's codebase and algorithm, adding test-time scaling with noise propagation, reward pruning, and memory sinking.
  • KVPO: Builds on LongLive and related AR video codebases to perform GRPO-style alignment through historical KV semantic exploration.
  • LoL: Builds on LongLive to study and mitigate sink-collapse for ultra-long AR streaming video generation.
  • TriAttention: Integrates trigonometric KV-cache compression into LongLive's causal inference pipeline, reducing KV memory inside LongLive's local-attention window.
  • StreamEdit: Provides a LongLive_StreamEdit implementation for training-free streaming video editing built on the LongLive 1.0 codebase.
  • LongLive-RAG: A general retrieval-augmented framework for long video generation.
  • Streaming Autoregressive Video Generation via Diagonal Distillation: Builds on the LongLive codebase and supports direct initialization from LongLive-1.3B checkpoints for streaming AR video distillation.
  • Forcing-KV: Adds hybrid KV-cache compression to LongLive, including LongLive inference and interactive-generation scripts.
  • Dummy Forcing: Unifies Self-Forcing, LongLive, and Causal-Forcing pipelines with LongLive inference, VBench, and interactive-generation configs.
  • MemRoPE: Uses LongLive as a supported base model for training-free infinite video generation with evolving memory tokens.
  • Astrolabe: Supports LongLive as a distilled autoregressive video backbone with LongLive-specific RL configs and LoRA initialization.
  • OPSD-V: Post-trains LongLive with cache-aware on-policy self-distillation to improve long-horizon visual quality and motion dynamics while preserving its few-step autoregressive inference pipeline.

License

Released under the Apache License 2.0. Each directory also carries its own copy of the license and, where applicable, its own third-party notices.

Citation

Please consider citing our work if you find it useful:

@misc{yang2026longliveplug,
  title  = {LongLive-Plug: Once-for-All Distillation for Video Generation},
  author = {Shuai Yang and Luozhou Wang and Wei Huang and ZhiFei Chen and
            Bohan Zhang and Xiao Fu and Qianli Ma and Chen-Hsuan Lin and
            Weian Mao and Bryan Chu and Song Han and Yukang Chen},
  year   = {2026}
}
@article{longlive_2.0,
  title={LongLive2.0: An NVFP4 Parallel Infrastructure for Long Video Generation},
  author={Chen, Yukang and Wang, Luozhou and Huang, Wei and Yang, Shuai and Zhang, Bohan and Xiao, Yicheng and Chu, Ruihang and Mao, Weian and Hu, Qixin and Liu, Shaoteng and Zhao, Yuyang and Mao, Huizi and Chen, Ying-Cong and Xie, Enze and Qi, Xiaojuan and Han, Song},
  journal={arXiv preprint arXiv: 2605.18739},
  year={2026}
}
@inproceedings{longlive,
    title={Longlive: Real-time interactive long video generation},
    author={Yang, Shuai and Huang, Wei and Chu, Ruihang and Xiao, Yicheng and Zhao, Yuyang and Wang, Xianbang and Li, Muyang and Xie, Enze and Chen, Yingcong and Lu, Yao and others},
    booktitle={ICLR},
    year={2026},
}

Related project:

@article{longlive_rag,
  title         = {LongLive-RAG: A General Retrieval-Augmented Framework for Long Video Generation},
  author        = {Hu, Qixin and Yang, Shuai and Huang, Wei and Han, Song and Chen, Yukang},
  journal       = {arXiv preprint arXiv:2606.02553},
  year          = {2026}
}

Acknowledgement

  • Self-Forcing: the AR training codebase and formulation we build upon.
  • Wan2.1: the video diffusion backbone used in LongLive 1.0 and LongLive-Plug.
  • Wan2.2: the base video diffusion model components used in this release.
  • MiniMax-H3: the audio-video generation backbone used in LongLive-Plug.

Releases

Packages

Used by

Contributors

Languages