LTX-2: the open-weight audio-video model that runs on your own GPU
Lightricks/LTX-2 · 9,525★ · 1,503 forks
The essentials on Lightricks/LTX-2: the official Python inference package and LoRA training toolkit for the LTX-2 audio-video generative model, paired with the model on Hugging Face, a family of pipelines, and a trainer.

What LTX-2 is
LTX-2 is a DiT (Diffusion Transformer)-based foundation model for audio and video generation. Unlike “silent” video models, LTX-2 generates video and audio synchronized within a single model: the soundtrack follows characters, environment, and effects — it isn’t generated separately. The README describes it as “the first DiT-based audio-video foundation model” that brings together in one model the core capabilities of modern video generation: synchronized audio and video, high fidelity, several performance modes, production-ready outputs, API access, and open weights.
The repository documented here (Lightricks/LTX-2) isn’t the model itself but its official software package: it’s a monorepo with three subpackages — ltx-core (model implementation and inference stack), ltx-pipelines (the high-level generation routes), and ltx-trainer (LoRA training and fine-tuning, full fine-tuning, and IC-LoRA) — plus an agent skill for assisted training. The model weights live on Hugging Face (Lightricks/LTX-2 and, in the currently recommended version, Lightricks/LTX-2.5).
The architecture, per the technical paper, is an asymmetric dual-stream transformer: a 14-billion-parameter video branch and a 5-billion-parameter audio branch, coupled via bidirectional cross-attention layers with temporal positional embeddings and a cross-modality AdaLN for shared timestep conditioning. Assigning more capacity to video than audio is a deliberate design decision. The model uses a multilingual text encoder (Gemma fine-tuned for LTX) and introduces a modality-CFG mechanism (modality-aware classifier-free guidance) to improve audio-video alignment and controllability.

Origin
The repository was created on January 3, 2026 by Lightricks, the Israeli company behind the LTX family. The LTX-2 model was formally introduced in the paper “LTX-2: Efficient Joint Audio-Visual Foundation Model” (arXiv:2601.03233, published January 6, 2026), with Yoav HaCohen as first author and Zeev Farbman — Lightricks’ co-founder and CEO — as last.
The most revealing launch anecdote is in the CEO’s Reddit AMA. On January 8, 2026, Farbman replied as u/ltx_model in the r/StableDiffusion thread 1q7dzq2: “I’m the co-founder and CEO of Lightricks. We just open-sourced LTX-2, a production audio-video model. AMA.” He clarified that the full release included weights, code, a trainer, benchmarks, LoRAs, and documentation, and that the model “runs locally on consumer GPUs and powers real products at Lightricks.” The thread reached 1,554 votes and 460 comments, and the CEO closed it acknowledging that “the volume of questions exceeded all expectations.”
The background story is Lightricks’ thesis: generative models “are evolving into full render engines,” with inputs like depth, normals, and motion vectors, and output feeding into compositing pipelines, VFX, animation tools, and game engines. Farbman argues that “static APIs can’t cover it” and that much of it needs to run at the edge. Hence his repeated line: “open weights aren’t a luxury, they’re the only path that works,” and Lightricks monetizes through licensing and a revenue share once someone building on top crosses the $10M annual revenue threshold.
Philosophy and principles
Open by architecture, not by charity: Farbman frames it explicitly — “we don’t think of open weights as charity or goodwill; it’s the heart of how we believe render engines should be built.” The goal is for the model to become integrated infrastructure in pipelines, not an API you consume.
Local and reproducible use: “open releases of multimodal models are rare, and when they happen they’re usually hard to run or reproduce. We built LTX-2 so you can actually use it: it runs locally on consumer GPUs.”

One model, composed capabilities: audio and video together, not two models bolted together; synchronization is intrinsic to the design. User-trainable: the base (dev) model is fully trainable, and the trainer is published so anyone can bring their own dataset and fine-tune for their case; for many adjustments, fine-tuning for motion, style, or likeness “can take less than an hour.” Efficiency as a priority: the distilled transformer (distilled) runs in very few steps, and the pipeline family explicitly separates the fast mode from the production-quality mode (DFR).
How it works
The repository is organized into three packages, each with its own README and documentation: ltx-core (model implementation, inference stack, and utilities), ltx-pipelines (the high-level generation routes: text-to-video, image-to-video, and other modes), and ltx-trainer (training and fine-tuning tools — LoRA, full fine-tuning, and IC-LoRA).
The available pipelines are the practical core:
DistilledPipeline— the fast starting point: faster text/image-to-video (8 predefined sigmas: 8 steps stage 1, 4 steps stage 2).DFRPipeline— production quality (DFR, Diffusion Fidelity Rendering): uses the same distilled transformer plus a detailing IC-LoRA; generates interior keyframes and a spatial detailing pass, optionally with 2×/4× fps.TI2VidTwoStagesPipeline/TI2VidTwoStagesHQPipeline— guided two-stage text/image-to-video with CFG/STG and 2× upscaling.TI2VidOneStagePipeline— single-stage generation for rapid prototyping.ICLoraPipeline— video-to-video and image-to-video transformation.KeyframeInterpolationPipeline— interpolation between keyframe images.A2VidPipelineTwoStage— video generation from audio, conditioned on an input audio file.RetakePipeline— regenerates a specific temporal region of an existing video.HDRICLoraPipeline— video-to-video with HDR output via IC-LoRA.DubItPipeline— dialogue rewriting while preserving the speaker’s identity and lip movements.- Native HDR/EXR — the standard pipelines accept EXR frames with
--hdr {SRGB_LINEAR,ACESCG,ACESCCT}and write half-float EXR frames plus a BT.2020/HLG master.

The trainer supports a wide battery of conditional modes: text-to-video, text-to-audio, image-to-video, video extension, audio extension, video and audio inpainting, video outpainting, IC-LoRA for video/audio/joint audio-video, audio-to-video, and video-to-audio. There’s also an agent skill (.claude/skills/train-model) that guides an end-to-end training run: explores data and hardware, chooses the mode, prepares/preprocesses the dataset, launches training, and monitors it.

The ecosystem
Lightricks repositories (same organization)
| Repository | Stars | Role |
|---|---|---|
Lightricks/LTX-Video | 10,950 | Official LTX-Video repository (the earlier 2024 real-time video model). |
Lightricks/LTX-2 | 9,411 | This repository: inference package and trainer for LTX-2. |
Lightricks/ComfyUI-LTXVideo | 4,129 | LTX-Video/LTX-2 support for ComfyUI. |
Lightricks/LTX-Desktop | 1,988 | Open-source desktop app for generating videos with LTX models. |
Lightricks/LTX-Video-Trainer | 468 | Community trainer for the LTX Video model. |
Lightricks/LTX-Video-Q8-Kernels | 82 | Q8 kernels for LTX-Video. |
Models on Hugging Face
Lightricks/LTX-2 — the original model (19B: 14B video + 5B audio). As checked: 333,852 downloads and 1,779 likes. image-to-video pipeline, integrated into diffusers as LTX2Pipeline. Lightricks/LTX-2.5 — the version currently recommended by the README (22B transformer + Gemma 4 12B). As checked: 1,559,653 downloads and 3,821 likes. Created July 23, 2026. Lightricks/LTX-2.5-22b-IC-LoRA-Pixel-Spatial-Upscaler — the detailing IC-LoRA required by the DFR pipeline. Weights are published “one file per component,” so only the pieces each pipeline needs get downloaded.
ComfyUI integration
The README and model card point to Lightricks/ComfyUI-LTXVideo and recommend using the LTXVideo nodes bundled in the ComfyUI Manager. There’s also a community project, LTXMac (a native Mac app for text-to-video generation, presented in a Show HN in January 2026), built on the LTX family.

Official / semi-official status
LTX-2 has a clear reputation as a de facto standard in the open-source video generation ecosystem: it’s integrated as LTX2Pipeline in the official diffusers library, with dedicated docs; LTXVideo nodes are available from the ComfyUI Manager; Lightricks offers LTX-Studio (online demo) and a paid API, and the CEO states the model “powers real products at Lightricks”; and the weights are on Hugging Face under the LTX Community License (not an OSI license; entities with annual revenue of at least $10,000,000 move to a branch of the license requiring commercial terms).
In practice, this means LTX-2 functions as a reference among open, locally-runnable audio-video models: it’s in diffusers, in ComfyUI, in a desktop app, and has a commercial API, but there’s no “formal standard designation” from any external body in the sources consulted.
Quick-start guide
Installation and first run
Prerequisites: Python ≥ 3.12, CUDA (>12.7 recommended), PyTorch ~= 2.7, and a GPU with plenty of VRAM.
git clone https://github.com/Lightricks/LTX-2.git
cd LTX-2
# The `natten` extra is the fastest backend for the video VAE; Linux+CUDA only.
uv sync --extra natten
# Download the model (LTX-2.5 repo, ~66 GiB)
hf auth login
hf download Lightricks/LTX-2.5 \
diffusion_models/ltx-2.5-22b-distilled-transformer-bf16.safetensors \
text_encoders/gemma4-12b-with-proj-ltx-2.5-bf16.safetensors \
vae/ltx-2.5-video-vae-bf16.safetensors \
vae/ltx-2.5-audio-vae-bf16.safetensors \
latent_upscale_models/ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors \
--local-dir models/ltx-2.5
On first run, if a 401/403 from Hugging Face appears, accept the model’s terms and log in with a read token.
Common workflows
1. Fast text-to-video (DistilledPipeline). To generate a starter clip:
uv run python -m ltx_pipelines.distilled \
--transformer-path models/ltx-2.5/diffusion_models/ltx-2.5-22b-distilled-transformer-bf16.safetensors \
--text-encoder-path models/ltx-2.5/text_encoders/gemma4-12b-with-proj-ltx-2.5-bf16.safetensors \
--video-vae-path models/ltx-2.5/vae/ltx-2.5-video-vae-bf16.safetensors \
--audio-vae-path models/ltx-2.5/vae/ltx-2.5-audio-vae-bf16.safetensors \
--spatial-upsampler-path models/ltx-2.5/latent_upscale_models/ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors \
--num-frames 121 --seed 42 --output-path output.mp4 \
--prompt "..."
The result lands in output.mp4. Default resolution is 1024×1536 at 24 fps.
2. Production quality (DFRPipeline). Reuse the same download and add the detailing IC-LoRA, running ltx_pipelines.dfr_pipeline with the extra --detailing-lora flag. Output in output_dfr.mp4; 4K UHD is --width 3840 --height 2176.
3. Video from audio (A2Vid). Use A2VidPipelineTwoStage to generate a video conditioned on an input audio file.
4. Dub/rewrite dialogue (DubIt). DubItPipeline rewrites dialogue while preserving the speaker’s identity and lips.

Essential configuration
--num-frames: divisible by 8 plus 1; can be omitted with--auto-duration.--quantization:fp8-cast,fp8-scaled-mm(Hopper+),nvfp4-cast/nvfp4-prequant(Blackwell).--offload {cpu,disk}: for GPUs with limited VRAM.--hdr {SRGB_LINEAR,ACESCG,ACESCCT}: color-space selection for EXR conditioning and HDR encoding.- Prompting guide: the README recommends chronological, literal prompts up to 200 words, “like a cinematographer describing a shot list.”
Common pitfalls and fixes
- 401/403 downloading from Hugging Face: accept the model’s terms and log in with a read token.
- Resolution and frames: width/height divisible by 32 and frame count by 8+1; 4K is 3840×2176, not 3840×2160.
nattenis Linux+CUDA only: on Windows/macOS it’s skipped, falling back to Triton/eager.- Don’t substitute the Gemma: the text encoder must be the LTX-fine-tuned Gemma 4; Google’s “stock” Gemma 4 isn’t a substitute.
- OOM decoding the VAE:
pipe.vae.enable_tiling()is recommended on the diffusers path. - Weights aren’t interchangeable between LTX-2.3 and LTX-2.5: a LoRA only works with the model it was trained on.
- The download is ~66 GiB: it’s downloaded per component so you don’t have to pull everything.
Integrations and migration
diffusers: LTX-2 is available as LTX2Pipeline and LTX2LatentUpsamplePipeline; direct migration with LTX2Pipeline.from_pretrained("Lightricks/LTX-2", torch_dtype=torch.bfloat16). ComfyUI: via Lightricks/ComfyUI-LTXVideo and the LTXVideo nodes from the Manager. LTX-Studio/API: for cloud use without managing hardware. HDR/EXR toward VFX: linear EXR outputs and the BT.2020/HLG master ease migration from traditional compositing workflows.
Current metrics
Measured: September 14, 2026, GitHub API.
| Metric | Value |
|---|---|
| Stars | 9,411 |
| Forks | 1,489 |
Subscribers (subscribers_count) | 99 |
| Open issues | 41 |
Commits on main | 50 (approximate) |
| Main language | Python |
| License | LTX Community License (proprietary, not OSI) |
| Created | January 3, 2026 |
| Last activity (push) | August 26, 2026 |
| Latest release | v1.3.0, August 26, 2026 |
Top contributors: michaellightricks (26), github-actions[bot] (10), AlexeyKravtsov1987 (1), and Bmantl (1). The GitHub API uses open_issues_count, which may include open pull requests. The general response’s watchers_count mirrors the stars, so the subscribers_count field is reported separately as real subscribers.
How to contribute
The trainer’s README documents an open, simple process: share results (interesting LoRAs or good results with the community), report issues (open a GitHub issue for a bug or suggestion), submit PRs (bug fixes or general improvements), and request features (via GitHub issues). The community coordinates mainly on the official Discord, where the dev team provides support and results are shared.
How the community received it
In the CEO’s AMA (r/StableDiffusion 1q7dzq2, January 8, 2026, 1,554 votes and 460 comments), Farbman’s top-voted reply was the justification for open weights: “models are evolving into full render engines… open weights aren’t a luxury, they’re the only path that works. We monetize through licensing and a revenue share when people build successful products on top (we draw the line at $10M in revenue).” He also announced an incremental 2.1 release and an architectural 2.5 leap.
u/Neex, who identifies as Niko from Corridor Digital, joined the thread and said “you’re nailing this response”; then u/That_Buddy_2928 added that Corridor’s Bullet Time remaster video “was instrumental in convincing some of my more skeptical friends of AI’s validity in the pipeline.” u/SvenVargHimmel raised whether the model can be forced to generate a single image and whether it can export normal or depth maps as video — a sign users were already treating it as a render engine.
On Hacker News, the most relevant Show HN is 47214472 “Show HN: Audio-to-Video with LTX-2” (by runshouse, March 2, 2026), with 23 points and 2 comments.
In the r/StableDiffusion subreddit in August 2026, LTX 2.5 dominates local-use threads: for example, “MiniMax H3 Model Copied LTX 2.5’s Best Feature… And It’s CRAZY Fast!” (159 votes, 96 comments) and “I Found a way to reduce LTX 2.5’s horrible smearing. Custom node + Workflow” (158 votes, 24 comments). These threads give two readings at once: enthusiasm for quality and a concrete criticism — motion “smearing” — that the community tries to compensate for with custom nodes.
The model card states explicit limitations: it isn’t meant for factual information, may amplify biases, may fail to follow the prompt (heavily dependent on prompt style), and audio without voice may be lower quality.
Comparison with similar projects
| Project | Verifiable overlap | Verifiable difference |
|---|---|---|
Lightricks/LTX-2.5 (same author) | Same family, same infrastructure; the README treats it as the starting point. | 22B transformer + Gemma 4, new latent space; more downloads (1.56M) than LTX-2 (333K). It’s the evolution, not a competitor. |
| Runway Gen-4 | Production video generator, listed in Lightricks’ own guide. | Closed and cloud-based; LTX-2 offers open weights and local execution. |
| Google Veo 3.1 | Video generator with audio, listed in the same guide. | Closed, from Google; LTX-2 is open and local. |
| Kling 3.0 | High-quality video, listed in the own guide. | Closed; LTX-2 integrates into diffusers/ComfyUI. |
| Pika 2.2 | Creative video generator, listed in the own guide. | Closed, API consumption; LTX-2 is local and trainable. |
| MiniMax H3 | Open-weight generator, for local use, repeatedly compared to LTX 2.5 on r/StableDiffusion. | Lightricks describes it as the one that “copied LTX 2.5’s best feature”; it’s a direct rival in the open field. |
The most useful comparison isn’t by popularity: LTX-2 stands out when you want an open, local audio-video model that integrates into diffusers/ComfyUI and is trainable; a closed cloud generator may be preferable when you lack your own hardware.
Use cases
- Creators and production houses wanting local, controlled audio-video: the synchronized video+audio pair (with
A2VidPipelineTwoStageandDubItPipeline) serves anyone making dialogue content without a separate dubbing or effects step. - VFX and post-production teams: HDR/EXR outputs and
RetakePipelineare designed to feed into color grading and compositing pipelines without a lossy intermediary.

- Researchers and ML teams fine-tuning models:
ltx-trainerand thetrain-modelagent skill let you bring your own dataset and fine-tune style, motion, or likeness, sometimes “in under an hour.” - Developers integrating generative models into software: diffusers and ComfyUI integration lets you embed LTX-2 into existing apps and flows.
- People with consumer GPUs: the distilled variant and the
--quantization/--offloadoptions are aimed at making the model run locally on consumer GPUs.
Resources
- Repository: https://github.com/Lightricks/LTX-2
- Hugging Face: https://huggingface.co/Lightricks/LTX-2 and https://huggingface.co/Lightricks/LTX-2.5
- Technical paper: https://arxiv.org/abs/2601.03233
- Community / Discord: https://discord.gg/ltxplatform
- Online demo: https://app.ltx.studio/ltx-2-playground/t2v
- CEO’s AMA: https://www.reddit.com/r/StableDiffusion/comments/1q7dzq2/
- HN: https://news.ycombinator.com/item?id=47214472
- diffusers: https://huggingface.co/docs/diffusers/main/en/api/pipelines/ltx2
- Blog / prompting guide: https://ltx.io/blog
Note: this article combines the repository README, the Hugging Face model card, the arXiv:2601.03233 paper, the CEO’s Reddit AMA, release notes (v1.2.0, v1.3.0), the GitHub API, and Hacker News and Reddit results consulted on September 14, 2026. Star, download, and vote figures change over time.
Comments