September 14, 2026 · By YasKad
Lightricks/LTX-2

LTX-2: the open-weight audio-video model that runs on your own GPU

Lightricks/LTX-2 · 9,525★ · 1,503 forks

The essentials on Lightricks/LTX-2: the official Python inference package and LoRA training toolkit for the LTX-2 audio-video generative model, paired with the model on Hugging Face, a family of pipelines, and a trainer.

A striking dark-mode cyberpunk hero image representing an open-weight audiovisual generation model: a massive translucent neural engine suspended in a black data cathedral, composed of two intertwined luminous streams, a cyan video stream made of moving cinematic frames, depth maps and motion vectors, and a magenta audio stream made of waveform crystals, speaker cones and spectral bars, the streams merge into a synchronized scene where a futuristic character speaks while sound waves ripple through the environment, producing lip-synced dialogue and ambient effects, surrounding GPU racks glow with VRAM modules, cooling fans and local inference indicators, holographic pipeline nodes, a Python terminal and a license vault float nearby, include a subtle abstract monogram suggesting an open-source audiovisual model, dark mode, cyberpunk/tech aesthetic, neon accents, ultra-detailed, 8K resolution, volumetric lighting, dramatic depth of field, cinematic composition

What LTX-2 is

LTX-2 is a DiT (Diffusion Transformer)-based foundation model for audio and video generation. Unlike “silent” video models, LTX-2 generates video and audio synchronized within a single model: the soundtrack follows characters, environment, and effects — it isn’t generated separately. The README describes it as “the first DiT-based audio-video foundation model” that brings together in one model the core capabilities of modern video generation: synchronized audio and video, high fidelity, several performance modes, production-ready outputs, API access, and open weights.

The repository documented here (Lightricks/LTX-2) isn’t the model itself but its official software package: it’s a monorepo with three subpackages — ltx-core (model implementation and inference stack), ltx-pipelines (the high-level generation routes), and ltx-trainer (LoRA training and fine-tuning, full fine-tuning, and IC-LoRA) — plus an agent skill for assisted training. The model weights live on Hugging Face (Lightricks/LTX-2 and, in the currently recommended version, Lightricks/LTX-2.5).

The architecture, per the technical paper, is an asymmetric dual-stream transformer: a 14-billion-parameter video branch and a 5-billion-parameter audio branch, coupled via bidirectional cross-attention layers with temporal positional embeddings and a cross-modality AdaLN for shared timestep conditioning. Assigning more capacity to video than audio is a deliberate design decision. The model uses a multilingual text encoder (Gemma fine-tuned for LTX) and introduces a modality-CFG mechanism (modality-aware classifier-free guidance) to improve audio-video alignment and controllability.

An ultra-detailed architectural cutaway of an asymmetric dual-stream Diffusion Transformer, dark mode cyberpunk style, the left tower is a video branch represented by a 14 billion-parameter scale, dense cyan frame lattices, depth maps, normal maps and motion vector ribbons, the right tower is an audio branch represented by a 5 billion-parameter scale, magenta waveform crystals, mel-spectrogram shards and speaker cones, bidirectional cross-attention bridges connect the towers with glowing temporal positional embedding rings and a shared AdaLN time-condition dial, floating labels are abstract, with no readable text, dark mode, cyberpunk/tech aesthetic, neon accents, ultra-detailed, 8K resolution

Origin

The repository was created on January 3, 2026 by Lightricks, the Israeli company behind the LTX family. The LTX-2 model was formally introduced in the paper “LTX-2: Efficient Joint Audio-Visual Foundation Model” (arXiv:2601.03233, published January 6, 2026), with Yoav HaCohen as first author and Zeev Farbman — Lightricks’ co-founder and CEO — as last.

The most revealing launch anecdote is in the CEO’s Reddit AMA. On January 8, 2026, Farbman replied as u/ltx_model in the r/StableDiffusion thread 1q7dzq2: “I’m the co-founder and CEO of Lightricks. We just open-sourced LTX-2, a production audio-video model. AMA.” He clarified that the full release included weights, code, a trainer, benchmarks, LoRAs, and documentation, and that the model “runs locally on consumer GPUs and powers real products at Lightricks.” The thread reached 1,554 votes and 460 comments, and the CEO closed it acknowledging that “the volume of questions exceeded all expectations.”

The background story is Lightricks’ thesis: generative models “are evolving into full render engines,” with inputs like depth, normals, and motion vectors, and output feeding into compositing pipelines, VFX, animation tools, and game engines. Farbman argues that “static APIs can’t cover it” and that much of it needs to run at the edge. Hence his repeated line: “open weights aren’t a luxury, they’re the only path that works,” and Lightricks monetizes through licensing and a revenue share once someone building on top crosses the $10M annual revenue threshold.

Philosophy and principles

Open by architecture, not by charity: Farbman frames it explicitly — “we don’t think of open weights as charity or goodwill; it’s the heart of how we believe render engines should be built.” The goal is for the model to become integrated infrastructure in pipelines, not an API you consume.

Local and reproducible use: “open releases of multimodal models are rare, and when they happen they’re usually hard to run or reproduce. We built LTX-2 so you can actually use it: it runs locally on consumer GPUs.”

A dark developer studio where a consumer GPU runs local inference for an audiovisual model, the GPU has a glowing heatsink, liquid cooling loops and neon VRAM indicators, on the desk a Python terminal shows abstract code, a dataset folder, a pipeline graph and a training log, a monitor displays a generated cinematic video with synchronized audio waveform, real-time preview, and edge-rendering indicators instead of cloud dependency, the scene emphasizes reproducible local execution, privacy, and low latency, dark mode, cyberpunk/tech aesthetic, neon accents, ultra-detailed, 8K resolution

One model, composed capabilities: audio and video together, not two models bolted together; synchronization is intrinsic to the design. User-trainable: the base (dev) model is fully trainable, and the trainer is published so anyone can bring their own dataset and fine-tune for their case; for many adjustments, fine-tuning for motion, style, or likeness “can take less than an hour.” Efficiency as a priority: the distilled transformer (distilled) runs in very few steps, and the pipeline family explicitly separates the fast mode from the production-quality mode (DFR).

How it works

The repository is organized into three packages, each with its own README and documentation: ltx-core (model implementation, inference stack, and utilities), ltx-pipelines (the high-level generation routes: text-to-video, image-to-video, and other modes), and ltx-trainer (training and fine-tuning tools — LoRA, full fine-tuning, and IC-LoRA).

The available pipelines are the practical core:

  • DistilledPipeline — the fast starting point: faster text/image-to-video (8 predefined sigmas: 8 steps stage 1, 4 steps stage 2).
  • DFRPipeline — production quality (DFR, Diffusion Fidelity Rendering): uses the same distilled transformer plus a detailing IC-LoRA; generates interior keyframes and a spatial detailing pass, optionally with 2×/4× fps.
  • TI2VidTwoStagesPipeline / TI2VidTwoStagesHQPipeline — guided two-stage text/image-to-video with CFG/STG and 2× upscaling.
  • TI2VidOneStagePipeline — single-stage generation for rapid prototyping.
  • ICLoraPipeline — video-to-video and image-to-video transformation.
  • KeyframeInterpolationPipeline — interpolation between keyframe images.
  • A2VidPipelineTwoStage — video generation from audio, conditioned on an input audio file.
  • RetakePipeline — regenerates a specific temporal region of an existing video.
  • HDRICLoraPipeline — video-to-video with HDR output via IC-LoRA.
  • DubItPipeline — dialogue rewriting while preserving the speaker’s identity and lip movements.
  • Native HDR/EXR — the standard pipelines accept EXR frames with --hdr {SRGB_LINEAR,ACESCG,ACESCCT} and write half-float EXR frames plus a BT.2020/HLG master.

A holographic control room for audiovisual generation pipelines, dark cyberpunk interface, multiple glowing route graphs branch from a central model core: a fast distilled route with minimal steps, a high-fidelity production route with detail enhancement and spatial upscaling, two-stage text-to-video and image-to-video routes, one-stage prototyping, video-to-video and image-to-video transformations, keyframe interpolation, temporal retake regeneration, audio-conditioned video generation, dialogue rewriting, and HDR/EXR output paths, each node is represented by abstract icons: film frames, waveforms, sliders, upscaling arrows, keyframe diamonds, retake loops, speaker masks, and color wheels, dark mode, cyberpunk/tech aesthetic, neon accents, ultra-detailed, 8K resolution

The trainer supports a wide battery of conditional modes: text-to-video, text-to-audio, image-to-video, video extension, audio extension, video and audio inpainting, video outpainting, IC-LoRA for video/audio/joint audio-video, audio-to-video, and video-to-audio. There’s also an agent skill (.claude/skills/train-model) that guides an end-to-end training run: explores data and hardware, chooses the mode, prepares/preprocesses the dataset, launches training, and monitors it.

A futuristic model training laboratory, dark mode cyberpunk aesthetic, a LoRA fine-tuning console displays adapter chips, learning curves, validation metrics and dataset preparation stages, shelves hold video clips, audio reels, reference portraits, voice samples and motion capture rigs, a digital agent assistant guides the workflow through stages: explore hardware, choose training mode, preprocess dataset, launch training, monitor convergence, emphasize fast fine-tuning for movement, style, voice likeness and appearance, with a subtle timer suggesting under one hour for many adjustments, dark mode, cyberpunk/tech aesthetic, neon accents, ultra-detailed, 8K resolution

The ecosystem

Lightricks repositories (same organization)

RepositoryStarsRole
Lightricks/LTX-Video10,950Official LTX-Video repository (the earlier 2024 real-time video model).
Lightricks/LTX-29,411This repository: inference package and trainer for LTX-2.
Lightricks/ComfyUI-LTXVideo4,129LTX-Video/LTX-2 support for ComfyUI.
Lightricks/LTX-Desktop1,988Open-source desktop app for generating videos with LTX models.
Lightricks/LTX-Video-Trainer468Community trainer for the LTX Video model.
Lightricks/LTX-Video-Q8-Kernels82Q8 kernels for LTX-Video.

Models on Hugging Face

Lightricks/LTX-2 — the original model (19B: 14B video + 5B audio). As checked: 333,852 downloads and 1,779 likes. image-to-video pipeline, integrated into diffusers as LTX2Pipeline. Lightricks/LTX-2.5 — the version currently recommended by the README (22B transformer + Gemma 4 12B). As checked: 1,559,653 downloads and 3,821 likes. Created July 23, 2026. Lightricks/LTX-2.5-22b-IC-LoRA-Pixel-Spatial-Upscaler — the detailing IC-LoRA required by the DFR pipeline. Weights are published “one file per component,” so only the pieces each pipeline needs get downloaded.

ComfyUI integration

The README and model card point to Lightricks/ComfyUI-LTXVideo and recommend using the LTXVideo nodes bundled in the ComfyUI Manager. There’s also a community project, LTXMac (a native Mac app for text-to-video generation, presented in a Show HN in January 2026), built on the LTX family.

An open-weight model vault in a dark data center, cyberpunk tech aesthetic, transparent containers hold component shards: transformer weights, multilingual text encoder, video VAE, audio VAE, spatial upscalers, temporal upscalers, and detail enhancement adapters, download streams flow into local nodes, diffusers-style integration plugs, ComfyUI-style node graphs, a desktop application window, and an API gateway, a license scroll glows with a revenue threshold marker, symbolizing open architecture and commercial licensing, no readable text, only abstract glyphs and icons, dark mode, cyberpunk/tech aesthetic, neon accents, ultra-detailed, 8K resolution

Official / semi-official status

LTX-2 has a clear reputation as a de facto standard in the open-source video generation ecosystem: it’s integrated as LTX2Pipeline in the official diffusers library, with dedicated docs; LTXVideo nodes are available from the ComfyUI Manager; Lightricks offers LTX-Studio (online demo) and a paid API, and the CEO states the model “powers real products at Lightricks”; and the weights are on Hugging Face under the LTX Community License (not an OSI license; entities with annual revenue of at least $10,000,000 move to a branch of the license requiring commercial terms).

In practice, this means LTX-2 functions as a reference among open, locally-runnable audio-video models: it’s in diffusers, in ComfyUI, in a desktop app, and has a commercial API, but there’s no “formal standard designation” from any external body in the sources consulted.

Quick-start guide

Installation and first run

Prerequisites: Python ≥ 3.12, CUDA (>12.7 recommended), PyTorch ~= 2.7, and a GPU with plenty of VRAM.

git clone https://github.com/Lightricks/LTX-2.git
cd LTX-2

# The `natten` extra is the fastest backend for the video VAE; Linux+CUDA only.
uv sync --extra natten

# Download the model (LTX-2.5 repo, ~66 GiB)
hf auth login
hf download Lightricks/LTX-2.5 \
    diffusion_models/ltx-2.5-22b-distilled-transformer-bf16.safetensors \
    text_encoders/gemma4-12b-with-proj-ltx-2.5-bf16.safetensors \
    vae/ltx-2.5-video-vae-bf16.safetensors \
    vae/ltx-2.5-audio-vae-bf16.safetensors \
    latent_upscale_models/ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors \
    --local-dir models/ltx-2.5

On first run, if a 401/403 from Hugging Face appears, accept the model’s terms and log in with a read token.

Common workflows

1. Fast text-to-video (DistilledPipeline). To generate a starter clip:

uv run python -m ltx_pipelines.distilled \
    --transformer-path   models/ltx-2.5/diffusion_models/ltx-2.5-22b-distilled-transformer-bf16.safetensors \
    --text-encoder-path  models/ltx-2.5/text_encoders/gemma4-12b-with-proj-ltx-2.5-bf16.safetensors \
    --video-vae-path     models/ltx-2.5/vae/ltx-2.5-video-vae-bf16.safetensors \
    --audio-vae-path     models/ltx-2.5/vae/ltx-2.5-audio-vae-bf16.safetensors \
    --spatial-upsampler-path models/ltx-2.5/latent_upscale_models/ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors \
    --num-frames 121 --seed 42 --output-path output.mp4 \
    --prompt "..."

The result lands in output.mp4. Default resolution is 1024×1536 at 24 fps.

2. Production quality (DFRPipeline). Reuse the same download and add the detailing IC-LoRA, running ltx_pipelines.dfr_pipeline with the extra --detailing-lora flag. Output in output_dfr.mp4; 4K UHD is --width 3840 --height 2176.

3. Video from audio (A2Vid). Use A2VidPipelineTwoStage to generate a video conditioned on an input audio file.

4. Dub/rewrite dialogue (DubIt). DubItPipeline rewrites dialogue while preserving the speaker’s identity and lips.

A close-up synchronized audiovisual generation scene, dark cyberpunk style, a photorealistic character's face is generated alongside an audio waveform, with lip movements, breath, and dialogue peaks aligned on a temporal grid, holographic tracks show speaker identity, voice timbre, ambient sound, foley effects and music layers, a dialogue rewriting interface preserves lip sync and character identity while changing speech content, neon cyan and magenta accents trace the alignment between audio events and video frames, dark mode, cyberpunk/tech aesthetic, neon accents, ultra-detailed, 8K resolution

Essential configuration

  • --num-frames: divisible by 8 plus 1; can be omitted with --auto-duration.
  • --quantization: fp8-cast, fp8-scaled-mm (Hopper+), nvfp4-cast/nvfp4-prequant (Blackwell).
  • --offload {cpu,disk}: for GPUs with limited VRAM.
  • --hdr {SRGB_LINEAR,ACESCG,ACESCCT}: color-space selection for EXR conditioning and HDR encoding.
  • Prompting guide: the README recommends chronological, literal prompts up to 200 words, “like a cinematographer describing a shot list.”

Common pitfalls and fixes

  • 401/403 downloading from Hugging Face: accept the model’s terms and log in with a read token.
  • Resolution and frames: width/height divisible by 32 and frame count by 8+1; 4K is 3840×2176, not 3840×2160.
  • natten is Linux+CUDA only: on Windows/macOS it’s skipped, falling back to Triton/eager.
  • Don’t substitute the Gemma: the text encoder must be the LTX-fine-tuned Gemma 4; Google’s “stock” Gemma 4 isn’t a substitute.
  • OOM decoding the VAE: pipe.vae.enable_tiling() is recommended on the diffusers path.
  • Weights aren’t interchangeable between LTX-2.3 and LTX-2.5: a LoRA only works with the model it was trained on.
  • The download is ~66 GiB: it’s downloaded per component so you don’t have to pull everything.

Integrations and migration

diffusers: LTX-2 is available as LTX2Pipeline and LTX2LatentUpsamplePipeline; direct migration with LTX2Pipeline.from_pretrained("Lightricks/LTX-2", torch_dtype=torch.bfloat16). ComfyUI: via Lightricks/ComfyUI-LTXVideo and the LTXVideo nodes from the Manager. LTX-Studio/API: for cloud use without managing hardware. HDR/EXR toward VFX: linear EXR outputs and the BT.2020/HLG master ease migration from traditional compositing workflows.

Current metrics

Measured: September 14, 2026, GitHub API.

MetricValue
Stars9,411
Forks1,489
Subscribers (subscribers_count)99
Open issues41
Commits on main50 (approximate)
Main languagePython
LicenseLTX Community License (proprietary, not OSI)
CreatedJanuary 3, 2026
Last activity (push)August 26, 2026
Latest releasev1.3.0, August 26, 2026

Top contributors: michaellightricks (26), github-actions[bot] (10), AlexeyKravtsov1987 (1), and Bmantl (1). The GitHub API uses open_issues_count, which may include open pull requests. The general response’s watchers_count mirrors the stars, so the subscribers_count field is reported separately as real subscribers.

How to contribute

The trainer’s README documents an open, simple process: share results (interesting LoRAs or good results with the community), report issues (open a GitHub issue for a bug or suggestion), submit PRs (bug fixes or general improvements), and request features (via GitHub issues). The community coordinates mainly on the official Discord, where the dev team provides support and results are shared.

How the community received it

In the CEO’s AMA (r/StableDiffusion 1q7dzq2, January 8, 2026, 1,554 votes and 460 comments), Farbman’s top-voted reply was the justification for open weights: “models are evolving into full render engines… open weights aren’t a luxury, they’re the only path that works. We monetize through licensing and a revenue share when people build successful products on top (we draw the line at $10M in revenue).” He also announced an incremental 2.1 release and an architectural 2.5 leap.

u/Neex, who identifies as Niko from Corridor Digital, joined the thread and said “you’re nailing this response”; then u/That_Buddy_2928 added that Corridor’s Bullet Time remaster video “was instrumental in convincing some of my more skeptical friends of AI’s validity in the pipeline.” u/SvenVargHimmel raised whether the model can be forced to generate a single image and whether it can export normal or depth maps as video — a sign users were already treating it as a render engine.

On Hacker News, the most relevant Show HN is 47214472 “Show HN: Audio-to-Video with LTX-2” (by runshouse, March 2, 2026), with 23 points and 2 comments.

In the r/StableDiffusion subreddit in August 2026, LTX 2.5 dominates local-use threads: for example, “MiniMax H3 Model Copied LTX 2.5’s Best Feature… And It’s CRAZY Fast!” (159 votes, 96 comments) and “I Found a way to reduce LTX 2.5’s horrible smearing. Custom node + Workflow” (158 votes, 24 comments). These threads give two readings at once: enthusiasm for quality and a concrete criticism — motion “smearing” — that the community tries to compensate for with custom nodes.

The model card states explicit limitations: it isn’t meant for factual information, may amplify biases, may fail to follow the prompt (heavily dependent on prompt style), and audio without voice may be lower quality.

Comparison with similar projects

ProjectVerifiable overlapVerifiable difference
Lightricks/LTX-2.5 (same author)Same family, same infrastructure; the README treats it as the starting point.22B transformer + Gemma 4, new latent space; more downloads (1.56M) than LTX-2 (333K). It’s the evolution, not a competitor.
Runway Gen-4Production video generator, listed in Lightricks’ own guide.Closed and cloud-based; LTX-2 offers open weights and local execution.
Google Veo 3.1Video generator with audio, listed in the same guide.Closed, from Google; LTX-2 is open and local.
Kling 3.0High-quality video, listed in the own guide.Closed; LTX-2 integrates into diffusers/ComfyUI.
Pika 2.2Creative video generator, listed in the own guide.Closed, API consumption; LTX-2 is local and trainable.
MiniMax H3Open-weight generator, for local use, repeatedly compared to LTX 2.5 on r/StableDiffusion.Lightricks describes it as the one that “copied LTX 2.5’s best feature”; it’s a direct rival in the open field.

The most useful comparison isn’t by popularity: LTX-2 stands out when you want an open, local audio-video model that integrates into diffusers/ComfyUI and is trainable; a closed cloud generator may be preferable when you lack your own hardware.

Use cases

  • Creators and production houses wanting local, controlled audio-video: the synchronized video+audio pair (with A2VidPipelineTwoStage and DubItPipeline) serves anyone making dialogue content without a separate dubbing or effects step.
  • VFX and post-production teams: HDR/EXR outputs and RetakePipeline are designed to feed into color grading and compositing pipelines without a lossy intermediary.

A production-grade HDR/EXR video pipeline in a dark post-production suite, cyberpunk tech aesthetic, holographic EXR frames display linear color data, BT.2020 color space, HLG masters, tonemapping curves, and half-float luminance histograms, side viewports show depth maps, normal maps, motion vectors, and compositing layers, dark mode, cyberpunk/tech aesthetic, neon accents, ultra-detailed, 8K resolution

  • Researchers and ML teams fine-tuning models: ltx-trainer and the train-model agent skill let you bring your own dataset and fine-tune style, motion, or likeness, sometimes “in under an hour.”
  • Developers integrating generative models into software: diffusers and ComfyUI integration lets you embed LTX-2 into existing apps and flows.
  • People with consumer GPUs: the distilled variant and the --quantization/--offload options are aimed at making the model run locally on consumer GPUs.

Resources


Note: this article combines the repository README, the Hugging Face model card, the arXiv:2601.03233 paper, the CEO’s Reddit AMA, release notes (v1.2.0, v1.3.0), the GitHub API, and Hacker News and Reddit results consulted on September 14, 2026. Star, download, and vote figures change over time.

Comments