September 12, 2026 · By YasKad
pathwaycom/pathway

Pathway: a live data framework for ETL, real-time analytics, and RAG

pathwaycom/pathway · 62,241★ · 1,681 forks

Everything worth knowing about pathwaycom/pathway: a Python ETL framework whose Rust engine runs the same code for both batch and stream processing, with an extension for LLM pipelines and RAG. This report focuses on the repository and its organization, which in 2025-2026 has refocused its public image toward a “post-transformer” AI architecture (BDH).


What Pathway is

Pathway Live Data Framework (abbreviated in code as pw) is a Python ETL framework for stream processing, real-time analytics, LLM pipelines, and RAG. It isn’t a database engine or a service: it’s a Python library the user writes as normal code, executed by a Rust engine.

The core premise, verifiable in the README, is batch/stream unification: the same code works for local development, CI/CD tests, batch jobs, stream replay, and live stream processing. The engine is built on Differential Dataflow and performs incremental computation, so when a piece of data changes, the framework only recalculates what’s affected.

The product includes three building blocks: a connector to sources, transformations (stateful and stateless: joins, windows, orderings), and an LLM extension with utilities (wrappers, parsers, embedders, splitters) and a live in-memory vector index, with documented integrations for LangChain and LlamaIndex.

A striking cyberpunk hero image of a living data framework: a neon dragon hatchling formed from streaming data, microchips, and differential dataflow graphs, emerging from a dark server room, its glowing body is a unified pipeline where batch files and live event streams merge into one luminous river, flowing through Python logic nodes, a Rust execution core, and a live vector index, holographic panels orbit the dragon showing real-time ETL transformations, joins, windows, reducers, and RAG retrieval, deep black and charcoal background, electric cyan, violet, and amber neon accents, volumetric light, ultra-detailed, cinematic, 8K

Origin

The repository was created on November 27, 2022 under the pathwaycom organization. The company is Pathway, based in Poland. The three founders listed on the official site are:

  • Zuzanna Stamirowska, CEO. PhD in complex systems, graduate of École Polytechnique, with network prediction models published in the National Academy of Sciences.
  • Jan Chorowski, CTO. Former researcher at MILA and Google Brain (under Samy Bengio), co-author of work with Geoffrey Hinton and of early attention models, co-author of the Theano library; the site cites over 12,000 citations and an h-index of 24. His GitHub username, janchorowski, is the author of the project’s Show HNs.
  • Adrian Kosowski, CSO. PhD in algorithms, former Inria, over 100 papers, h-index of 29; co-author of the BDH paper.

The organization’s verifiable timeline, per its own blog and press coverage: a $10 million round (TechCrunch, November 29, 2024); recognition as a “Emerging Visionary” in GenAI from Gartner (November 2024); a declared partnership with AWS and NVIDIA to deliver adaptive AI systems (businesswire, December 1, 2025); and a $500 million valuation announced in early August 2025.

The most notable narrative color is an identity shift: the repository is a data/ETL/RAG framework, but pathway.com’s homepage today presents itself as “Intelligence with no ceiling,” a frontier AI lab developing BDH (Dragon Hatchling), a brain-inspired “post-transformer” recurrent architecture. The paper, “The Dragon Hatchling: The Missing Link between the Transformer and Models of the Brain” (arXiv 2509.26507, 2025), is signed by Kosowski, Uznański, Chorowski, Stamirowska, and Bartoszkiewicz. The blog claims BDH was the second most popular AI paper of 2025. This dual focus — live data on one side, architecture research on the other — is the project’s real tension: the ETL framework remains the installable, code-based product, while BDH is the research bet that captures press attention (WSJ, Forbes, MIT Technology Review, Live Science, Semafor).

Philosophy and principles

The README and documentation surface these principles, all verifiable in the official text:

  • One code for batch and stream: the same pipeline definition runs local, in tests, in batch, in replay, and live.
  • Incremental computation: the engine doesn’t recompute everything when data arrives; it updates only what’s derived (built on Differential Dataflow).
  • Python for logic, Rust for execution: the user writes Python (and can call any Python library), but the Rust engine enables multithreading, multiprocessing, and distributed computation, dodging Python’s single-thread limit.
  • Persistence and consistency as a service: the framework saves state to restart after an update or failure, and manages time, updating results as late or out-of-order data arrives.
  • “Live” RAG: the vector index is built and maintained as documents change, so RAG answers reflect the source’s current state.

A dark-mode cyberpunk visualization of one unified pipeline across five environments: a laptop, CI/CD shield, batch warehouse, replay tape, and live stream beacon, all connected by the same glowing code graph, identical neon data packets travel through each environment without changing shape, emphasizing batch and stream unification, Python-like abstract script glyphs and Rust engine cores, holographic branches, deep charcoal background, cyan and magenta neon, ultra-detailed, 8K

How it works

The basic documented pattern (from the README example) is:

  1. pw.io.<source>.read(...) connects to the source and returns a table (with an optional pw.Schema typing the columns).
  2. Transformations are chained onto the table: .filter(...), .reduce(...) with pw.reducers.*, joins, windows, or any Python function.
  3. pw.io.<destination>.write(table, ...) loads the result to an external system.
  4. pw.run() starts the computation.

The project runs as a normal Python script: python main.py. The framework spins up a monitoring panel (per-connector messages, latency, and logs). Documented startup alternatives: pathway spawn python main.py and, for multithreading, pathway spawn --threads 3 python main.py.

An incremental differential dataflow image: a single small data delta enters a dense neural graph, causing only a precise chain of affected nodes to pulse and recalculate while the rest remains dim, ripple effects, glowing dependency edges, before-and-after state layers, time-shifted snapshots, dark cyberpunk lab, electric blue and amber accents, ultra-detailed, 8K

The engine keeps the pipeline in memory. The free version gives an “at least once” consistency guarantee; the enterprise version gives “exactly once.” Release v0.32.1 (August 1, 2026) added, among other things, the write connectors pw.io.chroma.write and pw.io.qdrant.write, which keep a vector database collection synced with the table.

A split composition of Python logic and Rust execution: a glowing Python code surface with flexible function blocks feeds into a hardened Rust engine core with multithreaded gears, distributed worker nodes, and parallel memory lanes, a broken cage of light symbolizes escaping the single-thread limit, high-performance multithreading, dark mode, neon cyan, orange, and violet accents, ultra-detailed, 8K

The LLM extension (called the LLM xpack in the docs) adds LLM-service wrappers, parsers, embedders, splitters, a live in-memory vector index, and LangChain/LlamaIndex integrations. The site’s listed “Featured collaborations” include Databento, LangChain, LlamaIndex, MinIO, PaddleOCR, and Redpanda.

A live RAG vector index scene: documents, PDFs, chat logs, and database records continuously stream into a pulsing in-memory vector index shaped like a crystal lattice, a user query beam retrieves the freshest embeddings and updates answers in real time, holographic retrieval paths, similarity clusters, and abstract integration nodes for LangChain and LlamaIndex, dark cyberpunk background, neon teal and pink accents, ultra-detailed, 8K

The ecosystem

Repositories in the pathwaycom organization

  • pathwaycom/llm-app: deploy-ready application templates (RAG, enterprise search, AI pipelines) built on the framework; 58,937 stars and 1,479 forks. It’s the largest sibling and where the template catalog lives: Question-Answering RAG App, Live Document Indexing, Multimodal RAG with GPT-4o, Unstructured-to-SQL, Adaptive RAG App, Private RAG with Mistral and Ollama, Slides AI Search, and Video RAG with TwelveLabs. That repository’s README states the default vector index uses usearch and hybrid full-text indexes use Tantivy, replacing separate vector database (Pinecone/Weaviate/Qdrant), cache (Redis), and API framework (FastAPI) modules.
  • pathwaycom/arc-task-gen: generates new ARC-AGI-1-style tasks to evaluate models like BDH-CQ; 10,625 stars and 68 forks.
  • pathwaycom/bdh: the official implementation of the Dragon Hatchling architecture (the 2025 paper); 3,545 stars and 252 forks.
  • pathwaycom/pathway-benchmarks: benchmarks against Spark, Flink, and Kafka Streams; 302 stars and 12 forks.
  • pathwaycom/cookiecutter-pathway: scaffolding template for bootstrapping production-ready Pathway projects; 277 stars and 17 forks.
  • pathwaycom/serviette: a RAG tool described as a “Universal RAG tool”; 4 stars.

A modular neon control board representing an LLM extension: glowing panels for wrappers, parsers, embedders, splitters, and live vector store connectors, pipelines connect language model cores to document parsers, embedding matrices, chunk splitters, and vector databases, abstract integration nodes for LangChain and LlamaIndex, dark mode, cyberpunk tech aesthetic, electric cyan, magenta, and amber accents, ultra-detailed, 8K

  • pathway-labs/pathway-examples: framework code examples; 43 stars and 6 forks (belong to a different organization, pathway-labs, not pathwaycom).
  • xyphoes0727/FraudDetection (5 stars): a real-time fraud-detection pipeline built with Pathway, an example of community adoption.

Forks

The main repository’s most-starred forks have few stars (at most 6, e.g. sheikhsajid69/pathway), suggesting the ecosystem grows more through official sibling repos and templates than through independent core forks. No community translation to Spanish or another language of the framework was found in this research.

Official and semi-official status

Pathway is a product with vendor backing and third-party recognition, not a named open standard:

  • Gartner recognized it as an Emerging Visionary in GenAI (November 2024).
  • There’s a declared partnership with AWS and NVIDIA for adaptive and continuous-learning AI systems (businesswire, December 1, 2025).
  • The README lists collaborations and integrations with Databento, LangChain, LlamaIndex, MinIO, PaddleOCR, and Redpanda. LangChain and LlamaIndex document Pathway as an integration (vector store / retriever).
  • The project appears in curated lists like vinta/awesome-python (“pathway” entry, line 497 of that list’s README).

In practice, this means Pathway is an installable, documented option within the LangChain and LlamaIndex ecosystems, and a commercial alternative to mature streaming engines (Flink, Spark, Kafka Streams) with a differentiator in live RAG. There’s no formal “de facto standard” designation in the sources consulted; its position is that of a vertical vendor (live data + RAG + AI) with press and partner traction.

A post-transformer Dragon Hatchling architecture: a brain-inspired recurrent neural core with a small glowing dragon hatchling at its center, surrounded by transformer attention layers and biological synapse lattices, the image shows a bridge between transformer geometry and brain-like recurrence, energy pulses looping through memory cells, dark laboratory, neon violet, cyan, and gold accents, ultra-detailed, 8K

Quick-start guide

Installation and first run

Prerequisites: Python 3.10 or higher and macOS or Linux (the README states other systems should use a virtual machine).

pip install -U pathway

The documented minimal example (sum positive values in real time) is:

import pathway as pw

class InputSchema(pw.Schema):
  value: int

input_table = pw.io.csv.read("./input/", schema=InputSchema)
filtered_table = input_table.filter(input_table.value >= 0)
result_table = filtered_table.reduce(
  sum_value = pw.reducers.sum(filtered_table.value)
)

pw.io.jsonlines.write(result_table, "output.jsonl")
pw.run()

Running python main.py opens a monitoring panel with per-connector messages, latency, and logs. To start with a specific thread count: pathway spawn --threads 3 python main.py. There’s a Google Colab notebook linked from the README and more examples in examples/.

Common workflows

  • Real-time ETL from Kafka: define a pipeline that reads a Kafka topic, transforms it, and writes to a destination; the README links the kafka-etl template and the “Switch from batch to streaming” guide. The result is a pipeline that, when a new message arrives, only updates the affected calculations.
  • Private, self-contained RAG: from pathwaycom/llm-app, the Private RAG with Mistral and Ollama template runs locally; syncing with a folder/Google Drive/SharePoint keeps the index live and answers reflecting the latest documents.
  • Unstructured-to-SQL: the unstructured_to_sql_on_the_fly template reads financial PDFs, structures them into SQL, loads them into PostgreSQL, and answers natural-language queries by translating them to SQL.
  • Running in Docker: for a single-file script, docker run -it --rm --name my-pathway-app -v "$PWD":/app pathwaycom/pathway:latest python my-pathway-app.py.

A real-time monitoring and deployment scene: a dark dashboard with glowing connector messages, latency graphs, log streams, and thread allocation panels, a Docker container and CLI terminal glow beside a multi-process execution grid, a pipeline runs with consistency shields shown as abstract emblems, dark cyberpunk background, neon green, cyan, and violet accents, ultra-detailed, 8K

Essential configuration

  • pw.Schema: types a table’s columns; optional but recommended for validation.
  • pw.io.* connectors: pick the source/destination (csv, jsonlines, kafka, postgresql, gdrive, chroma, qdrant, among others); the first thing most users touch.
  • pathway spawn --threads N: controls local multithreaded parallelism.
  • Persistence backend: the state backend (S3 or local file) is chosen via a code configuration parameter, per the team’s clarification in the Hacker News thread.
  • llm-app templates: for RAG, you start from a template and change the source or index type (vector → hybrid) with “a one-line change,” per its README.

Common pitfalls and fixes

  • Unsupported systems: the README warns Pathway is only available for macOS and Linux; other environments should use a virtual machine.
  • Consistency: the free version gives “at least once”; “exactly once” requires the enterprise version. It’s a license limit, not a bug.
  • RAM requirement in production: the team clarified on the Hacker News thread that the dominant factor in RAM usage is the size of the fed data, especially with an in-memory index. A user’s question about the “Community 8 GB / 4 cores” plan went unanswered with a specific minimum-requirements figure in the sources gathered.
  • Large forks: the CONTRIBUTING doc states a substantial PR without prior maintainer approval may be closed without detailed review; large contributions should open an issue first.

Integrations and migration

  • LangChain / LlamaIndex: Pathway integrates as a vector store / retriever (documented in both ecosystems); llm-app templates act as a retrieval backend for LangChain or LlamaIndex applications.
  • Vector databases: pw.io.chroma.write and pw.io.qdrant.write (v0.32.1) write the table to Chroma/Qdrant while keeping the collection synced.
  • Streaming: Kafka connectors (and, by extension, streaming ecosystems like Redpanda) and an Airbyte connector for over 300 sources.
  • Deployment: Docker (image pathwaycom/pathway:latest), Kubernetes, and clouds; there’s a Render deployment guide in the docs.
  • Migrating from separate vector database/cache/API: the llm-app README frames Pathway as replacing a stack of separate modules (vector DB + Redis + FastAPI) with a single framework with an integrated vector index (usearch) and full-text index (Tantivy).

Current metrics

Measured: September 5, 2026, GitHub API.

MetricValue
Stars62,344
Forks1,687
Subscribers120
Open issues per the API35
Commits≈2,233
Primary languagePython
LicenseBSL 1.1 (“Other” in the API)
Default branchmain
CreatedNovember 27, 2022
Latest releasev0.32.1, August 1, 2026

Top contributors returned by the API, by contribution count, were Pathway-Dev (837), zxqfd555 (436), pw-ppodhajski (136), szymondudycz (123), olruas (122), KamilPiechowiak (113), and berkecanrizai (92). The ≈2,233 commit count comes from the final page of the commits API’s pagination link. The general response’s watchers_count mirrors the star count, and open_issues_count (35) may include open pull requests. On PyPI, the current version is 0.32.1 with 80 published releases and requires_python >=3.10; pypistats reported 172 downloads in the last day, 927 in the week, and 7,273 in the month. On Docker Hub, the pathwaycom/pathway image logged 14,000 cumulative downloads and 2 stars (last updated July 31, 2026).

Community reception

The evidence gathered shows concrete interest, especially in live RAG, and practical questions to the team:

  • On the Hacker News thread 40669434 —“Show HN: Pathway – Build Mission Critical ETL and RAG in Python (NATO, F1 Used),” posted June 13, 2024 by janchorowski (CTO)— it reached 73 points. Among the verified comments: pipboyguy wrote “I’ve built DE and AI solutions based on Pathway for multiple clients. It’s robust and fast” (evidence of professional use, not an independent evaluation); snowpid joked ironically about the corporate usage (“Good old ‘Enterprise’ NATO!”); Arimbr asked whether, with everything kept in memory, Pathway persists state anywhere; threecheese asked about the “Community 8 GB RAM / 4 cores” plan and whether anything is always hosted; sriyansh7 asked what use cases it sees for RAG over streaming data.
  • Team members replied: dxtrous (Adrian, from Pathway) clarified that everything is RAM-based and that persistence/caching relies on configurable file backends (S3 or local), that RAM usage mostly depends on data size (especially with an in-memory index), and that the domain remains mostly unstructured data (messaging, live communications indexing, social media, news). janchorowski thanked and asked for use-case details.
  • Other Show HNs for the project, with smaller traction: 39852514 “Adaptive RAG – How we cut LLM costs without sacrificing accuracy” (8 points, 2024-03-28, by dxtrous); 38304483 “Alerting in realtime RAG: spot changes to LLM answers, using few tokens” (8 points, 5 comments, 2023-11-17); 42318221 “Build Live AI and RAG Pipelines in Minutes with YAML Templates” (8 points, 2024-12-04, by janchorowski); 40987194 “Private RAG with Mistral, Ollama and Pathway” (5 points, by berkecanrizai).

Reddit couldn’t be accessed during this research (the public API returned 403), so no Reddit threads are reported, nor was a Product Hunt page found or a verified YouTube video of the product. No broader discussion or consensus beyond what the sources show is inferred as a result. Press coverage of the organization (WSJ, Forbes, MIT Technology Review, Live Science, Semafor, TechCrunch) mostly focuses on the BDH architecture and valuation, not the ETL framework itself.

Pathway against other proposals

A floating neon ecosystem gallery: app cards for RAG question answering, live document indexing, multimodal RAG, enterprise search, private local RAG, slides search, video RAG, and unstructured-to-SQL, each card contains abstract icons of documents, videos, databases, and chat interfaces, connected to a central live-data framework hub, dark mode, cyberpunk aesthetic, electric blue and orange accents, ultra-detailed, 8K

ProjectVerified relationshipVerified difference
Apache Flink / Apache Spark / Kafka StreamsStream- and data-processing engines Pathway explicitly compares itself to in its README (claiming to outperform them) and in pathwaycom/pathway-benchmarks.Flink/Spark/Kafka Streams are mature, open streaming engines; Pathway offers a high-level Python API, batch/stream unification with one codebase, and an integrated LLM/RAG extension. Performance figures come from the project itself, not an independent evaluation.
Vector databases + Redis + FastAPI (e.g. Pinecone/Weaviate/Qdrant)A classic stack for building RAG with a vector index, cache, and separate API.The llm-app README frames Pathway as replacing those modules with a single framework with an integrated vector index (usearch) and full-text index (Tantivy), keeping the index live-synced.
LangChain / LlamaIndexFrameworks for building LLM/RAG applications.Not competitors: Pathway integrates with both as a vector store/retriever. The difference is layer-level: Pathway is the underlying live-data engine.
AirbyteData connectivity tool.Pathway includes an Airbyte connector for over 300 sources; it doesn’t compete, it uses it as a source.

The most useful comparison isn’t by popularity: Pathway stands out when you want a single Python codebase for batch and stream with live RAG; mature engines (Flink/Spark) may be preferable when the requirement is an open, widely-adopted streaming stack, and AI frameworks (LangChain/LlamaIndex) solve the agent-orchestration layer, not the data engine.

How to contribute

The process documented in CONTRIBUTING.md is explicit and has an unusual nuance:

  • The main channel is Discord (a forum channel for questions, design discussion, or RFCs) and the GitHub Issue Tracker for bugs, crashes, and performance problems.
  • Opening an issue is declared “the highest-value contribution”: the team asks contributors to describe how they use Pathway, what’s missing, friction points, bugs, and recurring usage patterns.
  • For code, a PR is accepted if (a) it’s small and self-contained (a fix, a correction, a documentation improvement) or (b) it has explicit maintainer approval (requested via issue or Discord, with prior design advice).
  • For large changes without prior approval, the team “reserves the right to close them without a detailed review,” arguing that reconstructing the context of a large external change costs more than solving the problem from scratch.
  • For libraries or connectors to be integrated, the README suggests publishing them first as a separate repository under MIT/Apache 2.0.

Use cases and who this repository can help

  • Data teams running pipelines over Kafka who want real-time analytics with a single codebase can use the framework so the same code serves local development, CI/CD tests, and streaming production, avoiding maintaining a separate batch and streaming implementation.
  • AI teams building RAG over documents that change (contracts, financial reports, communications) can start from pathwaycom/llm-app templates (e.g. Private RAG with Ollama and Mistral or Multimodal RAG with GPT-4o) so the vector index stays live and answers reflect current sources, without separately managing vector store, cache, and API.
  • People integrating RAG with LangChain or LlamaIndex can use Pathway as the underlying vector store/retriever, documented in both ecosystems, and as a retrieval backend for their applications.
  • Teams needing RAG over unstructured data at scale (billions of document pages, per the llm-app README) can use the templates optimized for accuracy or simplicity and deploy them on Docker, Kubernetes, or clouds (GCP, AWS, Azure, Render).
  • Researchers or reasoning-model evaluators can use pathwaycom/arc-task-gen to generate ARC-AGI-1-style tasks not in the public set, useful for evaluating models like BDH-CQ.

Resources


Methodology note: this article draws on the README, CONTRIBUTING.md, release notes, the GitHub API, PyPI, Docker Hub, the official site (pathway.com), and Hacker News threads consulted on September 5, 2026. Figures change over time. Data about the organization and its BDH pivot comes from its own site and the press cited; the framework’s performance claims come from the project itself and don’t constitute an independent evaluation. Reddit couldn’t be accessed (403) and no Product Hunt page or verified YouTube video of the product was found.

Comments