September 12, 2026 · By YasKad
pathwaycom/llm-app

llm-app: production-ready RAG templates for the Pathway framework

pathwaycom/llm-app · 58,898★ · 1,500 forks

llm-app is the collection of official AI application templates built on the Pathway Live Data Framework (pathwaycom/pathway, the sibling engine repository containing the library and Rust engine): RAG, document and video search, and real-time indexes that stay continuously synced with data sources. Each template is a deploy-ready Docker container exposing an HTTP API, with the vector index, hybrid index, and cache built into the library itself, no separate vector database, cache, or API framework required.

Origin

The repository was created on July 19, 2023 by the pathwaycom organization (the company Pathway, pathway.com), eight months after the base framework pathwaycom/pathway existed (created November 27, 2022). Eight days after creation, on July 27, 2023, Pathway launched it on Hacker News: “Show HN: LLM App – build a realtime LLM app in 30 lines, with no vector database” (thread 36894142), reaching 11 points and 8 comments.

The launch thread’s title sets the narrative: a real-time LLM app in ~30 lines with no vector database. In it, user Arimbr asked what exactly the Pathway package was, and janchorowski (Pathway) replied that it’s a data-processing framework unifying stream and batch processing of large datasets, letting developers focus on logic instead of manually tracking data changes.

The repository has evolved into a catalog: the original pitch was a single “30 lines” example; today (2026) it bundles ten templates. Subsequent Hacker News appearances follow the same per-template “Show HN” pattern: 38304483 (2023-11-17, real-time RAG alerting, 8 pts / 5 comments), 40987194 (2024-07-17, private RAG with Mistral and Ollama, 5 pts / 0 comments), and 42318221 (2024-12-04, YAML templates, 8 pts / 2 comments). In the base framework’s thread —40669434 “Show HN: Pathway – Build Mission Critical ETL and RAG in Python (NATO, F1 Used)” (73 pts, 19 comments, 2024-06-13)— the team itself claims the product is used by clients the headline summarizes as “NATO and F1”; it’s the family’s most visible public appearance.

Philosophy and principles

The README condenses the philosophy into several verifiable ideas:

  • Always-in-sync knowledge: the core value is that documents are re-indexed in real time on any addition, deletion, or change to the source. There’s no separate ETL step or “nightly batch”; the index reflects the source’s current state.
  • No auxiliary infrastructure: the README states this literally by striking through the usual stack: “Vector Database (e.g. Pinecone/Weaviate/Qdrant) + Cache (e.g. Redis) + API Framework (e.g. Fast API).” The vector index, hybrid index, and cache are integrated into the library (by default usearch for the KNN index and Tantivy for hybrid full-text search).
  • “One-line” templates: the READMEs claim that changing a pipeline step is “a one-line change,” and that templates differ only in their app.yaml file.
  • Declared scale: the README states templates scale “to millions of document pages.”

Incrementality is a cross-cutting principle: any change to source files immediately flows through parsing, embedding, indexing, and LLM responses.

A detailed dark-mode cyberpunk visualization of a luminous data engine with multiple data sources converging as neon streams — a local folder, a cloud drive, an S3 bucket, a Kafka stream, a PostgreSQL cylinder, a live API node — into a central engine emitting a holographic answer stream toward a chat-like interface, a small glowing Docker capsule and an HTTP API gateway orbiting nearby, deep navy and black background, electric cyan, violet, amber, and magenta neon accents, ultra-detailed micro-circuitry, volumetric lighting, 8K

How it works

Each template runs as a Docker container exposing an HTTP API (some include an optional Streamlit UI). Everything depends on the Pathway Live Data Framework, a Python library with an integrated Rust engine that syncs sources and serves requests.

The pipeline documented for the default template (question_answering_rag) has five stages:

  1. Data ingestion: sources are defined in app.yaml (local directories, Google Drive, SharePoint, S3, Kafka, PostgreSQL, APIs). The code chains these sources at configurable intervals and re-indexes in real time.
  2. Parsing and chunking: with Docling (via DoclingParser) and TokenCountSplitter.
  3. Embedding: OpenAIEmbedder by default; swappable for another embedder.
  4. Indexing: USearchKnnFactory for the vector index; optionally a hybrid index combining USearchKNN with TantivyBM25 via HybridIndex.
  5. Serving: SummaryQuestionAnswerer handles questions and summaries, and QASummaryRestServer exposes the endpoints.

A detailed cyberpunk visualization of real-time data synchronization for a live RAG system, a central holographic document index receiving simultaneous updates from multiple glowing source nodes — local directory, cloud drive, S3 bucket, Kafka message stream, PostgreSQL database, live API endpoint — each emitting a thin neon pulse on every document change, dark backgrounds, cyan and magenta light trails, subtle grid planes, volumetric haze, ultra-detailed tech panels, 8K

Default template’s default endpoints:

CategoryEndpointUse
Document indexing/v1/retrieveSimilarity search
/v1/statisticsBasic index statistics
/v2/list_documentsMetadata for all processed files
LLM / RAG/v2/answerAsk the documents or talk to the LLM
/v2/summarizeSummarize a list of texts

A conceptual dark tech illustration of an integrated AI backend architecture that removes the need for separate infrastructure, a central crystalline module with three internal chambers — vector index, hybrid keyword index, cache layer — fused into one compact engine, faint red holographic strikethroughs crossing out abstract external components like a vector database server and a standalone cache, dark mode palette with black and deep blue with neon teal highlights, clean futuristic lines, glowing internal chambers, ultra-detailed 3D rendering, 8K

The repository’s ten templates (verified in templates/):

TemplateDescription (per its README)
question_answering_ragBasic end-to-end RAG: answers over PDF/DOCX from a live source
document_indexingLive indexing acting as a vector store / retriever (backend for LangChain or LlamaIndex)
multimodal_ragMultimodal RAG with GPT-4o; extracts tables and charts from financial documents
unstructured_to_sql_on_the_flyConverts unstructured financial data into SQL + loads into PostgreSQL
adaptive_ragRAG with Adaptive RAG to reduce token cost up to 4× while keeping accuracy
private_ragFully private/local RAG with Mistral and Ollama
slides_ai_searchMultimodal indexing of slides (PowerPoint and PDF) with a live index
video_rag_twelvelabsVideo RAG with TwelveLabs
document_store_mcp_serverExposes live indexing as an MCP server (three tools, JMESPath filtering)
drive_alertRAG over Google Drive with real-time alerts (Slack) when answers change

A five-stage AI data pipeline rendered as a glowing cyberpunk conveyor system: live data ingestion from multiple sources, document parsing and chunking with abstract page fragments splitting into structured blocks, embedding generation represented as dense clusters of glowing vector points, indexing shown as a combined semantic and keyword lattice, and serving shown as an API response stream flowing toward a chat interface, each stage a distinct neon module connected by flowing data streams, dark background, electric cyan, amber, and violet accents, ultra-detailed industrial-tech aesthetic, cinematic depth, 8K

The ecosystem

Main sibling repository (the engine). pathwaycom/llm-app is the template catalog for pathwaycom/pathway, the base framework containing the library and Rust engine: “Python ETL framework for stream processing, real-time analytics, LLM pipelines, and RAG”; 62,344 stars, 1,686 forks, created November 27, 2022, custom license (flagged NOASSERTION by the API), latest release v0.32.1 (August 1, 2026), distributed on PyPI as the pathway package.

A futuristic Docker container acting as a self-contained AI application service, a sleek glowing capsule with translucent walls revealing an internal Pathway-style engine, vector index, hybrid index, cache, and LLM serving layer, outside several HTTP API endpoints appear as illuminated ports or gates sending and receiving neon request-response streams, the capsule sits on a dark server rack with subtle cyberpunk lighting, black, graphite, electric blue, cyan, and magenta, ultra-detailed 3D rendering, soft glow, reflective surfaces, 8K

Other repositories in the pathwaycom organization (star counts per the GitHub API, highest to lowest): pathwaycom/arc-task-gen — 10,565 stars — “Generates original ARC-AGI-1-style tasks”; pathwaycom/bdh — 3,544 stars — “BDH (Dragon Hatchling) – Architecture and Code”; pathwaycom/pathway-benchmarks — 302 stars; pathwaycom/cookiecutter-pathway — 277 stars; pathwaycom/serviette — 4 stars — “Universal RAG tool.”

A holographic catalog of ten AI application templates arranged in a dark cyberpunk dashboard, ten glowing cards in a grid, each with a distinct abstract icon — chat bubble, vector lattice, document with charts, database-to-SQL flow, cost-saving adaptive loop, local secure node, presentation slide, video timeline, protocol server, warning pulse — floating above a dark reflective surface with neon edges, dark background, cyan and violet neon accents, 8K

External libraries the architecture incorporates: usearch (unum-cloud/usearch) for the default KNN index, Tantivy (quickwit-oss/tantivy) for the hybrid index, and Docling (docling.ai) for document parsing. Documented integrations: LangChain and LlamaIndex (the document_indexing template can serve as a retrieval backend for both), TwelveLabs (video RAG), the Model Context Protocol (MCP), Slack (alerts in drive_alert), and LLM providers OpenAI, Mistral, and Ollama.

An ecosystem diagram of a software organization's AI framework family, rendered as a dark cosmic-tech constellation, a powerful Rust-based central core glowing with deep blue and cyan energy representing the main framework, orbiting satellite repositories and tools — a template catalog, a benchmark suite, a project scaffolding tool, a universal RAG utility, an experimental task-generation project — thin luminous lines connecting the nodes, a prominent stream flowing from PyPI-like packaging infrastructure into the central engine, starfield details, circuit-like constellations, volumetric light, ultra-detailed dark cyberpunk style, 8K

Official and semi-official status

llm-app is a first-party Pathway product: the README ends with a “Supported and maintained by” section explicitly crediting Pathway, and the repository is the official template channel (pathway.com/developers/templates/). There’s no third-party “marketplace acceptance”; its status is that of the company’s official reference repository for building RAG apps on its own framework.

In practice: it’s the set of reference implementations Pathway maintains and backs (bugs are filed against the base framework’s tracker, not llm-app); the underlying framework ships as the official pathway package on PyPI; the SharePoint connector is gated behind a Scale or Enterprise license key (Pathway offers a free Scale key); and the base framework appears on Trendshift as a trending repository.

A secure cyberpunk data-vault scene for a fully private, local, air-gapped RAG system, at the center a compact server node contains local model engines, a vector index, a hybrid index, and a cache, all fully self-contained, documents entering from a local folder or local cloud mirror, parsed and embedded inside the vault, answers returned to a local interface without crossing an external cloud boundary, shield-like hexagonal barriers and faint red "no external cloud" exclusion fields, dark graphite surfaces, neon green and cyan accents, subtle lock icons, reflective floor, volumetric lighting, ultra-detailed tech rendering, 8K

Quick-start guide

Installation and first boot

# 1. Clone the repository
git clone https://github.com/pathwaycom/llm-app.git

# 2. Enter the template
cd templates/question_answering_rag

# 3. Install dependencies (local run only, not Windows)
pip install -r requirements.txt

# 4. Set the OpenAI key in .env  ->  OPENAI_API_KEY=sk-...

# 5. Run
python app.py

Or with Docker (the only documented route for Windows):

docker compose build
docker compose up

On first boot you’ll see parsing and embedding of the sample documents in the logs. Once you see 0 entries (x minibatch(es)) have been... and updates stop arriving, the app is ready:

curl -X POST http://localhost:8000/v1/statistics \
  -H 'accept: */*' -H 'Content-Type: application/json'

Common workflows

  • List indexed documents: POST /v2/list_documents → metadata for all processed files.
  • Similarity search: POST /v1/retrieve -d '{"query": "your question", "k": 6}' → most similar chunks (supports metadata_filter via JMESPath and filepath_globpattern).
  • Ask with RAG: POST /v2/answer and POST /v2/summarize.
  • Change the model: in app.yaml, edit the model field of the $llm block (default gpt-4.1-mini; examples show gpt-5, gpt-4.1, gpt-4o, or LiteLLMChat pointing to local Ollama).

Essential configuration

  • app.yaml: defines data sources, LLM, embedder, and index; the only file that differs between templates.
  • .env: OPENAI_API_KEY required (and for drive_alert, slack_alert_channel_id and slack_alert_token).
  • $sources in app.yaml: local, Google Drive (default refresh_interval 30s), or SharePoint (requires a license key).
  • host / port: default 0.0.0.0 / 8000.
  • Cache / persistence: persistence_mode, persistence_backend (e.g. Backend.filesystem).

Common pitfalls and fixes

  • Windows doesn’t support local execution: the README says to use Docker.
  • Querying before initial indexing finishes: wait for the 0 entries (x minibatch(es)) have been... message.
  • SharePoint requires a license: without a Scale/Enterprise key, that connector doesn’t work.
  • Google Drive requires a service account: needs object_id and valid credentials.
  • LLM-based alert deduplication: on thread 38304483, user pstorm suggested comparing embedding similarity instead of spending LLM calls to detect changes; janchorowski replied the criterion is easily swappable and the LLM route is faster to start with, but over time an embedding-based criterion can be trained.

A futuristic representation of a live document store exposed as an MCP server with real-time alerting capabilities, a central server module connected to three glowing tool gates — retrieval queries, statistics queries, input queries — a filter panel shaped like a JMESPath-style query structure beside it, on the right a real-time alert pulse traveling from a Google Drive-like source through the index into a Slack-like notification channel, triggered when a document change alters an answer, dark mode, cyberpunk aesthetic, neon cyan, amber, and magenta, ultra-detailed 3D panels, reflective surfaces, 8K

Integrations and migration

  • LlamaIndex / LangChain retrieval backend: the document_indexing template serves as a vector store / retriever from any frontend or LlamaIndex/LangChain app.
  • MCP: document_store_mcp_server exposes live indexing with three tools (retrieve_query, statistics_query, inputs_query).
  • Slack: drive_alert sends notifications when an answer changes.
  • Cloud deployment: deployment guides for GCP, AWS Fargate, Azure (ACI), and Render.
  • Swapping components: migrating from a vector index to a hybrid one, or from OpenAI to local Ollama/Mistral, is a YAML change without touching app.py.

Current metrics

Measured: September 4, 2026, GitHub API.

MetricValue
Stars58,937
Forks1,478
Subscribers94
Commits253
Open issues per the API8
Primary languageJupyter Notebook (409,401 bytes); secondary: Python (68,370), Dockerfile (3,888)
LicenseMIT
CreatedJuly 19, 2023
Last pushJuly 5, 2026
Latest releaseNone: the repository doesn’t use releases or versioned tags

Top contributors were berkecanrizai (48), szymondudycz (45), pw-ppodhajski (33), dxtrous (19), olruas (14), Boburmirzo (9), janchorowski (9), and embe-pw (9). The 253 commit count came from the commits endpoint’s pagination header. open_issues_count may include open pull requests. watchers_count mirrors the star count, so subscribers_count is reported separately as real subscribers.

Community reception

The evidence gathered comes from several Hacker News threads, all with small-to-moderate audiences, plus one base-framework thread that did reach a notable figure:

Recognition:

  • In the base framework’s thread 40669434 (73 pts, 19 comments), articsputnik wrote it’s “impressive to see a Python tool for ETL and RAG tasks with such solid features” and that the focus on security and performance, “especially with self-hosted RAG pipelines, is fantastic.”
  • pipboyguy, in the same thread, contributed a real-world use case: “I’ve built DE and AI solutions based on Pathway for multiple clients. It’s robust and fast.”
  • devnull777: “Your site’s examples look easy to reproduce. By the way, what a beautiful and clear website!”
  • In 36894142 (11 pts, 8 comments), Arimbr showed early interest and anupsurendran closed the exchange saying they “understand much better” the index/database distinction after the team’s reply.

Concrete criticism and skepticism:

  • In 40669434, snowpid reacted skeptically to the “NATO” label: “Good old ‘Enterprise’ NATO! Always good for a surprise,” and suziemanul jokingly replied “some say it’s not Fortune 100 but Fortune 1.” The team didn’t deny the headline’s claim.
  • threecheese, in the same thread, asked about RAM on the “Community” plan (8 GB) and noted the Rust engine is “shipped as compiled executable, with a license check”; dxtrous (Pathway) replied the main driver of RAM usage is the size of the fed data.
  • In 42318221 (8 pts, 2 comments), Arimbr asked to “see all supported fields and values in one place”; janchorowski replied they would.
  • In 38304483 (8 pts, 5 comments), pstorm questioned the deduplication approach, and bobur_umurzokov broadened the scope by proposing it for fraud detection, customer support, medical diagnosis, and predictive maintenance.

The balance: technical enthusiasm and real use cases, with mostly modest audiences on the templates, but a base-framework thread with 73 points and 19 comments introducing reasoned skepticism about the corporate headline and the hosting/licensing model.

Comparison with similar projects

ProjectVerifiable overlapVerifiable difference
Classic RAG stack: Pinecone / Weaviate / Qdrant / pgvector + Redis + FastAPIAll solve “RAG over corporate documents” with vector store + cache + API.The README states it does not require assembling those components separately: vector index, hybrid index, and cache are integrated into the library.
LlamaIndex / LangChainBoth are document-retrieval frameworks for LLMs.llm-app treats them as clients of its index, and adds real-time source synchronization.
pathwaycom/pathway (sibling repository)Same underlying framework; llm-app only holds templates built on it.pathway is the library/engine (Python + Rust, PyPI); llm-app is the “deploy-ready” app catalog.
TwelveLabs (video RAG)Offers video RAG.llm-app consumes TwelveLabs inside its video_rag_twelvelabs template; it doesn’t compete, it integrates.

The most useful comparison: llm-app stands out when the requirement is RAG that stays synced in real time with corporate sources, without assembling separate vector/cache/API infrastructure.

How to contribute

  • Accepts contributions for documentation, features, fixes, code cleanup, tests, and reviews.
  • The recommended route for larger contributions is to flag it in Pathway’s Discord server (#get-help channel) before starting.
  • Bugs and feedback are filed against the base framework’s tracker (pathwaycom/pathway/issues), not llm-app.
  • Since templates each have a single app.yaml, the practical way to contribute a variation is to start from an existing template and adjust its YAML.

No CONTRIBUTING.md or specific PR template was found in the repository.

Use cases and who this repository can help

  • Teams building RAG over live corporate documents (SharePoint, Google Drive, S3, Kafka, PostgreSQL, APIs): question_answering_rag plus the right connector gives a self-updating index. document_indexing reuses that index as a backend for an existing LlamaIndex/LangChain app.
  • Finance departments / analysts: multimodal_rag extracts tables and charts from financial documents; unstructured_to_sql_on_the_fly structures PDF reports into PostgreSQL.
  • Teams with privacy / data-residency constraints: private_rag (Mistral + Ollama) enables a fully local pipeline.
  • Those looking to cut token cost: adaptive_rag claims up to a 4× reduction while keeping accuracy.
  • LLM agent / MCP integrators: document_store_mcp_server exposes live indexing as an MCP server.
  • Teams monitoring knowledge-base changes: drive_alert sends Slack alerts after a document change.
  • Content teams / search over slides and video: slides_ai_search and video_rag_twelvelabs.

Resources


Methodology note: this article draws on the pathwaycom/llm-app README, the template READMEs, the GitHub API, Hacker News threads, and PyPI, consulted on September 4, 2026. Star and version figures change over time.

Comments