llm-app: production-ready RAG templates for the Pathway framework
pathwaycom/llm-app · 58,898★ · 1,500 forks
llm-app is the collection of official AI application templates built on the Pathway Live Data Framework (pathwaycom/pathway, the sibling engine repository containing the library and Rust engine): RAG, document and video search, and real-time indexes that stay continuously synced with data sources. Each template is a deploy-ready Docker container exposing an HTTP API, with the vector index, hybrid index, and cache built into the library itself, no separate vector database, cache, or API framework required.
Origin
The repository was created on July 19, 2023 by the pathwaycom organization (the company Pathway, pathway.com), eight months after the base framework pathwaycom/pathway existed (created November 27, 2022). Eight days after creation, on July 27, 2023, Pathway launched it on Hacker News: “Show HN: LLM App – build a realtime LLM app in 30 lines, with no vector database” (thread 36894142), reaching 11 points and 8 comments.
The launch thread’s title sets the narrative: a real-time LLM app in ~30 lines with no vector database. In it, user Arimbr asked what exactly the Pathway package was, and janchorowski (Pathway) replied that it’s a data-processing framework unifying stream and batch processing of large datasets, letting developers focus on logic instead of manually tracking data changes.
The repository has evolved into a catalog: the original pitch was a single “30 lines” example; today (2026) it bundles ten templates. Subsequent Hacker News appearances follow the same per-template “Show HN” pattern: 38304483 (2023-11-17, real-time RAG alerting, 8 pts / 5 comments), 40987194 (2024-07-17, private RAG with Mistral and Ollama, 5 pts / 0 comments), and 42318221 (2024-12-04, YAML templates, 8 pts / 2 comments). In the base framework’s thread —40669434 “Show HN: Pathway – Build Mission Critical ETL and RAG in Python (NATO, F1 Used)” (73 pts, 19 comments, 2024-06-13)— the team itself claims the product is used by clients the headline summarizes as “NATO and F1”; it’s the family’s most visible public appearance.
Philosophy and principles
The README condenses the philosophy into several verifiable ideas:
- Always-in-sync knowledge: the core value is that documents are re-indexed in real time on any addition, deletion, or change to the source. There’s no separate ETL step or “nightly batch”; the index reflects the source’s current state.
- No auxiliary infrastructure: the README states this literally by striking through the usual stack: “
Vector Database (e.g. Pinecone/Weaviate/Qdrant) + Cache (e.g. Redis) + API Framework (e.g. Fast API).” The vector index, hybrid index, and cache are integrated into the library (by default usearch for the KNN index and Tantivy for hybrid full-text search). - “One-line” templates: the READMEs claim that changing a pipeline step is “a one-line change,” and that templates differ only in their
app.yamlfile. - Declared scale: the README states templates scale “to millions of document pages.”
Incrementality is a cross-cutting principle: any change to source files immediately flows through parsing, embedding, indexing, and LLM responses.

How it works
Each template runs as a Docker container exposing an HTTP API (some include an optional Streamlit UI). Everything depends on the Pathway Live Data Framework, a Python library with an integrated Rust engine that syncs sources and serves requests.
The pipeline documented for the default template (question_answering_rag) has five stages:
- Data ingestion: sources are defined in
app.yaml(local directories, Google Drive, SharePoint, S3, Kafka, PostgreSQL, APIs). The code chains these sources at configurable intervals and re-indexes in real time. - Parsing and chunking: with Docling (via
DoclingParser) andTokenCountSplitter. - Embedding:
OpenAIEmbedderby default; swappable for another embedder. - Indexing:
USearchKnnFactoryfor the vector index; optionally a hybrid index combiningUSearchKNNwithTantivyBM25viaHybridIndex. - Serving:
SummaryQuestionAnswererhandles questions and summaries, andQASummaryRestServerexposes the endpoints.

Default template’s default endpoints:
| Category | Endpoint | Use |
|---|---|---|
| Document indexing | /v1/retrieve | Similarity search |
/v1/statistics | Basic index statistics | |
/v2/list_documents | Metadata for all processed files | |
| LLM / RAG | /v2/answer | Ask the documents or talk to the LLM |
/v2/summarize | Summarize a list of texts |

The repository’s ten templates (verified in templates/):
| Template | Description (per its README) |
|---|---|
question_answering_rag | Basic end-to-end RAG: answers over PDF/DOCX from a live source |
document_indexing | Live indexing acting as a vector store / retriever (backend for LangChain or LlamaIndex) |
multimodal_rag | Multimodal RAG with GPT-4o; extracts tables and charts from financial documents |
unstructured_to_sql_on_the_fly | Converts unstructured financial data into SQL + loads into PostgreSQL |
adaptive_rag | RAG with Adaptive RAG to reduce token cost up to 4× while keeping accuracy |
private_rag | Fully private/local RAG with Mistral and Ollama |
slides_ai_search | Multimodal indexing of slides (PowerPoint and PDF) with a live index |
video_rag_twelvelabs | Video RAG with TwelveLabs |
document_store_mcp_server | Exposes live indexing as an MCP server (three tools, JMESPath filtering) |
drive_alert | RAG over Google Drive with real-time alerts (Slack) when answers change |

The ecosystem
Main sibling repository (the engine). pathwaycom/llm-app is the template catalog for pathwaycom/pathway, the base framework containing the library and Rust engine: “Python ETL framework for stream processing, real-time analytics, LLM pipelines, and RAG”; 62,344 stars, 1,686 forks, created November 27, 2022, custom license (flagged NOASSERTION by the API), latest release v0.32.1 (August 1, 2026), distributed on PyPI as the pathway package.

Other repositories in the pathwaycom organization (star counts per the GitHub API, highest to lowest): pathwaycom/arc-task-gen — 10,565 stars — “Generates original ARC-AGI-1-style tasks”; pathwaycom/bdh — 3,544 stars — “BDH (Dragon Hatchling) – Architecture and Code”; pathwaycom/pathway-benchmarks — 302 stars; pathwaycom/cookiecutter-pathway — 277 stars; pathwaycom/serviette — 4 stars — “Universal RAG tool.”

External libraries the architecture incorporates: usearch (unum-cloud/usearch) for the default KNN index, Tantivy (quickwit-oss/tantivy) for the hybrid index, and Docling (docling.ai) for document parsing. Documented integrations: LangChain and LlamaIndex (the document_indexing template can serve as a retrieval backend for both), TwelveLabs (video RAG), the Model Context Protocol (MCP), Slack (alerts in drive_alert), and LLM providers OpenAI, Mistral, and Ollama.

Official and semi-official status
llm-app is a first-party Pathway product: the README ends with a “Supported and maintained by” section explicitly crediting Pathway, and the repository is the official template channel (pathway.com/developers/templates/). There’s no third-party “marketplace acceptance”; its status is that of the company’s official reference repository for building RAG apps on its own framework.
In practice: it’s the set of reference implementations Pathway maintains and backs (bugs are filed against the base framework’s tracker, not llm-app); the underlying framework ships as the official pathway package on PyPI; the SharePoint connector is gated behind a Scale or Enterprise license key (Pathway offers a free Scale key); and the base framework appears on Trendshift as a trending repository.

Quick-start guide
Installation and first boot
# 1. Clone the repository
git clone https://github.com/pathwaycom/llm-app.git
# 2. Enter the template
cd templates/question_answering_rag
# 3. Install dependencies (local run only, not Windows)
pip install -r requirements.txt
# 4. Set the OpenAI key in .env -> OPENAI_API_KEY=sk-...
# 5. Run
python app.py
Or with Docker (the only documented route for Windows):
docker compose build
docker compose up
On first boot you’ll see parsing and embedding of the sample documents in the logs. Once you see 0 entries (x minibatch(es)) have been... and updates stop arriving, the app is ready:
curl -X POST http://localhost:8000/v1/statistics \
-H 'accept: */*' -H 'Content-Type: application/json'
Common workflows
- List indexed documents:
POST /v2/list_documents→ metadata for all processed files. - Similarity search:
POST /v1/retrieve -d '{"query": "your question", "k": 6}'→ most similar chunks (supportsmetadata_filtervia JMESPath andfilepath_globpattern). - Ask with RAG:
POST /v2/answerandPOST /v2/summarize. - Change the model: in
app.yaml, edit themodelfield of the$llmblock (defaultgpt-4.1-mini; examples showgpt-5,gpt-4.1,gpt-4o, orLiteLLMChatpointing to local Ollama).
Essential configuration
app.yaml: defines data sources, LLM, embedder, and index; the only file that differs between templates..env:OPENAI_API_KEYrequired (and fordrive_alert,slack_alert_channel_idandslack_alert_token).$sourcesinapp.yaml: local, Google Drive (defaultrefresh_interval30s), or SharePoint (requires a license key).host/port: default0.0.0.0/8000.- Cache / persistence:
persistence_mode,persistence_backend(e.g.Backend.filesystem).
Common pitfalls and fixes
- Windows doesn’t support local execution: the README says to use Docker.
- Querying before initial indexing finishes: wait for the
0 entries (x minibatch(es)) have been...message. - SharePoint requires a license: without a Scale/Enterprise key, that connector doesn’t work.
- Google Drive requires a service account: needs
object_idand valid credentials. - LLM-based alert deduplication: on thread 38304483, user
pstormsuggested comparing embedding similarity instead of spending LLM calls to detect changes;janchorowskireplied the criterion is easily swappable and the LLM route is faster to start with, but over time an embedding-based criterion can be trained.

Integrations and migration
- LlamaIndex / LangChain retrieval backend: the
document_indexingtemplate serves as a vector store / retriever from any frontend or LlamaIndex/LangChain app. - MCP:
document_store_mcp_serverexposes live indexing with three tools (retrieve_query,statistics_query,inputs_query). - Slack:
drive_alertsends notifications when an answer changes. - Cloud deployment: deployment guides for GCP, AWS Fargate, Azure (ACI), and Render.
- Swapping components: migrating from a vector index to a hybrid one, or from OpenAI to local Ollama/Mistral, is a YAML change without touching
app.py.
Current metrics
Measured: September 4, 2026, GitHub API.
| Metric | Value |
|---|---|
| Stars | 58,937 |
| Forks | 1,478 |
| Subscribers | 94 |
| Commits | 253 |
| Open issues per the API | 8 |
| Primary language | Jupyter Notebook (409,401 bytes); secondary: Python (68,370), Dockerfile (3,888) |
| License | MIT |
| Created | July 19, 2023 |
| Last push | July 5, 2026 |
| Latest release | None: the repository doesn’t use releases or versioned tags |
Top contributors were berkecanrizai (48), szymondudycz (45), pw-ppodhajski (33), dxtrous (19), olruas (14), Boburmirzo (9), janchorowski (9), and embe-pw (9). The 253 commit count came from the commits endpoint’s pagination header. open_issues_count may include open pull requests. watchers_count mirrors the star count, so subscribers_count is reported separately as real subscribers.
Community reception
The evidence gathered comes from several Hacker News threads, all with small-to-moderate audiences, plus one base-framework thread that did reach a notable figure:
Recognition:
- In the base framework’s thread 40669434 (73 pts, 19 comments),
articsputnikwrote it’s “impressive to see a Python tool for ETL and RAG tasks with such solid features” and that the focus on security and performance, “especially with self-hosted RAG pipelines, is fantastic.” pipboyguy, in the same thread, contributed a real-world use case: “I’ve built DE and AI solutions based on Pathway for multiple clients. It’s robust and fast.”devnull777: “Your site’s examples look easy to reproduce. By the way, what a beautiful and clear website!”- In 36894142 (11 pts, 8 comments),
Arimbrshowed early interest andanupsurendranclosed the exchange saying they “understand much better” the index/database distinction after the team’s reply.
Concrete criticism and skepticism:
- In 40669434,
snowpidreacted skeptically to the “NATO” label: “Good old ‘Enterprise’ NATO! Always good for a surprise,” andsuziemanuljokingly replied “some say it’s not Fortune 100 but Fortune 1.” The team didn’t deny the headline’s claim. threecheese, in the same thread, asked about RAM on the “Community” plan (8 GB) and noted the Rust engine is “shipped as compiled executable, with a license check”;dxtrous(Pathway) replied the main driver of RAM usage is the size of the fed data.- In 42318221 (8 pts, 2 comments),
Arimbrasked to “see all supported fields and values in one place”;janchorowskireplied they would. - In 38304483 (8 pts, 5 comments),
pstormquestioned the deduplication approach, andbobur_umurzokovbroadened the scope by proposing it for fraud detection, customer support, medical diagnosis, and predictive maintenance.
The balance: technical enthusiasm and real use cases, with mostly modest audiences on the templates, but a base-framework thread with 73 points and 19 comments introducing reasoned skepticism about the corporate headline and the hosting/licensing model.
Comparison with similar projects
| Project | Verifiable overlap | Verifiable difference |
|---|---|---|
| Classic RAG stack: Pinecone / Weaviate / Qdrant / pgvector + Redis + FastAPI | All solve “RAG over corporate documents” with vector store + cache + API. | The README states it does not require assembling those components separately: vector index, hybrid index, and cache are integrated into the library. |
| LlamaIndex / LangChain | Both are document-retrieval frameworks for LLMs. | llm-app treats them as clients of its index, and adds real-time source synchronization. |
pathwaycom/pathway (sibling repository) | Same underlying framework; llm-app only holds templates built on it. | pathway is the library/engine (Python + Rust, PyPI); llm-app is the “deploy-ready” app catalog. |
| TwelveLabs (video RAG) | Offers video RAG. | llm-app consumes TwelveLabs inside its video_rag_twelvelabs template; it doesn’t compete, it integrates. |
The most useful comparison: llm-app stands out when the requirement is RAG that stays synced in real time with corporate sources, without assembling separate vector/cache/API infrastructure.
How to contribute
- Accepts contributions for documentation, features, fixes, code cleanup, tests, and reviews.
- The recommended route for larger contributions is to flag it in Pathway’s Discord server (#get-help channel) before starting.
- Bugs and feedback are filed against the base framework’s tracker (
pathwaycom/pathway/issues), notllm-app. - Since templates each have a single
app.yaml, the practical way to contribute a variation is to start from an existing template and adjust its YAML.
No CONTRIBUTING.md or specific PR template was found in the repository.
Use cases and who this repository can help
- Teams building RAG over live corporate documents (SharePoint, Google Drive, S3, Kafka, PostgreSQL, APIs):
question_answering_ragplus the right connector gives a self-updating index.document_indexingreuses that index as a backend for an existing LlamaIndex/LangChain app. - Finance departments / analysts:
multimodal_ragextracts tables and charts from financial documents;unstructured_to_sql_on_the_flystructures PDF reports into PostgreSQL. - Teams with privacy / data-residency constraints:
private_rag(Mistral + Ollama) enables a fully local pipeline. - Those looking to cut token cost:
adaptive_ragclaims up to a 4× reduction while keeping accuracy. - LLM agent / MCP integrators:
document_store_mcp_serverexposes live indexing as an MCP server. - Teams monitoring knowledge-base changes:
drive_alertsends Slack alerts after a document change. - Content teams / search over slides and video:
slides_ai_searchandvideo_rag_twelvelabs.
Resources
- Repository: https://github.com/pathwaycom/llm-app
- Base framework (engine): https://github.com/pathwaycom/pathway
- Documentation / official site: https://pathway.com/developers/templates/
- YAML configuration guide: https://pathway.com/developers/templates/configure-yaml
- Templates (code): https://github.com/pathwaycom/llm-app/tree/main/templates
- Package registries: PyPI
pathway(latest 0.32.1): https://pypi.org/project/pathway/ - Trendshift: https://trendshift.io/repositories/4400
- Community / Discord: https://discord.gg/pathway
- Relevant Hacker News threads: https://news.ycombinator.com/item?id=40669434 , https://news.ycombinator.com/item?id=36894142 , https://news.ycombinator.com/item?id=38304483 , https://news.ycombinator.com/item?id=40987194 , https://news.ycombinator.com/item?id=42318221
Methodology note: this article draws on the pathwaycom/llm-app README, the template READMEs, the GitHub API, Hacker News threads, and PyPI, consulted on September 4, 2026. Star and version figures change over time.
Comments