August 13, 2026 · By YasKad
HKUDS/DeepTutor

DeepTutor: an agent-oriented personalized learning workspace

HKUDS/DeepTutor · 40,284★ · 5,073 forks

Everything worth knowing about HKUDS/DeepTutor: an open-source application that brings tutoring, practice, research, document retrieval, and student memory together in a single workspace.


What DeepTutor is

DeepTutor is an agent-oriented personalized learning workspace. Its README presents it as an extensible system for tutoring, problem solving, quiz generation, research, visualization, and domain practice. It is an application, not just a retrieval library: it includes a web interface, a command-line interface, knowledge bases, tools, and learning workspaces.

The design keeps a single agent loop across the conversation, quiz, research, visualization, solving, and domain-path modes. Materials and context — libraries, books, drafts, notebooks, question banks, profiles, and memory — can move between those modes. The core proposal is to personalize the next learning step without splitting each activity into a separate application.

The origin: from a community launch to a technical report

The repository was created on December 28, 2025, and the README places the official launch on December 29, 2025. The project is led as open source by Bingxi Zhao within the HKUDS group; the README itself thanks, among others, Chao Huang, director of HKU’s Data Intelligence Lab, along with Jiahao Zhang, Zirui Guo, and Xubin Ren.

The technical context was later formalized in the arXiv report “DeepTutor: Towards Agentic Personalized Tutoring,” submitted on April 10, 2026 by Bingxi Zhao, Jiahao Zhang, Xubin Ren, Zirui Guo, Tianzhe Chu, Yi Ma, and Chao Huang. The paper poses a concrete tension: language models start from static knowledge, and conventional retrieval systems are not enough to deliver guided feedback adapted to each student. DeepTutor proposes combining source-grounded knowledge with dynamic student memory.

The project’s public narrative is deliberately community-driven: the README states that it does not offer paid online products and that issues, pull requests, and discussions shape the product. It also records two adoption milestones: 10,000 stars in 39 days and 20,000 in 111 days; these are the project’s own announcements, not an independent audit.

Philosophy and principles

The verifiable philosophy rests on four product decisions:

  • Visible personalization: memory is split into L1 traces, L2 summaries, and L3 synthesis; the README says the memory graph lets you trace a claim back to its evidence and edit it.

Visualization of DeepTutor's three-layer memory architecture: L1 traces, L2 summaries, and L3 synthesis connected by a beam of light tracing back to the original evidence.

  • A connected context: books, notes, libraries, quizzes, and profiles are not treated as silos per learning mode.
  • Interchangeable knowledge: the retrieval layer supports LlamaIndex, PageIndex, GraphRAG, LightRAG, or a linked Obsidian vault, rather than locking into a single engine.

Illustration of interchangeable retrieval engines — LlamaIndex, GraphRAG, LightRAG, and an Obsidian vault — connected by glowing cables to a central processing core.

  • Extensibility over a closed catalog: learning resources use the open Agent Skills format, based on a folder with SKILL.md, YAML metadata, and optional reference files.

The paper backs this philosophy with TutorBench, an interactive benchmark with student profiles and university curricula from five domains. Its reported average improvements of 10.8% on personalized metrics and 29.4% on agent reasoning are results the authors declare in an ongoing technical report; they should not be read as independent external validation.

How it works

The documented usage flow is: choose a working directory, install, run deeptutor init, and start with deeptutor start. There are four installation paths: PyPI for the local web app and command-line interface, installation from source, a Docker container, and a command-line-only distribution. Configuration is stored under data/user/settings/, or at the path indicated by DEEPTUTOR_HOME or deeptutor start --home.

Terminal showing the commands deeptutor init and deeptutor start in bright green text, with structured JSON output scrolling down the screen.

The interface offers conversation, persistent companions, importable agents, cooperative Markdown editing, living books, a knowledge hub, a learning space, memory, and settings.

DeepTutor dashboard panel with floating views for the Knowledge Hub, the Living Book, the Memory graph, and the Learning Space.

The command-line interface provides both an interactive session and structured JSON output so that another agent can drive DeepTutor with the same capabilities, tools, and knowledge bases.

Agent-oriented learning loop shown as a continuous cycle between a conversation interface, a practice quiz, and a document research mode.

Internally, the knowledge layer uses versioned, pluggable retrieval and document-analysis libraries. The README also documents MCP servers, command-line applications, image, video, and voice models, and a sandboxed environment for running generated code for office documents. Version v1.5.9, released on August 4, 2026, added Gemini Embedding 2 via its native endpoint, per-model reasoning-effort control, a Novita AI gateway, retrieval roles for queries, and Compose deployments that preserve data/.

Official and semi-official status

No evidence was found that DeepTutor has been accepted into an official marketplace run by a model provider, or of a certification or endorsement from any company. It should therefore not be described as an official product of OpenAI, Anthropic, Google, GitHub, or HKU.

Its semi-official standing, in a technical and non-commercial sense, comes from two integrations the README declares: it ships with EduHub, the project’s default educational registry, and it is compatible with ClawHub and the open Agent Skills format. This allows importing skills from registries that speak the same format, but it is not equivalent to an endorsement of those registries or a formal standard. The adoption reflected by 32,408 stars shows community interest, not a de facto standard designation backed by a provider.

The ecosystem

The README identifies several HKUDS projects as technical building blocks or inspiration. The following figures come from a GitHub API query on August 4, 2026:

  • HKUDS/LightRAG — 38,507 stars; appears as one of the retrieval options and as inspiration for simple, fast retrieval.
  • HKUDS/AutoAgent — 9,682 stars; the README cites it as a no-code agent framework.
  • HKUDS/AI-Researcher — 5,646 stars; cited as an automated research channel.
  • HKUDS/nanobot — 46,618 stars; the README notes that its lightweight engine powered HKUDS’s original TutorBot.
  • HKUDS/RAG-Anything — 22,632 stars; a project from the same group whose GitHub description presents it as a comprehensive retrieval framework.
  • HKUDS/OpenSpace — 7,274 stars; its description presents it as a skills-management layer for agents.

Open-source ecosystem with the DeepTutor repository as a central core surrounded by satellite nodes representing forks, community guides, and related HKUDS projects such as AutoAgent and AI-Researcher.

No independent public EduHub repository was found in the organization query; the official source documents it as DeepTutor’s integrated registry, not as a separate repository that can be safely attributed.

Community forks, guides, and extensions

The forks API returns 4,233 forks in total. Among the highest-starred retrieved are 0xSojalSec/AI-Tutor (27), hacksider/DeepTutor (12), oalanicolas/deeptutor (12), and cloudtoolbox/deeptutor (7). Their descriptions essentially reproduce DeepTutor’s proposal, so they are classified as forks rather than technically differentiated ports.

Community materials in Chinese were found outside the main fork listing: chencore/deeptutor-guide and chencore/deeptutor-complete-guide. Their descriptions present themselves as complete installation and usage guides for DeepTutor in Chinese, and GitHub search returned 0 stars for each at the time of this measurement. They are community translations or documentation, not official HKUDS repositories.

The README also mentions LlamaIndex, PageIndex, GraphRAG, Obsidian, OpenClaw, Codex, Claude Code, and ManimCat as conceptual dependencies, tools, or inspiration. That mention does not demonstrate that all of them are bundled modules, sponsors, or maintained extensions of DeepTutor.

Repo numbers

Measured: August 4, 2026, via the GitHub API and web page.

MetricValue
Stars32,408
Forks4,233
Real subscribers169
Commits1,186
Open issues per the API88
Issues visible on the page52
Open pull requests visible on the page36
Primary languagePython
LicenseApache-2.0
CreatedDecember 28, 2025
Last metadata updateAugust 4, 2026, 23:53 UTC
Latest releasev1.5.9, August 4, 2026

The top contributors returned by the API were pancacake (562 contributions), Pinkllow (116), github-actions[bot] (108), tusharkhatriofficial (20), scrrlt (16), and wedone (16). The total of 1,186 commits comes from the repository page. watchers_count in GitHub’s general response duplicates the 32,408 stars; that is why subscribers_count is reported here as the real subscriber count. The API’s open_issues_count field can include open pull requests, which explains why it does not necessarily match the 52 issues shown separately.

How to contribute

The repository provides CONTRIBUTING.md and a roadmap where items can be voted on or proposed. The README explicitly points to that guide for branching strategy, code standards, and getting-started steps. The pull request template asks contributors to describe the changes, link related issues, identify affected modules, and confirm that the standards were followed.

There are signs of active collaboration: pull request #13, by kushalgarg101, added Docker support and closed request #3 from lovingfish, which had asked for a maintained, reproducible Compose deployment. The guide and templates are the documented process; it does not follow from this that every proposal will be accepted.

How the community received it

The verifiable reception combines strong GitHub adoption with practical problems reported by users. No extensive discussion directly about HKUDS/DeepTutor was found on Hacker News: submission 47703911, posted by wslh on April 9, 2026, got 2 points and 0 comments. It documents that the repository was discovered, but it does not support inferring praise, criticism, or consensus.

  • Concrete acknowledgment: in issue #351, with 10 comments, StarNight2021 explicitly thanked jiakeboge and pancacake for pull request #319, saying their changes eliminated severe freezing during long conversations. The same issue qualifies the praise: they still observed input lag when typing quickly, so it is not a global evaluation of the product.
  • Performance criticism: in issue #137, with 9 comments, yepyhun reported that the CPU-only Docker image made MinerU’s PDF processing extremely slow; as an example, they cited about 40 seconds per page on an RTX 3070 Ti when using CPU inference. This is one user’s reported experience and configuration, not a universal comparative measurement.
  • Compatibility criticism: in closed issue #15, with 17 comments, tusharkhatriofficial described a Smart Solver failure when a model generated invalid multiline Python strings inside JSON. The report identifies a limitation in the integration with model outputs, not an objection to DeepTutor’s educational purpose.

Benchmark panel showing a circular gauge with a 29.4% improvement in agent reasoning, next to a terminal with a multiline Python string error being fixed by the Smart Solver.

  • Deployment demand: request #3, with 12 comments, asked for prebuilt images and an official Compose setup so users would not need to install Python and Node dependencies on the host machine. Its closure via #13 shows a maintenance response to a community need, but it does not measure subsequent satisfaction.

DeepTutor versus other approaches

ApproachVerifiable overlapVerifiable difference
HKUDS/LightRAGDeepTutor can choose LightRAG as its retrieval engine.LightRAG presents itself as retrieval augmentation; DeepTutor adds tutoring, memory, quiz, and learning-workspace surfaces.
LlamaIndexDeepTutor documents LlamaIndex as an option for its knowledge libraries.In DeepTutor it is an interchangeable retrieval component; the README does not present LlamaIndex as a substitute for its tutoring and memory modes.
GraphRAGBoth relate to graph-based retrieval, per the options the README declares.GraphRAG appears as a selectable engine; DeepTutor additionally coordinates interaction, personalization, and tooling.
HKUDS/AutoAgentBoth are agent projects from the same group, and DeepTutor cites it as inspiration.AutoAgent is described on GitHub as a no-code agent framework; DeepTutor specializes in personalized learning and tutoring.

No independent benchmark comparing these tools on cost, accuracy, or pedagogical outcomes was found. The differences above therefore describe the scope declared by the sources, not a performance ranking.

Use cases and who this repository can help

  • Students and teachers working with their own materials can bring documents, books, quizzes, and a knowledge library into one space, and switch between conversation, practice, research, and domain path without losing the context the README describes.
  • Education teams that need traceable adaptation can use the L1/L2/L3 memory and the memory graph to inspect and edit the basis for a personalization, instead of treating the student profile as a black box.
  • People researching with heterogeneous document collections can choose between LlamaIndex, PageIndex, GraphRAG, LightRAG, or Obsidian, and combine retrieval with tools, MCP, and agents or persistent companions.
  • Administrators and developers who prefer local deployment have paths via PyPI, source, Docker, or the command line; the issue findings suggest it’s particularly worth validating GPU, container, and model configuration before using it with heavy documents.
  • Developers of educational skills can use the Agent Skills format and EduHub to bring in reusable resources, reviewing imported sources before adding them to a learning library.

Agent Skills folder containing a SKILL.md file plugged into an AI agent's head via a glowing fiber-optic port, illuminating its neural pathways.

Resources


Note: this article combines DeepTutor’s README, contribution guide, releases and issues, the GitHub API and page, the arXiv report, and Hacker News, retrieved on August 4, 2026. The figures change over time.

Comments