July 30, 2026 · By YasKad
garrytan/gstack

gstack: a role-based software factory for agents

garrytan/gstack · 134,177★ · 19,991 forks

Everything you need to know about garrytan/gstack: a set of skills and utilities that structures the work of programming agents as a product, engineering, quality, and release cycle.


What is gstack?

gstack is an open-source project by Garry Tan to turn programming agents into a coordinated set of specialists. The repository presents it as 23 tools or commands, mostly written as skills in Markdown, for functions such as product direction, architecture review, design, code review, security, quality control, and release.

It’s not a model or hosted service. It installs locally, detects different agent environments, and exposes commands like /office-hours, /autoplan, /review, /qa, /ship, and /land-and-deploy. Its stated goal is to allow a person to go through an entire process instead of just asking for snippets of code from an assistant.

The origin: publishing the YC president’s workflow

The GitHub API dates the repository creation as March 11, 2026. Its author, Garry Tan, identifies himself in the README as chairman and CEO of Y Combinator; his GitHub profile lists Y Combinator as a company and San Francisco as location.

The launch narrative comes from the README itself. Tan starts with a quote by Andrej Karpathy about stopping manual coding and Peter Steinberger’s ability to build OpenClaw with agents. From there, he presents gstack as his practical response: opening up his own setup, which he claims to use daily, instead of selling a commercial layer. The text argues that in 60 days, he shipped three production services and over 40 features while working part-time on products and running YC; these figures are the author’s assertions, not an independent audit.

The central tension is with using an agent natively as a textbox: Tan proposes chained roles and artifacts so that the session questions the problem, writes a plan, reviews it, tests the result, and prepares deployment. It also addresses a common criticism of assisted programming: the README itself acknowledges that unnormalized code lines are inflated by AI and links to its methodology to defend a measure of logical changes.

Philosophy and principles

The guiding idea is summarized in the README as a sequence: think, plan, build, review, test, ship, and reflect. Each skill delivers an artifact that the next can use: /office-hours drafts a design document, plan reviews refine it, /review detects issues, and /ship verifies before opening a change request.

A glowing circular diagram with the seven stages of the flow — think, plan, build, review, test, ship, and reflect — connected by streams of data.

Its verifiable operating principles are:

  • Reframe vague requests before writing code, with questions that force product and scope decisions.
  • Separate specialties: architecture, design, developer experience, security, testing, and release are not treated as the same task.
  • Require operational evidence: /qa tests an application in a real browser, and /ship reviews tests and coverage before preparing a change request.
  • Turn failures and preferences into local memory: /learn manages learnings per project and session.
  • Prioritize security checks when acting on code or browser: /careful warns about destructive operations, and /freeze limits editable paths.

It’s a methodology with strong opinions, not a quality guarantee. Its results depend on the models, permissions, tests, human reviews, and cost of calls used by each installation.

How it works

The documented base installation clones the repository and runs ./setup; for Claude Code, the suggested path is ~/.claude/skills/gstack. Team mode runs gstack-team-init required, writes configuration in .claude/ and CLAUDE.md, and checks updates limited to once per hour. The optional mode does not block those who don’t use it.

The typical flow can start with /office-hours, continue with /autoplan or plan reviews, implement the change, and close with /review, /qa, and /ship. Some specific functions are:

A holographic specialist seated at a dark glass desk projects a design document with diagrams, invoked by the /office-hours command.

CommandDocumented function
/office-hoursFormulates six product questions, challenges assumptions, and generates a design document.
/autoplanChains direction, design, and engineering reviews and leaves the user with decision points.
/review and /codexReview changes; the second requests a second opinion via Codex CLI.
/qa and /browseOpen Chromium, go through flows, capture visual tests, and generate regression tests when fixing bugs.
/csoApplies a security review based on OWASP Top 10 and STRIDE.
/ship, /land-and-deploy, and /canaryPrepare the change request, wait for CI and deployment, and monitor the application after deployment.
/learn and /retroPreserve learnings per project and perform engineering retrospectives.

A robotic hand manipulates a holographic Chromium browser window while neon markers flag visual bugs and regression tests.

The project also documents useful checkpoints for long runs: the optional checkpoint mode creates local WIP: commits with decisions and pending work; /context-restore can rebuild the state, and /ship compacts those changes before the change request. For multiple work fronts, the README recommends isolated workspaces and describes using Conductor for parallel sessions, but Conductor is not part of gstack.

A crystalline neural structure absorbs fragments of data next to a retrospective panel with charts and engineering notes, representing /learn and /retro.

The installer declares compatibility with Claude Code, Codex CLI, OpenCode, Cursor, Factory Droid, Slate, Kiro, Hermes, and a mode for GBrain. Contribution files specify that templates are generated for eight hosts: Claude, Codex, Factory, Kiro, OpenCode, Slate, Cursor, and OpenClaw. This difference reflects evolving documentation, not certification of each provider.

Official and semi-official status

No evidence was found in the README or GitHub API that gstack has been accepted as a plugin in an official marketplace from Anthropic, OpenAI, Cursor, or Hermes. Its main distribution is the repository itself and the ./setup installer.

There are two semi-official integrations in the technical sense, not institutional: the installer generates skills for compatible hosts, and the README indicates four native OpenClaw skills installable from ClawHub (gstack-openclaw-office-hours, gstack-openclaw-ceo-review, gstack-openclaw-investigate, and gstack-openclaw-retro). Installing a skill from ClawHub does not equal approval by Anthropic, OpenAI, or another manufacturer. With 125,699 stars in the measurement below and numerous derivatives, it can be described as a de facto reference for this style of workflows, but the recovered sources do not attribute a formal standard to it.

The ecosystem

A luminous central monolith with "125k Stars" surrounded by satellite nodes connected by beams of light representing gbrain, alphaclaw, and gstack-auto.

Author’s repositories

  • garrytan/gbrain: persistent knowledge base for OpenClaw and Hermes agents; gstack incorporates /setup-gbrain and /sync-gbrain to initialize it and index repositories. In the consulted API, it had 27,538 stars and 4,036 forks.
  • garrytan/gbrain-evals: repository of evaluations by the same author, with 329 stars and 57 forks in the list of repositories of his account.
  • garrytan/alphaclaw: installation harness for OpenClaw, with 142 stars and 29 forks.
  • garrytan/openclaw-render-template: deployment template for OpenClaw on Render, with 15 stars and 5 forks.

GBrain is the closest functional link: the gstack README allows using local PGLite, Supabase, or a remote MCP server to preserve memory and defines read-write, read-only, or denial policies per repository.

Community forks, ports, and extensions

  • XLearnity/gstack is the most prominent fork returned by the API of forks: 111 stars and 8 forks.
  • kimjin8/gstack-antigravity ports gstack to Google Antigravity: 42 stars and 6 forks.
  • bulyaki/gstackplusplus adapts the approach to C++ development: 20 stars and 8 forks.
  • TMFNK/gstack-OpenCode is an adapter for OpenCode with 21 engineering skills: 5 stars.
  • fustackat/gstack-skill-translations-zh-tw is identified in the GitHub search as a translation of skills into traditional Chinese: 0 stars in the query. It’s a community translation, not by the original author.
  • loperanger7/gstack-auto proposes a semi-automatic orchestration based on gstack: 243 stars and 24 forks.
  • mr-daedalium/ostack-saas declares itself as a fork of gstack for an AI-assisted engineering team: 107 stars and 19 forks.
  • fagemx/gstack-game adapts the methodology to game production: 55 stars and 5 forks.
  • MikeChongCan/cfo-stack applies the idea to accounting and personal finance: 50 stars and 11 forks.

These relationships come from the metadata and descriptions of GitHub recovered in this research. They do not prove shared maintenance, current compatibility, or endorsement by Garry Tan for each derivative.

Repo numbers

Measurement: August 1, 2026, GitHub API.

MetricValue
Stars125,699
Forks18,854
Real subscribers769
Commits360
Open issues indicated by the API860
Main languageTypeScript
LicenseMIT
CreationMarch 11, 2026
Last commit to the repositoryJuly 15, 2026
Last metadata updateAugust 1, 2026
Last documented version1.60.1.0, July 9, 2026
GitHub publicationsnone in the consulted endpoint

The main contributors in the API response were garrytan (319 contributions), test22345 (17), 16francej (7), and time-attack (6). The total of 360 commits comes from the last pagination link of the API. watchers_count replicates the number of stars in the general GitHub response; that’s why subscribers_count is reported as real subscribers. The open_issues_count field may include open change requests, so 860 is not necessarily an exclusive count of issues.

How to contribute

The CONTRIBUTING.md guide documents a detailed contribution flow. You should clone the entire repository, run bun install and bin/dev-setup, which links the working tree so that Claude Code can test local changes immediately. The document advises a full fork, not a shallow clone, to allow git log, git blame, and git bisect.

Skills are edited from SKILL.md.tmpl templates, not from the generated Markdowns. The indicated flow is to modify the template, run bun run gen:skill-docs --host all, check it with bun run skill:check, test the skill in real work, and open a change request from a fork.

A mechanical arm assembles SKILL.md.tmpl files on a glowing pipeline while a terminal in the foreground runs the ./setup command.

The project defines three levels of testing: bun test for free static validation; bun run test:e2e for full execution using claude -p; and bun run test:evals, which combines full execution and model evaluation. The guide warns that the last two consume API and that evaluation children run in an isolated environment to prevent configuration, MCP, or local memory from altering the result. A GitHub action verifies on each push and change request that the generated skill documentation is up-to-date.

How the community received it

The recovered reception mixes interest in the disciplined flow with strong objections to its metrics, cost, and autonomy:

  • The main Hacker News thread, 47418576, was posted by alienreborn on March 17, 2026. The Algolia API recovered assigns it 74 points; the reference comment in another post cites 87 comments, but the detail endpoint did not return a comment count and therefore that figure is not presented as confirmed measurement. josh2600 praised that drafts of answers to product and engineering questions improved their quality and development speed. Conversely, MaxLeiter considered the planning mode interesting, although heavy on tokens; rileymichael questioned whether code lines are a useful metric, and input_sh asked about monthly API costs and possible discounts. These are opinions from participants, not independent measurements.
  • In the same thread, observationist called it a powerful setup but argued that better custom tools exist for specific tasks. This criticism goes to scope: the project combines many functions and may not be the simplest option for an isolated problem.
  • The post 47668746, published by thisisfatih, presents tonone-ai/tonone as inspired by gstack. It had 3 points and 1 comment. Its author stated that assigning a single function to each agent reduced context and token consumption and improved output; it’s evidence of an inspired extension, not a controlled evaluation of gstack.
  • The post 47355173, by jumploops, reached 15 points and 15 comments according to the Algolia search. zippolyon summarized the operational reservation clearly: speed without limits is dangerous when the agent works autonomously between repositories. That user linked his own product as a response, so it should be read as an interested recommendation, not proof of a gstack failure.

A neon AI entity scans a floating block of code while red warning glyphs and holographic shields flag OWASP and STRIDE findings.

gstack versus other proposals

ProposalVerifiable matchVerifiable difference
tonone-ai/tononeIts author states that it was inspired by gstack and also organizes agents by functions.The author of Tonone claims to have created product and engineering teams with leaders and subagents, as well as its own marketplace; gstack is distributed as local skills and utilities.
browser-use/browser-harness-jsBoth allow an agent to interact with Chrome using the DevTools protocol.The gstack README presents it as a thin transport alternative without controls; gstack uses a whitelist, mutual blocking, and untrusted output wrapping for sensitive operations.
forrestchang/andrej-karpathy-skillsThe gstack README cites it as a set of rules for assisted programming failures.gstack positions itself as a layer that enforces a multi-stage flow during a sprint, while the README only attributes to the other project rules about assumptions, complexity, external changes, and declarative goals.
loperanger7/gstack-autoDeclares orchestration based on gstack.It is presented as a semi-automatic extension that starts from a specification; it is not the original repository nor an official verified integration.

The practical difference is not just the number of agents: gstack tries to link decisions, tests, quality control, and release through state files and commands. A browser tool or a rule library may be more suitable if only one of those layers is needed.

Use cases and who this repository can help

  • People using agents to take a function from idea to delivery can chain /office-hours to clarify the product, /autoplan to review the plan, /review to inspect changes, /qa to go through the application in Chromium, and /ship to check the status before opening a change request.
  • Teams that want to make controls visible during long tasks can use WIP: checkpoints and /context-restore to recover decisions and pending work. In delivery flows, /land-and-deploy waits for CI and deployment, while /canary monitors the application after; /careful and /freeze provide restrictions for sensitive operations.
  • Documentation and project learning maintainers can use /document-release to compare changes with READMEs, architecture, and guides, /document-generate to structure documentation with Diataxis, and /learn or /retro to preserve patterns and problems per project. The skill templates themselves can be validated with bun run skill:check, full tests, and model evaluations.

Gstack organizes a workflow and its local artifacts; it does not replace branch protections, CI/CD policies, or permission reviews. Full runs and evaluations consume API, so it is advisable to measure cost and results before extending the flow to an entire team.

Resources


Note: this article combines the README, contribution guide, and change history of gstack, the GitHub and Hacker News APIs consulted on August 1, 2026. The figures change over time.

Comments