gstack: a role-based software factory for agents
garrytan/gstack · 134,177★ · 19,991 forks
Everything you need to know about garrytan/gstack: a set of skills and utilities that structures the work of programming agents as a product, engineering, quality, and release cycle.
What is gstack?
gstack is an open-source project by Garry Tan to turn programming agents into a coordinated set of specialists. The repository presents it as 23 tools or commands, mostly written as skills in Markdown, for functions such as product direction, architecture review, design, code review, security, quality control, and release.
It’s not a model or hosted service. It installs locally, detects different agent environments, and exposes commands like /office-hours, /autoplan, /review, /qa, /ship, and /land-and-deploy. Its stated goal is to allow a person to go through an entire process instead of just asking for snippets of code from an assistant.
The origin: publishing the YC president’s workflow
The GitHub API dates the repository creation as March 11, 2026. Its author, Garry Tan, identifies himself in the README as chairman and CEO of Y Combinator; his GitHub profile lists Y Combinator as a company and San Francisco as location.
The launch narrative comes from the README itself. Tan starts with a quote by Andrej Karpathy about stopping manual coding and Peter Steinberger’s ability to build OpenClaw with agents. From there, he presents gstack as his practical response: opening up his own setup, which he claims to use daily, instead of selling a commercial layer. The text argues that in 60 days, he shipped three production services and over 40 features while working part-time on products and running YC; these figures are the author’s assertions, not an independent audit.
The central tension is with using an agent natively as a textbox: Tan proposes chained roles and artifacts so that the session questions the problem, writes a plan, reviews it, tests the result, and prepares deployment. It also addresses a common criticism of assisted programming: the README itself acknowledges that unnormalized code lines are inflated by AI and links to its methodology to defend a measure of logical changes.
Philosophy and principles
The guiding idea is summarized in the README as a sequence: think, plan, build, review, test, ship, and reflect. Each skill delivers an artifact that the next can use: /office-hours drafts a design document, plan reviews refine it, /review detects issues, and /ship verifies before opening a change request.

Its verifiable operating principles are:
- Reframe vague requests before writing code, with questions that force product and scope decisions.
- Separate specialties: architecture, design, developer experience, security, testing, and release are not treated as the same task.
- Require operational evidence:
/qatests an application in a real browser, and/shipreviews tests and coverage before preparing a change request. - Turn failures and preferences into local memory:
/learnmanages learnings per project and session. - Prioritize security checks when acting on code or browser:
/carefulwarns about destructive operations, and/freezelimits editable paths.
It’s a methodology with strong opinions, not a quality guarantee. Its results depend on the models, permissions, tests, human reviews, and cost of calls used by each installation.
How it works
The documented base installation clones the repository and runs ./setup; for Claude Code, the suggested path is ~/.claude/skills/gstack. Team mode runs gstack-team-init required, writes configuration in .claude/ and CLAUDE.md, and checks updates limited to once per hour. The optional mode does not block those who don’t use it.
The typical flow can start with /office-hours, continue with /autoplan or plan reviews, implement the change, and close with /review, /qa, and /ship. Some specific functions are:

| Command | Documented function |
|---|---|
/office-hours | Formulates six product questions, challenges assumptions, and generates a design document. |
/autoplan | Chains direction, design, and engineering reviews and leaves the user with decision points. |
/review and /codex | Review changes; the second requests a second opinion via Codex CLI. |
/qa and /browse | Open Chromium, go through flows, capture visual tests, and generate regression tests when fixing bugs. |
/cso | Applies a security review based on OWASP Top 10 and STRIDE. |
/ship, /land-and-deploy, and /canary | Prepare the change request, wait for CI and deployment, and monitor the application after deployment. |
/learn and /retro | Preserve learnings per project and perform engineering retrospectives. |

The project also documents useful checkpoints for long runs: the optional checkpoint mode creates local WIP: commits with decisions and pending work; /context-restore can rebuild the state, and /ship compacts those changes before the change request. For multiple work fronts, the README recommends isolated workspaces and describes using Conductor for parallel sessions, but Conductor is not part of gstack.

The installer declares compatibility with Claude Code, Codex CLI, OpenCode, Cursor, Factory Droid, Slate, Kiro, Hermes, and a mode for GBrain. Contribution files specify that templates are generated for eight hosts: Claude, Codex, Factory, Kiro, OpenCode, Slate, Cursor, and OpenClaw. This difference reflects evolving documentation, not certification of each provider.
Official and semi-official status
No evidence was found in the README or GitHub API that gstack has been accepted as a plugin in an official marketplace from Anthropic, OpenAI, Cursor, or Hermes. Its main distribution is the repository itself and the ./setup installer.
There are two semi-official integrations in the technical sense, not institutional: the installer generates skills for compatible hosts, and the README indicates four native OpenClaw skills installable from ClawHub (gstack-openclaw-office-hours, gstack-openclaw-ceo-review, gstack-openclaw-investigate, and gstack-openclaw-retro). Installing a skill from ClawHub does not equal approval by Anthropic, OpenAI, or another manufacturer. With 125,699 stars in the measurement below and numerous derivatives, it can be described as a de facto reference for this style of workflows, but the recovered sources do not attribute a formal standard to it.
The ecosystem

Author’s repositories
garrytan/gbrain: persistent knowledge base for OpenClaw and Hermes agents; gstack incorporates/setup-gbrainand/sync-gbrainto initialize it and index repositories. In the consulted API, it had 27,538 stars and 4,036 forks.garrytan/gbrain-evals: repository of evaluations by the same author, with 329 stars and 57 forks in the list of repositories of his account.garrytan/alphaclaw: installation harness for OpenClaw, with 142 stars and 29 forks.garrytan/openclaw-render-template: deployment template for OpenClaw on Render, with 15 stars and 5 forks.
GBrain is the closest functional link: the gstack README allows using local PGLite, Supabase, or a remote MCP server to preserve memory and defines read-write, read-only, or denial policies per repository.
Community forks, ports, and extensions
XLearnity/gstackis the most prominent fork returned by the API of forks: 111 stars and 8 forks.kimjin8/gstack-antigravityports gstack to Google Antigravity: 42 stars and 6 forks.bulyaki/gstackplusplusadapts the approach to C++ development: 20 stars and 8 forks.TMFNK/gstack-OpenCodeis an adapter for OpenCode with 21 engineering skills: 5 stars.fustackat/gstack-skill-translations-zh-twis identified in the GitHub search as a translation of skills into traditional Chinese: 0 stars in the query. It’s a community translation, not by the original author.loperanger7/gstack-autoproposes a semi-automatic orchestration based on gstack: 243 stars and 24 forks.mr-daedalium/ostack-saasdeclares itself as a fork of gstack for an AI-assisted engineering team: 107 stars and 19 forks.fagemx/gstack-gameadapts the methodology to game production: 55 stars and 5 forks.MikeChongCan/cfo-stackapplies the idea to accounting and personal finance: 50 stars and 11 forks.
These relationships come from the metadata and descriptions of GitHub recovered in this research. They do not prove shared maintenance, current compatibility, or endorsement by Garry Tan for each derivative.
Repo numbers
Measurement: August 1, 2026, GitHub API.
| Metric | Value |
|---|---|
| Stars | 125,699 |
| Forks | 18,854 |
| Real subscribers | 769 |
| Commits | 360 |
| Open issues indicated by the API | 860 |
| Main language | TypeScript |
| License | MIT |
| Creation | March 11, 2026 |
| Last commit to the repository | July 15, 2026 |
| Last metadata update | August 1, 2026 |
| Last documented version | 1.60.1.0, July 9, 2026 |
| GitHub publications | none in the consulted endpoint |
The main contributors in the API response were garrytan (319 contributions), test22345 (17), 16francej (7), and time-attack (6). The total of 360 commits comes from the last pagination link of the API. watchers_count replicates the number of stars in the general GitHub response; that’s why subscribers_count is reported as real subscribers. The open_issues_count field may include open change requests, so 860 is not necessarily an exclusive count of issues.
How to contribute
The CONTRIBUTING.md guide documents a detailed contribution flow. You should clone the entire repository, run bun install and bin/dev-setup, which links the working tree so that Claude Code can test local changes immediately. The document advises a full fork, not a shallow clone, to allow git log, git blame, and git bisect.
Skills are edited from SKILL.md.tmpl templates, not from the generated Markdowns. The indicated flow is to modify the template, run bun run gen:skill-docs --host all, check it with bun run skill:check, test the skill in real work, and open a change request from a fork.

The project defines three levels of testing: bun test for free static validation; bun run test:e2e for full execution using claude -p; and bun run test:evals, which combines full execution and model evaluation. The guide warns that the last two consume API and that evaluation children run in an isolated environment to prevent configuration, MCP, or local memory from altering the result. A GitHub action verifies on each push and change request that the generated skill documentation is up-to-date.
How the community received it
The recovered reception mixes interest in the disciplined flow with strong objections to its metrics, cost, and autonomy:
- The main Hacker News thread, 47418576, was posted by alienreborn on March 17, 2026. The Algolia API recovered assigns it 74 points; the reference comment in another post cites 87 comments, but the detail endpoint did not return a comment count and therefore that figure is not presented as confirmed measurement. josh2600 praised that drafts of answers to product and engineering questions improved their quality and development speed. Conversely, MaxLeiter considered the planning mode interesting, although heavy on tokens; rileymichael questioned whether code lines are a useful metric, and input_sh asked about monthly API costs and possible discounts. These are opinions from participants, not independent measurements.
- In the same thread, observationist called it a powerful setup but argued that better custom tools exist for specific tasks. This criticism goes to scope: the project combines many functions and may not be the simplest option for an isolated problem.
- The post 47668746, published by thisisfatih, presents
tonone-ai/tononeas inspired by gstack. It had 3 points and 1 comment. Its author stated that assigning a single function to each agent reduced context and token consumption and improved output; it’s evidence of an inspired extension, not a controlled evaluation of gstack. - The post 47355173, by jumploops, reached 15 points and 15 comments according to the Algolia search. zippolyon summarized the operational reservation clearly: speed without limits is dangerous when the agent works autonomously between repositories. That user linked his own product as a response, so it should be read as an interested recommendation, not proof of a gstack failure.

gstack versus other proposals
| Proposal | Verifiable match | Verifiable difference |
|---|---|---|
tonone-ai/tonone | Its author states that it was inspired by gstack and also organizes agents by functions. | The author of Tonone claims to have created product and engineering teams with leaders and subagents, as well as its own marketplace; gstack is distributed as local skills and utilities. |
browser-use/browser-harness-js | Both allow an agent to interact with Chrome using the DevTools protocol. | The gstack README presents it as a thin transport alternative without controls; gstack uses a whitelist, mutual blocking, and untrusted output wrapping for sensitive operations. |
forrestchang/andrej-karpathy-skills | The gstack README cites it as a set of rules for assisted programming failures. | gstack positions itself as a layer that enforces a multi-stage flow during a sprint, while the README only attributes to the other project rules about assumptions, complexity, external changes, and declarative goals. |
loperanger7/gstack-auto | Declares orchestration based on gstack. | It is presented as a semi-automatic extension that starts from a specification; it is not the original repository nor an official verified integration. |
The practical difference is not just the number of agents: gstack tries to link decisions, tests, quality control, and release through state files and commands. A browser tool or a rule library may be more suitable if only one of those layers is needed.
Use cases and who this repository can help
- People using agents to take a function from idea to delivery can chain
/office-hoursto clarify the product,/autoplanto review the plan,/reviewto inspect changes,/qato go through the application in Chromium, and/shipto check the status before opening a change request. - Teams that want to make controls visible during long tasks can use
WIP:checkpoints and/context-restoreto recover decisions and pending work. In delivery flows,/land-and-deploywaits for CI and deployment, while/canarymonitors the application after;/carefuland/freezeprovide restrictions for sensitive operations. - Documentation and project learning maintainers can use
/document-releaseto compare changes with READMEs, architecture, and guides,/document-generateto structure documentation with Diataxis, and/learnor/retroto preserve patterns and problems per project. The skill templates themselves can be validated withbun run skill:check, full tests, and model evaluations.
Gstack organizes a workflow and its local artifacts; it does not replace branch protections, CI/CD policies, or permission reviews. Full runs and evaluations consume API, so it is advisable to measure cost and results before extending the flow to an entire team.
Resources
- Repository: https://github.com/garrytan/gstack
- Documentation and installation: https://github.com/garrytan/gstack#quick-start
- Skills and architecture: https://github.com/garrytan/gstack/blob/main/docs/skills.md and https://github.com/garrytan/gstack/blob/main/ARCHITECTURE.md
- Contribution and tests: https://github.com/garrytan/gstack/blob/main/CONTRIBUTING.md
- Complementary GBrain repository: https://github.com/garrytan/gbrain
- Native OpenClaw skills: https://github.com/garrytan/gstack#native-openclaw-skills-via-clawhub
- Conversations and reviews: https://news.ycombinator.com/item?id=47418576, https://news.ycombinator.com/item?id=47355173, https://news.ycombinator.com/item?id=47668746
Note: this article combines the README, contribution guide, and change history of gstack, the GitHub and Hacker News APIs consulted on August 1, 2026. The figures change over time.
Comments