August 11, 2026 · By YasKad
karpathy/autoresearch

autoresearch: an overnight lab for an agent to iterate on a model

karpathy/autoresearch · 96,768★ · 13,524 forks

Everything you need to know about karpathy/autoresearch: a minimal language-model training experiment in which an agent modifies a single file, measures the result, and keeps or reverts each attempt.


What autoresearch is

autoresearch is a repository by Andrej Karpathy that lets an AI agent run language-model training experiments for hours. The official description sums it up as automated research on training nanochat on a single GPU; the README spells out the cycle: modify code, train for five minutes, measure, keep or discard, and repeat.

It does not try to be a general AutoML system or a distributed lab product. It is a deliberate reduction of the problem to a small, measurable environment: one NVIDIA GPU, one GPT model, one metric (val_bpb, validation bits per byte; lower is better), and a controlled surface for modification.

Futuristic circular diagram showing the four-step cycle: modify code, train for five minutes, measure the metric, and keep or revert.

The origin: a swarm fiction and a minimal prototype

GitHub records the creation of karpathy/autoresearch on March 6, 2026. Its author is Andrej Karpathy (karpathy), whose GitHub account describes him as someone interested in training deep neural nets on large datasets, with a profile based at Stanford. The README is narratively dated March 2026: it opens by imagining a future of autonomous agent swarms modifying code beyond human comprehension, and presents this repository as the beginning of that story.

The project is born as a simplification of karpathy/nanochat: the README states explicitly that the training code is a reduced version of that repository. The tension it tries to resolve is not a gap in any specific agent, but the cost of doing iterative research by hand: the human stops editing Python files each round and instead designs program.md, the document that gives the agent context and operating instructions.

The first Hacker News submission, 47291123, linked the repository and reached 208 points and 20 comments. The discussion that would follow the launch was already present there: whether this represents agent-guided research or a form of hyperparameter tuning wrapped in a natural-language loop.

Philosophy and principles

The documented design rests on very concrete principles:

  • A single success variable. val_bpb is the source of truth and is compared under a fixed five-minute budget; the README argues it does not depend on vocabulary size and makes it easier to compare architectural changes.
  • Narrow-scope changes. The agent only modifies train.py; prepare.py, which holds preparation, data loading, evaluation, and constants, is left untouched.
  • Comparability on the same machine, not across machines. Every run consumes the same training time, but results are optimized for the specific platform and are not directly comparable to those from another GPU.
  • Keep evidence, revert the rest. Every attempt is recorded with a Git commit and in results.tsv; a result that is worse or equal rolls the branch back to its previous state.
  • Simplicity as an additional criterion. program.md asks the agent to weigh complexity: a small improvement that adds fragile code may not be worth keeping, while simplifying without hurting the metric does count as progress.

Neon visualization of the val_bpb metric descending as a single bar of light, with alternative metrics crossed out in the background.

This turns the repository into an example of bounded empirical optimization, not a demonstration that the agent produces generalizable scientific knowledge on its own.

Git tree stylized as a neon cityscape, with a robotic hand editing train.py while prepare.py stays shielded behind a red barrier.

How it works

The main flow has three pieces:

FileDocumented role
prepare.pyDownloads data, trains the BPE tokenizer, and provides data loading and evaluation; read-only for the agent.
train.pyContains the GPT model, the Muon + AdamW optimizer, and the training loop; the only file the agent may modify.
program.mdInstructions for the agent; the file the human edits to define the research organization.

The documented manual start is uv sync, uv run prepare.py, and uv run train.py. It requires Python 3.10 or newer, uv, and an NVIDIA GPU; the README states it was tested on H100. After that, Claude, Codex, or another agent can be launched in the repository, unpermissioned, and asked to read program.md and set up an experiment.

Split illustration: a human hand writes program.md on a digital tablet while an AI entity reads the instructions and manipulates a matrix of code.

program.md prescribes a precise protocol: create an autoresearch/<label> branch, check the cache, initialize results.tsv, run the baseline first, and then repeat edit → commit → uv run train.py > run.log 2>&1 → extract val_bpb and VRAM → log → keep or revert. A run that exceeds ten minutes is treated as a failure. If a change fails from out-of-memory or a code error, it is logged as crash; the document asks the agent to keep going autonomously until a person interrupts it.

The recovered train.py confirms a PyTorch GPT implementation for a single GPU, with Flash Attention 3 and a kernel-repository choice based on CUDA capability. This is not a portability claim: the README states that, in its main form, the code requires an NVIDIA GPU and prefers not to expand the core with CPU or MPS support.

Neon GPU cluster centered on a single NVIDIA H100, with holographic readouts showing VRAM usage and a five-minute countdown.

Official and semi-official status

No evidence was retrieved that autoresearch has been accepted into an official vendor marketplace, or of a formal certification or endorsement from Anthropic, OpenAI, NVIDIA, or another company. The README does name Claude and Codex as agents that can run the protocol, but that describes usage compatibility, not an official integration.

Its practical status is that of a de facto reference for a pattern: the GitHub search retrieved numerous extensions explicitly described as Karpathy-inspired, and the repository documents notable ports. That spread does not amount to formal standardization, nor does it validate the quality of results produced by each derivative.

The ecosystem

  • karpathy/nanochat is the technical parent repository according to the README; autoresearch simplifies its training. The GitHub API showed it with 56,896 stars and 7,876 forks.
  • karpathy/nanoGPT, described by its author as a simple, fast repository for training or fine-tuning medium-sized GPT models, had 61,807 stars and 10,638 forks.
  • karpathy/llm.c, language-model training in C/CUDA, had 30,711 stars and 3,716 forks.

These are projects from the same author, not components required to run autoresearch; the narrative and technical dependency the README directly confirms is with nanochat.

Ports, forks, and community extensions

The official README highlights four ports: miolini/autoresearch-macos for macOS (2,349 stars, 336 forks); trevin-creator/autoresearch-mlx, an Apple Silicon/MLX port found in the search (1,775, 359); jsegov/autoresearch-win-rtx for Windows (703, 143); and andyluo7/autoresearch for AMD (68, 17).

Futuristic network map with the central karpathy/autoresearch node connected to satellite nodes representing the macOS, Apple Silicon, Windows, and AMD ports.

The fork search also retrieved:

  • sanbuphy/autoresearch-cn, a Chinese-language adaptation for training and fine-tuning a language model with an agent on one GPU (190 stars, 18 forks).
  • mishig25/hf-autoresearch, a fork for Hugging Face infrastructure (177, 17).
  • eli-labz/ResearchSwarm, described as an autonomous swarm for roughly a hundred overnight experiments and as a fork of karpathy/autoresearch (359, 1).
  • ncdrone/autoresearch-ANE, a variant for Apple’s Neural Engine (57, no forks in the retrieved response).

The global GitHub search also surfaces tools that extend the pattern rather than fork the base code: davebcn87/pi-autoresearch (7,307 stars, 431 forks) as an extension of the loop for Pi; uditgoenka/autoresearch (5,664, 429) as a skill for Claude Code; leo-lilinxiao/codex-autoresearch (2,045, 117) as a skill for Codex; RightNow-AI/autokernel (1,496, 155) for optimizing Triton kernels; and evo-hq/evo (1,358, 102) for turning a codebase into a measure-and-search loop with subagents. Their descriptions come from the GitHub search; no compatibility, sponsorship, or quality is inferred beyond what they state.

The expansion has also produced catalogs, such as webfuse-com/awesome-autoresearch (2,345, 176) and WecoAI/awesome-autoresearch (1,027, 75), which serve as an ecosystem signal, not an official project source.

Repo numbers

Measured: August 3, 2026, GitHub API.

MetricValue
Stars92,838
Forks13,218
Real subscribers721
Commits36
Open issues reported by the API196
Primary languagePython
LicenseNot indicated by the API
CreatedMarch 6, 2026
Latest code pushMarch 26, 2026
Latest metadata updateAugust 3, 2026
GitHub releasesNone retrieved

The top contributors returned by the API are karpathy with 28 contributions, followed by nishantpurohit04, dipeshbabu, hughdbrown, marcinbogdanski, dumko2001, haosenwang1018, indianspeedster, and kaizen-38, each with one contribution in the retrieved list. The total of 36 commits comes from the last-page link of the API pagination.

Two caveats apply: watchers_count in the general response mirrors the star count, which is why subscribers_count is reported here as the real subscriber figure. Also, open_issues_count may include open pull requests, so 196 is not necessarily an issues-only count. While updated_at matches the date of this measurement, the latest code push the API returns is from March 26, 2026; a metadata update does not prove later code activity.

How to contribute

No contribution file or human guide for branches, tests, or a pull-request template was retrieved. The API allows forking the repository and accepts pull requests, and the README invites creating forks or discussions for other platforms; that describes an open path, not a guaranteed maintenance process.

To contribute to the experiment as designed, program.md does define the agent’s procedure: a new branch, exclusive editing of train.py, a commit for every attempt, an unversioned log in results.tsv, and reverting results that are not better. Anyone proposing changes to the repository should separately review current discussions and pull requests, since no official acceptance policy was retrieved.

How the community received it

The retrieved reception combines technical curiosity, reuse, and concrete methodological criticism:

  • In the initial thread 47291123, submitted by simonpure (208 points, 20 comments), mikert89 praised environments where a model learns by trial and error with objective verification. By contrast, abeppu asked whether the apparent improvements were actually hyperparameter changes and asked for a comparison against BayesOpt with the same number of five-minute trials. gmerc objected that confusing brute-force search with research turns the metric into a possible case of Goodhart’s law. elikoga flagged the risk that the best five-minute bpb might favor models too small for certain emergent behaviors.

Conceptual illustration of a robotic eye watching a shifting mathematical formula, surrounded by holographic warnings about BayesOpt and Goodhart's law.

  • SkyPilot’s scaling analysis was discussed in 47442435, submitted by hopechong (237 points, 17 comments retrieved; the story search reported 94 comments). zhwu noted that the agent, without being told to, used H100 to filter ideas and H200 to validate candidates. augment_me’s objection was that the comparison favored wall-clock time: at equal GPU-hours, sequential execution appeared roughly twice as efficient. herf questioned whether selecting on five-minute early progress could hurt the asymptotic result. These are participant opinions, not results from an independent evaluation.
  • The practical spread is also visible in 47343935: austinbaggio presented Autoresearch@home and got 79 points and 19 comments. It is evidence of later derivation, not proof that the base repository outperforms other optimizers.

autoresearch versus other proposals

ProposalVerifiable overlapVerifiable difference
karpathy/nanochatBoth belong to Karpathy, and the README confirms autoresearch simplifies its training.nanochat is the parent training project; autoresearch adds the agent protocol, the instruction file, and the keep/revert rule.
davebcn87/pi-autoresearchThe search describes it as an autonomous extension of the experiment loop.It targets Pi; it is not the single-GPU training environment the original repository documents.
RightNow-AI/autokernelThe search describes it as applying the pattern to automatic iterations against a measurable objective.It optimizes PyTorch models’ Triton kernels, while the original modifies the model and training loop itself.
evo-hq/evoBoth use change, measurement, and repetition to improve code.Evo states codebase instrumentation and tree search with subagents; autoresearch fixes a single agent, one editable file, and a linear keep-or-revert path.

The central comparison is not between conversational assistants, but between search mechanisms. autoresearch contributes explicit constraints so that a natural-language exploration is auditable on a local metric; it does not by itself demonstrate that it is more efficient than BayesOpt, random search, or other optimization methods. That comparison was precisely an explicit criticism from the community.

Use cases and who this repository can help

  • Researchers and model-engineering people with an NVIDIA GPU can turn a night of compute into a sequence of trials on architecture, optimizer, batch size, or hyperparameters, as long as they accept val_bpb under five minutes as the local objective.
  • Teams wanting to prototype a verifiable agent harness can study the split between program.md, the editable surface train.py, the locked evaluation in prepare.py, and the results.tsv log. It is a concrete template for keeping the agent from changing both the experiment and the meter at once.
  • macOS, Windows, AMD, or Apple Silicon users can start from the ports the README or the fork search identifies, but should treat their metrics as machine-specific and review the derivative’s code before adopting it.
  • Anyone optimizing other quantifiable domains can take the keep/revert pattern as inspiration, as the extensions for Codex, kernels, or codebases do. They need to define a manipulation-resistant metric and independent checks beforehand: the community criticism shows that a single early marker can incentivize narrow or misleading improvements.

Futuristic terminal running uv run train.py, with holographic panels showing a new Git branch being created and the run.log file filling with data.

Resources


Note: this article combines the README, program.md, and source code retrieved from karpathy/autoresearch, the GitHub API, repository searches, and Hacker News, consulted on August 3, 2026. Figures change over time.

Comments