Let AI write its own paper: How Karpathy’s autoresearch trained a better model in a weekend
AI Agent, automatic research, model training, Karpathy, open source projects, LLM
Slug
karpathy-autoresearch-autonomous-ai-research
GitHub URL
https://github.com/karpathy/autoresearch
Stable Diffusion Prompt
A futuristic laboratory scene with autonomous AI agents in a glowing server room, multiple screens showing training loss curves and neural network architectures, a scientist sleeping peacefully in a chair while AI agents work around them, holographic displays with code and training metrics, cyberpunk color palette with blues and purples, cinematic lighting, digital art style
Cover processing
Use Stable Diffusion to generate a cover, and update the Notion Cover field after uploading to the image bed.
Let AI write its own paper: How Karpathy’s autoresearch trains a better model in a weekend
What is autoresearch?
In March 2026, Andrej Karpathy tweeted on X:
One day, cutting-edge AI research will be completed by computers made of meat that eat, sleep, and play, and occasionally synchronize through the sonic interconnection ritual of "group meetings." That era has passed. Research now belongs entirely to swarms of autonomous AI agents running on megastructures of computer clusters in the sky.
Then he lost three files: prepare.py, train.py, program.md.
These three files are the entire autoresearch - an open source project that allows AI agents to conduct machine learning research independently.
Core Concept
The traditional research process is as follows:
1. You come up with an idea
1. You change the code
1. You train the model
1. Look at the results
1. You judge good or bad
1. Repeat
autoresearch automates this entire process. You set up program.md (equivalent to the instruction file for the agent), and then the agent will:
- Modify
train.py(change model architecture, hyperparameters, optimizer, batch size)
- Training 5 minutes (fixed time budget)
- Evaluation val_bpb (validation set bits per byte, the lower the better)
- Judge whether the results have improved
- Keep or Discard Modify
- Repeat - More than 100 experiments can be done in a day
When you wake up in the morning, what you see is not a model, but a complete experiment log.
Why are there only three files?
The beauty of autoresearch is its minimalist design. It is intentionally kept to a minimum scope:
prepare.py — Prepare once and for all
This file is responsible for:
- Download training data set
- Training BPE tokenizer
- Provides data loader and evaluation tools
Rule: Never change this file. It is a set-and-forget infrastructure.
train.py — agent’s only battlefield
This single file contains:
- Complete GPT model definition
- Muon + AdamW optimizer
- Complete training cycle
The agent will modify every line of this file - model depth, attention mode, learning rate, batch size, all of which are the agent's experimental space.
program.md — Directives from the Research Director
This is the most critical document. You're not writing Python, you're writing about the culture of your organization. program.md defines:
- The agent's role and goals
- Experimental evaluation criteria
- How to think and make decisions
- How to record experimental results
This is a new programming paradigm - you no longer write code directly, but write instructions that let the agent write code.
Technical details
Fixed five-minute budget
autoresearch made a smart design decision: Strictly limit each training session to 5 minutes.
This means:
- Whether you use H100 or RTX 4090, the experimental results can be directly compared
- Approximately 12 experiments per hour, approximately 100 per night
- All experimental results are compared fairly on the same baseline
val_bpb: fair comparison baseline
Bits per byte (bpb) is a metric that is independent of vocabulary size. This means that when the agent changes the model architecture (for example, from 8 layers to 4 layers), the evaluation results will still be fair and comparable.
Muon Optimizer
autoresearch uses the Muon optimizer - a relatively new optimizer that generally performs better than AdamW for the same training time. This also reflects the philosophy of autoresearch: let the agent explore what is optimal, rather than you specifying it in advance.
How to run in your environment
Prerequisites
- Single NVIDIA GPU (officially tested on H100)
-Python 3.10+
- uv package manager
Startup steps
# 1. Install uv (if not installed yet)
curl -LsSf https://astral.sh/uv/install.sh | sh
# 2. Synchronization dependencies
uv sync
# 3. Prepare information (one-time, about 2 minutes)
uv run prepare.py
# 4. Run a training manually (about 5 minutes, confirm that the environment is normal)
uv run train.py
Start independent research mode
Now, start your Claude, Codex or other coding agent:
Hi, have a look at program.md and let's kick off a new experiment!
Let's do the setup first.
Then you just have to watch. The Agent will read program.md by itself, start modifying train.py, train the model, evaluate the results, and then start the next experiment.
Runs on consumer grade hardware
Don’t have an H100? no problem. Karpathy offers some advice:
- Dataset: Using TinyStories (children's stories generated by GPT-4), entropy is lower, and small models can be trained
- Vocabulary size: reduced from 8192 to 2048 or even 1024
- Sequence Length: Depending on your hardware, it can be as low as 256
- Model Depth: reduced from 8 layers to 4 layers
- Batch size: reduced to
2**14(~16K tokens)
The community has also launched multiple forks to support MacOS (MLX), Windows (RTX), AMD GPU and other platforms.
Why is this project important?
autoresearch is more than just a tool. It represents a research paradigm shift:
From "writing code" to "arranging agent"
In the past, researchers spent 80% of their time adjusting parameters, running experiments, and looking at results. autoresearch changes this process from "manual" to "automated". The role of the researcher changes from "executor" to "orchestrator" - you design the research organization and the agent is responsible for execution.
Human intuition + agent’s physical strength
What humans are best at is asking good questions and designing experimental directions. What AI agents are best at is a lot of repetitive work and fine-tuning. autoresearch combines the two: human intuition (research policy in program.md) + agent's physical strength (100+ experiments/night).
Reproducible independent research
Traditional autonomous research (such as Google's AlphaDev) is closed. autoresearch is open source and reproducible. Anyone can build their own autonomous research system based on it.
Actual case: What did autoresearch find?
In early runs of autoresearch, the agent discovered the following improvements:
1. Attention mode optimization: The agent automatically selects a more efficient attention mode (such as SSSL with alternating attention) instead of the standard Transformer structure preset by humans.
1. Optimizer combination: The combination of Muon + AdamW was found by the agent to be better than using either optimizer alone.
1. Learning rate scheduling: The agent automatically adjusts the learning rate scheduling strategy and has a more strategic attenuation mode in the later stages of training.
These are not the results of manual attempts by human researchers, but are automatically discovered by the agent in 100+ experiments.
Comparison with related projects
- Differences in project positioning
- ------------------
Core Concept
The traditional research process is as follows:
1. You come up with an idea
1. You change the code
1. You train the model
1. Look at the results
1. You judge good or bad
1. Repeat
autoresearch automates this entire process. you put
Program.md (equivalent to the instruction file for the agent) is set, and then the agent will:
- Modify train.py (change model architecture, hyperparameters, optimizer, batch size)
- Training for 5 minutes (fixed time budget)
- Evaluate val_bpb (validation set bits per byte, lower is better)
- Determine whether the results have improved
- Keep or discard Modify
- Repeat - can do more than 100 experiments in a day
When you wake up in the morning, what you see is not a model, but a complete experiment log.
Why are there only three files?
The beauty of autoresearch is its minimalist design. It is intentionally kept to a minimum scope:
prepare.py — Prepare once and for all
This file is responsible for:
- Download training data set
- Training BPE tokenizer
- Provides data loader and evaluation tools
Rule: Never change this file. It's set-and-forget infrastructure.
train.py — agent’s only battlefield
This single file contains:
- Complete GPT model definition
- Muon + AdamW optimizer
- Complete training cycle
The agent will modify every line of this file - model depth, attention mode, learning rate, batch size, all of which are the agent's experimental space.
program.md — Directives from the Research Director
This is the most critical document. You're not writing Python, you're writing about studying the culture of an organization.
program.md defines:
- The agent's role and goals
- Experimental evaluation criteria
- How to think and make decisions
- How to record experimental results
This is a new programming paradigm - you no longer write code directly, but write instructions for the agent to write code.
Technical details
Fixed five-minute budget
autoresearch made a smart design decision: each training session is strictly limited to 5 minutes.
This means:
- Whether you use H100 or RTX 4090, the experimental results can be directly compared
- Approximately 12 experiments per hour, approximately 100 per night
- All experimental results are compared fairly on the same baseline
val_bpb: fair comparison baseline
Bits per byte (bpb) is a metric that is independent of vocabulary size. This means that when the agent changes the model architecture (for example, from 8 layers to 4 layers), the evaluation results will still be fair and comparable.
Muon Optimizer
Core Concept
The traditional research process is as follows:
1. You come up with an idea
1. You change the code
1. You train the model
1. Look at the results
1. You judge good or bad
1. Repeat
autoresearch automates this entire process. you put
Program.md (equivalent to the instruction file for the agent) is set, and then the agent will:
- Modify train.py (change model architecture, hyperparameters, optimizer, batch size)
- Training for 5 minutes (fixed time budget)
- Evaluate val_bpb (validation set bits per byte, lower is better)
- Determine whether the results have improved
- Keep or discard Modify
- Repeat - can do more than 100 experiments in a day
When you wake up in the morning, what you see is not a model, but a complete experiment log.
Why are there only three files?
The beauty of autoresearch is its minimalist design. It is intentionally kept to a minimum scope:
prepare.py — Prepare once and for all
This file is responsible for:
- Download training data set
- Training BPE tokenizer
- Provides data loader and evaluation tools
Rule: Never change this file. It's set-and-forget infrastructure.
train.py — agent’s only battlefield
This single file contains:
- Complete GPT model definition
- Muon + AdamW optimizer
- Complete training cycle
The agent will modify every line of this file - model depth, attention mode, learning rate, batch size, all of which are the agent's experimental space.
program.md — Directives from the Research Director
This is the most critical document. You're not writing Python, you're writing about studying the culture of an organization.
program.md defines:
- The agent's role and goals
- Experimental evaluation criteria
- How to think and make decisions
- How to record experimental results
This is a new programming paradigm - you no longer write code directly, but write instructions for the agent to write code.
Technical details
Fixed five-minute budget
autoresearch made a smart design decision: each training session is strictly limited to 5 minutes.
This means:
- Whether you use H100 or RTX 4090, the experimental results can be directly compared
- Approximately 12 experiments per hour, approximately 100 per night
- All experimental results are compared fairly on the same baseline
val_bpb: fair comparison baseline
Bits per byte (bpb) is a metric that is independent of vocabulary size. This means that when the agent changes the model architecture (for example, from 8 layers to 4 layers), the evaluation results will still be fair and comparable.
Muon Optimizer
autoresearch uses the Muon optimizer - a relatively new optimizer that generally performs better than AdamW for the same training time. This also reflects the philosophy of autoresearch: let the agent explore what is optimal, rather than you specifying it in advance.
How to run in your environment
Prerequisites
- Single NVIDIA GPU (officially tested on H100)
-Python 3.10+
- uv package manager
Startup steps
# 1. Install uv (if not installed yet)
curl -LsSf https://astral.sh/uv/install.sh | sh
# 2. Synchronization dependencies
uv sync
# 3. Prepare information (one-time, about 2 minutes)
uv run prepare.py
# 4. Run a training manually (about 5 minutes, confirm that the environment is normal)
uv run train.py
Start independent research mode
Now, start your Claude, Codex or other coding agent:
Hi, have a look at program.md and let's kick off a new experiment!
Let's do the setup first.
Then you just have to watch. The Agent will read program.md by itself, start modifying train.py, train the model, evaluate the results, and then start the next experiment.
Runs on consumer grade hardware
Don’t have an H100? no problem. Karpathy offers some advice:
- Dataset: Using TinyStories (children's stories generated by GPT-4), entropy is lower and small models can be trained
- Vocabulary size: reduced from 8192 to 2048 or even 1024
- Sequence length: adjusted according to your hardware, can be as low as 256
- Model depth: reduced from 8 layers to 4 layers
- Batch size: down to 2^14 (~16K tokens)
The community has also launched multiple forks to support MacOS (MLX), Windows (RTX), AMD GPU and other platforms.
Why is this project important?
autoresearch is more than just a tool. It represents a research paradigm shift:
From "writing code" to "arranging agent"
In the past, researchers spent 80% of their time adjusting parameters, running experiments, and looking at results. autoresearch changes this process from "manual" to "automated". The role of the researcher changes from "executor" to "orchestrator" - you design the research organization and the agent is responsible for execution.
Human intuition + agent’s physical strength
What humans are best at is asking good questions and designing experimental directions. What AI agents are best at is a lot of repetitive work and fine-tuning. autoresearch combines the two: human intuition (research policy in program.md) + agent's physical strength (100+ experiments/night).
Reproducible independent research
Traditional autonomous research (such as Google's AlphaDev) is closed. autoresearch is open source and reproducible. Anyone can build their own autonomous research system based on it.
Actual case: What did autoresearch find?
In early runs of autoresearch, the agent discovered the following improvements:
1. Attention mode optimization: The agent automatically selects a more efficient attention mode (such as SSSL with alternating attention) instead of the standard Transformer structure preset by humans.
1. Optimizer combination: The combination of Muon + AdamW was found by the agent to be better than using any optimizer alone
1. Learning rate scheduling: The agent automatically adjusts the learning rate scheduling strategy and has a more strategic attenuation mode in the later stages of training.
These are not the results of manual attempts by human researchers, but are automatically discovered by the agent in 100+ experiments.
Comparison with related projects
autoresearch: Let agents modify code, train, and evaluate | NexusGenome: Let agents write complete applications | Autogen/CrewAI: Multi-agent division of labor, not independent research | AI Engineer: General agent tools, not dedicated to research
Summary
autoresearch is one of the most noteworthy open source projects in 2026. Its core insight is simple but extremely powerful: You don’t need to write every line of code by hand to do research—you just need to design the research organization and let AI do it.
Three profiles, one weekend, one better model. This is how future research will look.
Original source: karpathy/autoresearch
Tags: AI Agent, automatic research, model training, Karpathy, open source project, LLM
Status: Draft