Building ChatGPT from Scratch: A Complete Step-by-Step Guide to Implementing Large Language Models with PyTorch
Why should we learn LLM from scratch?
In 2026, when the AI wave has entered its third wave, I believe many developers have had similar feelings: there are more and more frameworks and APIs are becoming more and more convenient, but the more you use them, the more you feel that you are "using" AI instead of "understanding" AI.
I recently spent more than a month following the rasbt/LLMs-from-scratch project and hand-carved a complete GPT-like language model from scratch. This project has received more than 99k stars on GitHub, and updates continued until a week ago. Its companion book "Build a Large Language Model (From Scratch)" is written by Sebastian Raschka, who is one of the contributors to the "Transformer" entry on Wikipedia and has a very solid academic background in the field of deep learning.
This article is not to introduce "how to use Hugging Face to quickly train a model", but to share with you a perspective that I think is valuable to every developer who wants to do AI: what kind of insights will you gain when you personally go through the complete process from tokenizer to attention mechanism, to pre-training and fine-tuning.
What is this project doing?
Simply put, the goal of this project is to use PyTorch from scratch to implement a language model that can generate text step by step. Instead of using a ready-made library to wrap everything up for you, you start from the bottom and write every mathematical operation yourself.
The entire project is divided into 7 main chapters and 4 appendices, covering every stage of building a complete LLM:
- Chapter 1: Understand the basic concepts of large language models
- Chapter 2: Process text data and implement tokenizer
- Chapter 3: Writing Attention Mechanism
- Chapter 4: Implementing the GPT model architecture from scratch
- Chapter 5: Pre-training on unlabeled data
- Chapter 6: Fine-tuning for text classification tasks
- Chapter 7: Fine-tuning command-following tasks
Each chapter comes with a Jupyter Notebook and executable Python script, as well as answers to the exercises.
Why do I think this project deserves attention?
1. It solves a real pain point
After learning a lot of frameworks and APIs, many people are actually still vague about the "internal operations" of LLM. I know Transformer is awesome, but what exactly is the attention mechanism? Why is Position Embedding important? What exactly does Softmax do when generating text? Using the API will never force you to figure these issues out.
The approach of this project is: don’t give you a black box. Each line of code comes with mathematical formulas and text explanations, allowing you to "see" what the model is doing.
2. Very low hardware requirements
One of the most amazing things about this project for me is that it is designed to run all chapters on an ordinary laptop. No high-end GPU required, no cloud resources required. The pre-training step in Chapter 5 only takes a few hours on a CPU, even without a GPU. This means that any developer can fully experience the entire process.
3. The code quality is extremely high
Sebastian Raschka's code has a distinctive feature: it is not written to "run", but to "people can understand". Each function has a clear docstring, each mathematical operation has a corresponding formula description, and there is a clear connection between each chapter.
4. Supporting books and video courses
The project is paired with a published book (published by Manning) and a 17-hour video course. The code in the book is updated simultaneously with the open source code on GitHub, so you can choose the learning method that best suits you.
Practical hands-on: How to get started
Next let's take a look at actual usage. The structure of the entire project is very clear.
Environment settings
# clone project
git clone --depth 1 https://github.com/rasbt/LLMs-from-scratch.git
cd LLMs-from-scratch
# Use uv to install dependencies (recommended)
uv venv
uv pip install -r requirements.txt
# Or use conda
conda create -n llms python=3.11
conda activate llms
pip install -r requirements.txt
The first verifiable task: Run the tokenizer of Chapter 2
After completing the installation, the most direct way to get started is to start from Chapter 2. The code of Chapter 2 is in ch02/01_main-chapter-code/ch02.ipynb, and you can execute it step by step in Jupyter Notebook.
This Notebook will take you to implement a basic Byte-Pair Encoding (BPE) tokenizer, which is the tokenization strategy used by the GPT series models. You will see:
1. How to count token frequency from text data
1. How to gradually merge the most frequent token pairs
1. How to convert arbitrary text into token ID sequence
After executing Chapter 2, you will be able to convert any English sentence into a form that the model can understand. This is a very concrete, verifiable outcome.
Advanced tasks: Complete pre-training
If you want to challenge a higher level of difficulty, you can go all the way to Chapter 5. The code of Chapter 5 is in ch05/01_main-chapter-code/ch05.ipynb, which will take you to train a complete GPT model.
After training is complete, you can use the gpt_generate.py script to generate text:
python gpt_generate.py --prompt "The future of AI is" --max_new_tokens 50
This will output the text your model generated based on the prompts. Watching the model you trained from scratch produce the first complete sentence of text, the sense of accomplishment is very unique.
Implementation insights of several key concepts
In the process of implementing this project, I think there are three concepts that are particularly worthy of in-depth understanding:
Tokenizer is more than just "cutting characters"
Many frameworks will automatically handle tokenizers for you, but when you hand-carve a BPE tokenizer, you will understand that tokenization is actually a compression strategy. It doesn't just cut up the text, but learns "which character combinations occur most frequently in the language" and then assigns independent tokens to these combinations first. This directly affects the efficiency and performance of the model.
The mathematical beauty of Attention mechanism
Chapter 3 is the most exciting part of the entire project. You will personally implement Multi-Head Attention, which involves the calculation of Q (query), K (key), and V (value) matrices. When you see with your own eyes that Softmax assigns attention weights to different locations, you will truly understand why Transformer can capture long-distance dependencies.
Where is the "magic" of pre-training?
The pre-training process in Chapter 5 is probably the most exciting of the entire project. When you see the loss curve gradually decreasing with training steps, when your model starts to be able to complete sentences and even produce meaningful paragraphs, you will understand why everyone calls this process "training". It's not magic, but it's a very beautiful process.
Limitations and things to note
Of course, this project also has its limitations:
- It trains a "mini" model: The goal of this project is for educational purposes, so the number of parameters of the model is much smaller than that of commercial-grade LLM. It's not comparable to GPT-4 or Claude.
- Time investment required: Going through all chapters may take weeks to months, depending on your background knowledge.
- Code use academic license: The project uses a non-commercial license. If it is for commercial use, you need to pay attention.
- Mainly for English: Tokenizer and training materials are mainly in English. If you want to handle other languages, additional adjustments are required.
Who is suitable to start this project?
I think the following types of developers will gain the most from this project:
- AI engineers who want to understand LLM in depth: If you are already using Hugging Face but want to understand the underlying principles, this project is the best supplement.
- Students with PyTorch basics: This book can be used in conjunction with the project to build solid deep learning implementation capabilities.
- Technical Decision Maker: Understanding how models work can help you make more informed judgments when selecting models and designing systems.
If you're curious about the inner workings of LLM, or if you feel like you're always "using" models but never really "understanding" them, then this project is for you.
Conclusion
rasbt/LLMs-from-scratch is not just a GitHub project, it is the most complete LLM implementation tutorial on the market. Sebastian Raschka uses his profound academic skills and extremely high code quality to break down a seemingly complex topic into steps that every developer can follow.
Today, with the rapid development of AI tools, being able to understand the operation of the model from the bottom is no longer a "plus", but increasingly becoming a "necessary". This project provides not only code, but a way of thinking - a way of thinking that allows you to truly understand how AI works.
References