LlamaFactory in Practice: Fine-Tune, Infer, and Deploy LLMs with One CLI and Web UI
LlamaFactory in Practice: One Toolchain for Fine-Tuning, Inference, and Deployment
Why it matters
Turning an open model into a useful domain model involves more than starting a training job. Teams must handle model and tokenizer loading, dataset formats, training methods, GPU memory, inference checks, weight merging, and serving. LlamaFactory brings these steps together in one Python project with a CLI, a Gradio Web UI, and an OpenAI-style API with vLLM or SGLang inference backends.
As of 2026-08-05, the repository has 73,782 GitHub stars and was pushed on 2026-08-04. Its latest official release, v0.9.5, was published on 2026-05-30. This is an implementation-oriented AI Engineering tool, not a resource list or a learning-only repository.
Core capabilities
The official README describes support for Llama, Qwen, DeepSeek, Gemma, Mistral, GLM, Phi, and multimodal models. Training approaches include pre-training, supervised fine-tuning, reward modeling, PPO, DPO, KTO, ORPO, and SimPO. Parameter-efficient options include freeze-tuning, LoRA, and 2/3/4/5/6/8-bit QLoRA through several quantization backends. Experiment tracking can use LlamaBoard, TensorBoard, W&B, MLflow, or SwanLab, while inference can use the CLI, Gradio, vLLM, or SGLang.
The main value is integration: models, datasets, templates, trainers, and inference interfaces share one configuration ecosystem. That reduces the manual handoffs between an experiment and an API service.
The minimum path: LoRA, chat, and export
The official Quickstart uses Qwen3-4B-Instruct and exposes three commands:
llamafactory-cli train examples/train_lora/qwen3_lora_sft.yaml
llamafactory-cli chat examples/inference/qwen3_lora_sft.yaml
llamafactory-cli export examples/merge_lora/qwen3_lora_sft.yaml
train validates the dataset, template, memory settings, and hyperparameters. chat checks the adapter with the base model before merging. export produces a checkpoint suitable for the next deployment step. Keeping the chat validation before export is a useful guard against shipping a broken template or dataset.
Data preparation is the first engineering problem
Custom datasets must be registered in data/dataset_info.json and formatted according to data/README.md. Sources can include Hugging Face, ModelScope, Modelers Hub, local storage, S3, or GCS. Validate structure, chat templates, and data quality separately. The framework handles loading and formatting, but deduplication, privacy filtering, label consistency, and evaluation-set isolation remain application responsibilities.
LoRA, QLoRA, and GPU trade-offs
The README estimates roughly 16GB for a 7B model with 16-bit Freeze/LoRA-style tuning and roughly 6GB with 4-bit QLoRA, while full tuning requires much more memory. These are estimates; sequence length, batch size, optimizer, gradient accumulation, and activation checkpointing change the actual footprint.
Start with a small LoRA smoke test, move to QLoRA when memory is constrained, and evaluate distributed options such as DeepSpeed, FSDP, or Megatron-core only when the single-device path is understood. Compare the base, adapter, and merged models on a held-out evaluation set instead of relying on training loss alone.
Moving from experiments to serving
Use llamafactory-cli webui for the LlamaBoard interface, an OpenAI-style API with vLLM for application integration, or Docker images and Compose to pin the CUDA, PyTorch, Python, and FlashAttention environment. Before serving, constrain input length and concurrency, version the adapter together with the tokenizer and chat template, and add audit and safety controls. API compatibility does not automatically provide production security.
What v0.9.5 signals
The v0.9.5 release notes mention primary support for Qwen3.5/Qwen3.6 and Gemma 4, Transformers v5 compatibility, an FP8 Transformer Engine backend, Megatron-related integration, and a CLI sampler. The project is moving quickly with model and infrastructure changes, so reproducibility requires pinning release tags or commits and running a small smoke test before upgrades.
Conclusion
LlamaFactory connects dataset registration, template selection, LoRA/QLoRA training, interactive validation, checkpoint export, and API deployment in one repeatable toolchain. It does not replace data-quality or evaluation design, but it can reduce the glue code between research experiments and application integration. Teams building private knowledge, customer-support, code, or multimodal models should start with a small model and dataset, validate the end-to-end path, and then scale it deliberately.
References
- GitHub repository: https://github.com/hiyouga/LlamaFactory
- Official README: https://github.com/hiyouga/LlamaFactory/blob/main/README.md
- v0.9.5 release: https://github.com/hiyouga/LlamaFactory/releases/tag/v0.9.5
- Official documentation: https://llamafactory.readthedocs.io/en/latest/
- Paper: https://arxiv.org/abs/2403.13372