Chapter 7: Finetuning
Finetuning Foundation Models: A Practical Guide
How to decide when to finetune, memory-efficient techniques (PEFT, LoRA, QLoRA), RAG vs finetuning, model merging, and a practical workflow — distilled from Chapter 7 of AI Engineering. (Kindle clipping supplied by the author.)
TL;DR
Finetuning adapts a pre-trained foundation model to a specific task by updating some or all of its parameters. It can dramatically improve domain-specific performance and control model behavior, but it’s resource-intensive. Today, parameter-efficient techniques (PEFT) like LoRA and quantized workflows such as QLoRA make finetuning accessible without a data-center-sized budget. For many problems, a structured approach — try prompts first, add RAG for information gaps, and only finetune when behavior or format problems persist — yields the best cost/benefit.

1. What is finetuning and why does it matter
Finetuning means continuing training on a pre-trained (foundation) model so it performs better on your target task. Unlike prompt engineering or RAG, finetuning changes model weights and permanently alters model behavior. It’s transfer learning applied after pre-training: you leverage broad, general knowledge learned during pre-training and push the model toward the specific patterns, style, or formats your application needs.
Use finetuning when:
The base model consistently fails on the desired task or format (semantic parsing, strict output formats, domain-specific reasoning).
You need consistent behavioral changes (tone, safety constraints, or output structure) that are hard to force with prompts alone.
Avoid finetuning when prompt engineering or retrieval (RAG) solves the issue more cheaply — finetuning requires data, compute, and ongoing maintenance.
2. Types of finetuning
Continued pretraining (self-supervised): keep pretraining on domain-relevant unlabeled text to improve representations.
Supervised finetuning: train on (input, output) examples — great for instruction-following or structured output tasks.
Reinforcement/preference finetuning (RLHF/RL from preference data): optimize for human preference using comparative labels (winning vs losing responses).
Infilling / next-token finetuning: train for tasks like code editing or text completion, where the model must fill gaps.
Each has its place — pick the one that matches the data you can collect.
3. RAG vs Finetuning: when to use which
RAG (Retrieval-Augmented Generation) supplies external facts at inference time. It’s best when failures are information-based (missing or outdated facts). Start with RAG for factual or timely content.
Finetuning addresses behavioural failures: output format, syntax, style, or tasks where the model must reliably produce a specific structure.
A practical pipeline: 1) Prompt engineering → 2) Add simple retrieval (BM25) → 3) Move to vector retrieval if needed → 4) Consider finetuning for persistent behavior/format problems. Combining RAG + finetuning can give additive benefits in many benchmarks.
4. Memory bottlenecks: why finetuning can be expensive
Finetuning costs scale with: the model’s parameters, the number of trainable parameters, and numerical precision. Training memory roughly equals:
Training memory = model weights + activations + gradients + optimizer states
Examples and implications:
Full finetuning (update all weights) requires gradients for every parameter — big memory cost.
Numerical representation matters: storing 13B parameters in FP32 needs ~52 GB for weights; halving bits halves the footprint.
Techniques like gradient checkpointing (recompute activations) and CPU offloading (DeepSpeed-style paging) help reduce peak GPU memory.
Key principle: reduce the number of trainable parameters and the precision at which values are stored to make finetuning practical.

5. Parameter-Efficient Finetuning (PEFT)
PEFT is the dominant approach to make finetuning affordable. Two broad families:
Adapter-based (additive) — insert small modules into the network and only train those. Example: Houlsby adapters.
Soft prompt-based — append trainable vectors (continuous tokens) to inputs that steer the model.
Both reduce memory and data needs. They make finetuning possible on modest hardware and simplify serving multiple task-specific variants.

LoRA (Low-Rank Adaptation)
LoRA factorizes weight updates into two small matrices (A and B). Instead of updating a huge weight matrix, you learn low-rank updates. Benefits:
Very parameter-efficient and sample-efficient.
Modular: LoRA adapters can be kept separate from the base model and swapped in/out per task.
Can be merged back into base weights for low-latency serving or kept separate to support multi-adapter setups.
LoRA tends to hit the sweet spot of memory efficiency vs performance, which is why it’s widely used.
QLoRA (Quantized LoRA)
QLoRA stores base weights in 4-bit formats (e.g., NF4) and dequantizes to BF16 during training, combined with paged optimizers to move data between CPU/GPU. This lets practitioners finetune very large models (e.g., 65B) on a single 48 GB GPU.
Caveats: QLoRA reduces memory but can increase training time due to quantization overhead and requires careful tooling.
6. Model merging & multi-task finetuning
When you need several task-specialized models, consider model merging instead of training monolithic multi-skill models.
Approaches include:
Summing / linear combinations of weight deltas (good when models share a base and architecture).
Spherical linear interpolation (SLERP) for smoother interpolation between two models.
Layer stacking (frankenmodels) and concatenation — useful for novel architectures but can increase parameter counts.
Merging supports multi-task use-cases and can be cheaper than simultaneous finetuning, though it comes with alignment and compatibility challenges.

7. Practical finetuning tactics & hyperparameters
Frameworks: use an API if you want simplicity and limited model choices; use DeepSpeed / PyTorch Distributed / ColossalAI for large-scale or multi-machine runs.
Paths to finetune:
Progression path: test code on a cheap model → try middling model for data sanity → run best model experiments → map price/performance.
Distillation path: finetune a strong model with small curated data → use it to generate more training data → train a cheaper model.
Key hyperparameters:
Learning rate: small, tuned to loss curves. If loss is noisy, lower the rate; if too slow, increase until instability appears. Use schedules.
Batch size: affects stability; tiny batches (<8) can be unstable. Use gradient accumulation to simulate larger batches when memory is limited.
Epochs: watch training vs validation loss to avoid overfitting.
Prompt loss weight: for instruction finetuning, weigh response tokens higher than prompt tokens (defaults often around 10%).
Monitoring: log training and validation loss, and track human-aligned metrics whenever possible.
8. Serving and deployment
Merge adapters when you need low-latency single-task endpoints.
Keep adapters separate when you need multiple task adapters on top of one base model — trades latency for storage efficiency.
QLoRA and other quantized approaches may require special runtime support.
9. Checklist: Should you finetune? A minimal decision flow
Can prompts + evaluation pipeline reach your SLA? If yes → don’t finetune.
Is the model failing due to missing facts? Try RAG (start with BM25, move to vector search if needed).
If failures are behavioral (format, style, irrelevance) → try PEFT (LoRA) on a small dataset.
Run experiments on a smaller model first (progression path) or distill from a strong model if data is limited.
Track metrics and watch for catastrophic forgetting when doing sequential finetuning.
10. Example quick-start recommendations
Start with LoRA adapters on attention & feed-forward matrices. Use a small rank (r = 4–16) initially.
Learning rate: try 1e-4 → 5e-5 with an AdamW variant; use warmup + cosine scheduler.
Batch size: make per-GPU 8–32; use gradient accumulation to reach an effective batch of 128 where possible.
Epochs: 3–10 depending on dataset size; prioritize validation metrics and early stopping.
These are starting points — always tune to your dataset and compute.
Closing thoughts
Finetuning remains the most direct path to tailor foundation models to real products. The good news: a modern toolbox — PEFT, LoRA, QLoRA, gradient checkpointing, and CPU offloading — has dramatically lowered the barrier. Start with prompts and retrieval, build a solid evaluation pipeline, then adopt parameter-efficient methods when you need lasting behavioral change. With careful design and evaluation, finetuning is now within reach for many teams previously priced out of model customization.