1. Do you even need to fine-tune?
Fine-tuning is the heavy answer. Most problems have lighter ones. Run through this list before spending a dollar on training.
- Prompt engineering first. A good system prompt with a few examples (few-shot prompting) solves many style and format problems. It costs nothing and takes minutes. If the model can do the task when shown how, it does not need training.
- Retrieval (RAG) for knowledge. Fine-tuning is a bad way to teach facts. Models trained on new facts hallucinate them back with confidence. If your task needs private documents, wire retrieval into the prompt instead.
- Fine-tune for behavior, not information. The cases where training earns its keep: a consistent voice across thousands of outputs, a rigid output format (JSON schemas, tool calls), a domain workflow the model keeps getting wrong, or latency and cost (a small tuned model replacing a large general one).
If you can describe the desired behavior in under a page of instructions and the model follows it, stop there. Fine-tune when you need the behavior baked in: always-on, no prompt to maintain, cheaper to serve, and robust to prompt injection drift.
2. The method map
Post-training means changing a model's weights after pretraining. The methods below stack. A common pipeline is: base model → SFT on demonstrations → preference alignment (DPO) → quantize and serve.
| Method | What it does | When to use it | Compute |
|---|---|---|---|
| Full SFT | Updates every weight on input→output examples | Large domain shift; you have serious data and GPUs | High |
| LoRA / QLoRA | Trains small adapter matrices; base weights stay frozen | Default choice for tasks and style; single-GPU friendly | Low |
| DPO / ORPO / KTO | Aligns on chosen vs rejected response pairs, no reward model | Polishing tone, format, and preference after SFT | Low–medium |
| RLHF (PPO) | Trains a reward model, then optimizes against it | Large labs; rarely worth it for individuals now | High |
| GRPO / RLVR | Reinforcement learning with verifiable rewards (math, code, tools) | Reasoning tasks with checkable answers | Medium–high |
| Distillation | A small student learns from a large teacher's outputs | Compressing capability into a cheap servable model | Medium |
| Model merging | Combines weights of two tuned models, no training | Blending skills (e.g. style adapter + task adapter) | Near zero |
3. Supervised fine-tuning (SFT)
SFT is the foundation. You show the model thousands of input→output examples and train it to predict the outputs. Everything else in this guide either replaces SFT's full-weight update with something cheaper or polishes the result afterward.
How it works, in one paragraph
Each training example is a prompt plus the desired response. The model reads the prompt, predicts the response token by token, and the loss measures how far its predictions were from the actual tokens. Gradient descent nudges the weights to close the gap. Repeat for 1–3 passes over the data (epochs). That is the whole trick.
Full SFT vs adapters
Full SFT updates every parameter. For a 7B model in 16-bit precision that means ~14 GB just for the weights, plus gradients and optimizer state. Adam keeps a master copy plus two running statistics per parameter, so full fine-tuning needs roughly 16 bytes per parameter at 16-bit: about 112 GB for 7B. That is why individuals rarely do full SFT on anything above 3B without serious hardware. Section 4 covers the standard workaround.
Hyperparameters that actually matter
- Learning rate: 5e-6 to 2e-5 for full SFT, far below LoRA rates. Too high and the model collapses into repeating itself. Too low and nothing changes.
- Epochs: 1–3 for large datasets; up to 5–10 for tiny ones (under 1,000 examples). More epochs on small data = memorization, not learning.
- Batch size: effective batch of 32–128 via gradient accumulation. Larger batches stabilize training but need more memory.
- Warmup: ramp the learning rate over the first ~5% of steps, then cosine decay. Skipping warmup is the most common cause of early-training collapse.
- Weight decay: 0.01–0.1. Mild regularization against overfitting.
- Loss masking: compute loss on the assistant's tokens only, not the prompt. Pack short examples together to stop wasting compute on padding.
Catastrophic forgetting. SFT on narrow data erodes the model's general ability. A model trained only on your support tickets will get worse at everything else. Mix in 10–30% general instruction data to keep the base capabilities intact, and check both directions: your task metrics and a general benchmark.
4. LoRA and QLoRA: the default for everyone else
LoRA (Low-Rank Adaptation) freezes the base model and trains two small matrices per layer whose product approximates the weight change. Instead of updating a 4096×4096 weight matrix (16.8M numbers), you train a 4096×16 and a 16×4096 pair (131K numbers). Same architecture, 0.1–1% of the parameters.
Why LoRA usually wins for individuals
- Memory: the frozen base can be quantized to 4-bit. A 7B model fits in ~6 GB; the adapters and optimizer state for the adapters add a few GB more. Total: trainable on a single 24 GB consumer GPU (RTX 4090), often on 16 GB cards.
- Speed: fewer parameters to update means faster steps and cheaper experiments. You can try ten ideas in the time one full run takes.
- Modularity: adapters are small files (tens to hundreds of MB). Swap them at load time: one adapter for style, one for a task, one per client. The base model stays untouched.
- Quality: on instruction-following and style tasks, LoRA matches full fine-tuning. It falls behind only when the task needs deep knowledge changes (new languages, heavy domain shift).
QLoRA: the 4-bit variant
QLoRA combines LoRA with a 4-bit quantized base (usually NF4) plus paged optimizers. It is the reason a single GPU can fine-tune 7B–13B models at all. The finding that launched it: 4-bit quantization barely dents fine-tuning quality. Use QLoRA unless you have a reason not to.
Settings that work
- Rank (r): 16 is the default. Start at r=8 for style-only shifts; go 32–64 for harder tasks. Higher rank = more capacity, more memory, slower.
- Alpha: set to 2× rank (the classic recipe). It just scales the adapter's influence.
- Target modules: start with attention (query, key, value, output). If the model underfits, add the MLP projections (gate, up, down) or target all linear layers.
- Learning rate: 2e-4 with a cosine schedule, higher than full SFT because you are moving fewer parameters.
- Dropout: 0.05 on the adapter, cheap insurance on small datasets.
Overfitting is the number one LoRA failure mode: training loss falls while held-out scores stall. Fix in this order: reduce rank, raise dropout, cut epochs, add data. Do not reach for a bigger rank first.
DoRA splits magnitude and direction of the update and slightly improves quality at small extra cost. rsLoRA fixes a scaling quirk so higher ranks actually help. LoRA+ uses different learning rates for the two adapter matrices. None of these change the workflow; they are drop-in upgrades in most libraries.
5. Preference alignment: RLHF, DPO, and friends
SFT teaches the model what to say. Preference methods teach it which answer is better. You give pairs: prompt, chosen response, rejected response. The model learns to prefer the chosen one.
DPO killed the RLHF pipeline for most people
Classic RLHF trains a separate reward model on human preferences, then optimizes the policy against it with PPO. Three models, finicky hyperparameters, frequent reward hacking. DPO (Direct Preference Optimization) skips all of it: a single loss function that pushes probability toward the chosen response and away from the rejected one. Same data format, one training run, stable optimization. For style polish and format compliance, DPO after SFT is the standard stack.
The family
- DPO: the default. Key knob is beta (0.1), which controls how far the model may drift from the reference. Low beta = stronger preference following, higher risk of weirdness. Never reuse the policy as its own reference; the distribution collapses within an epoch or two.
- ORPO: folds SFT and preference learning into one step, no separate reference model, at roughly a quarter of PPO's cost. Good when you want a single-stage run.
- KTO: needs only good/bad labels, not pairs. Useful when your feedback is thumbs-up/thumbs-down rather than A-vs-B comparisons.
- SimPO: drops the reference model and normalizes for length, which reduces the classic failure where the model learns that longer = better.
DPO is a gentle nudge, not a capability builder. It re-ranks behavior the model already has; it does not teach new skills. If the base model cannot do the task at all, fix that with SFT or distillation first.
Preference data is full of an accidental signal: chosen responses tend to be longer. Models learn to ramble. Counter it by keeping chosen and rejected responses similar in length, or use a length-normalized method like SimPO.
Where do preference pairs come from? Generate two responses per prompt from your SFT model, then rank them: by hand for a few hundred (gold standard), with a stronger model as judge (cheap, slightly noisy), or from implicit signals like user edits and thumbs votes. A few hundred to a few thousand pairs is enough for a noticeable polish pass.
6. RL post-training: GRPO and RLVR
The biggest post-training story of the last two years is reinforcement learning with verifiable rewards. Instead of a learned reward model guessing what humans like, you use rewards you can check: the math answer is right, the code passes its tests, the tool call returned the right schema. DeepSeek-R1-Zero showed what this unlocks: pure RL on a base model took AIME math scores from 15.6% to 77.9%, with no human reasoning examples at all.
Why it matters
RLVR (reinforcement learning with verifiable rewards) is what made small models reason. DeepSeek-R1 showed that RL on checkable problems produces long chains of self-correction: trying, failing, backtracking. SFT can imitate reasoning traces, but RL discovers them. If your custom task has right answers you can verify programmatically, RL post-training beats more SFT.
GRPO in one paragraph
GRPO (Group Relative Policy Optimization) samples a group of responses to the same prompt, scores each with your verifier, and pushes the policy toward the better ones relative to the group average. No value network, no reward model, which makes it far cheaper than PPO. Libraries (TRL, OpenRLHF, veRL) now ship GRPO trainers you can point at your own reward function.
When to use it
- Your task has objectively checkable outputs: code, math, structured extraction, game play, tool use with defined success.
- You have already done SFT and hit a plateau. RL squeezes out the next 5–15%.
- You can afford the compute: RL needs many rollouts per prompt, typically 4–8× the cost of the SFT run that preceded it.
The model will find the cheapest way to score, not the way you meant. A code reward based on passing tests produces code that games the tests. Keep verifiers strict, hold out test cases the model never sees during training, and spot-check rollouts by hand. If your reward can be gamed, assume it will be.
7. Knowledge distillation
Distillation trains a small student on the outputs of a large teacher. The teacher generates answers (or reasoning traces, or full token probability distributions), and the student learns to match them. You get much of the teacher's skill in a model you can actually serve.
Three flavors
- Black-box distillation: prompt a strong model, collect its outputs, SFT your small model on them. This is how most open "reasoning" models were built: teacher writes long chains of thought, student learns the pattern. Check the teacher's terms of service first; some providers prohibit training competing models on their outputs.
- White-box distillation: train the student to match the teacher's token probabilities (logits), not just its final text. Richer signal per example, needs access to the teacher's weights.
- Data distillation: the teacher generates your training dataset (synthetic examples, preference pairs, edge cases), and you train normally. This is the most common real-world use.
Distillation transfers behavior, not knowledge the teacher does not reliably have. A student distilled from a teacher's shaky domain knowledge inherits the shakiness. Distill reasoning patterns and style freely; distill facts cautiously.
8. Choosing a base model
The base model sets your ceiling. Pick wrong and no amount of tuning recovers it. Three criteria dominate: capability at your size, license, and context length.
The current open-weight landscape
The 2026 default base is Qwen3: strongest performance per parameter in the open field, Apache 2.0 license, long context, and a large fine-tuning community. Llama 3.3 remains the safe generalist if you want the Meta ecosystem and tooling.
| Family | Sizes | License | Notes |
|---|---|---|---|
| Qwen3 | 0.5B–72B+ | Apache 2.0 | Default fine-tune base; coding, agents, multilingual |
| Llama 3.3 | 1B–70B | Llama Community License | Excellent ecosystem; license has use restrictions, read before productizing |
| DeepSeek | 1.5B–70B+ (distills) | Often MIT for distills | Reasoning distills are the best cheap starting point for RL work |
| Gemma | 1B–27B | Google terms; check the model card | Strong at small sizes; licensing is inconsistently reported, verify per checkpoint |
| Mistral | 7B–24B+ | Apache 2.0 / MNPL | Efficient architectures; check license per release |
| gpt-oss (OpenAI) | 20B / 120B | Apache 2.0 | OpenAI's open entry; surprisingly strong |
| Phi | 3B–14B | MIT | Microsoft's small models; punchy for their size |
How to choose
- Size: 7–8B is the sweet spot for individuals. Big enough to follow instructions well, small enough to train with QLoRA on one GPU and serve cheaply. Go 1–3B if latency or edge deployment dominates. Go 32–70B only if you have the hardware budget and the task truly needs it.
- License: Apache 2.0 or MIT means no strings attached commercially. Community licenses (Llama, Gemma) are fine for most businesses but read the terms; some restrict competitive use or trigger obligations at scale.
- Instruct vs base: start from the instruct version. It already follows instructions; your fine-tune specializes it. Training from a raw base model means re-teaching instruction-following from scratch.
- Context length: match it to your task. Long-document workflows need 32K+ context. Fine-tuning does not extend context well; pick a model that already has the window you need.
Model releases move monthly. Before committing, check the current leaderboards (LMArena, Open LLM Leaderboard) at your size class. A model from six months ago is often strictly worse than this month's release at the same size.
9. Building the dataset
Data quality decides the outcome more than any hyperparameter. The repeated finding across the literature: a few thousand excellent examples beat a hundred thousand mediocre ones. The LIMA result is the landmark: 1,000 meticulously curated pairs on a 65B model beat 52,000 lower-quality examples. The refinement since: style saturates at around 100 examples, while new capabilities keep improving with more data.
Format
Most trainers expect instruction pairs or conversations. The modern standard is the chat template: a list of messages with roles.
{"messages": [
{"role": "user", "content": "Summarize this quarter's churn drivers."},
{"role": "assistant", "content": "Three drivers explain 80% of churn..."}
]}
Match the base model's own chat template exactly. Every major library applies it automatically; the mistake is hand-rolling a format the model never saw.
How much data
- Style and format: style itself saturates near ~100 examples; 200–2,000 examples give a robust, generalizing voice shift.
- A focused task: 1,000–10,000 examples for reliable behavior. Below 200–500, models tend to memorize rather than generalize.
- Broad domain adaptation: 10,000–100,000+ examples, with replay data mixed in.
Where the data comes from
- Your own work. For style mimicry this is the gold: your emails, docs, posts, code. Nothing beats it.
- Synthetic generation. Prompt a strong model to produce examples in your format, then curate hard. Self-Instruct and Evol-Instruct are the classic recipes: seed with a few examples, ask the teacher to invent variations, filter ruthlessly.
- Distillation. Have the teacher solve your actual task inputs and keep the good outputs (section 7).
- Human annotation. Slowest, best. Worth it for the final few hundred examples that define quality.
Cleaning checklist
- Deduplicate aggressively. Near-duplicates teach memorization.
- Cut anything the model already does well. Training on solved cases wastes capacity.
- Fix the labels by hand on a sample. If 5% of your examples are wrong, the model learns 5% wrongness.
- Keep prompt diversity high: vary phrasing, length, and edge cases so the model learns the task, not the template.
- Check training data against your eval benchmarks for overlap (n-gram contamination checks). Training on the test is the oldest cheat in the book.
- If one model generates your synthetic data, use a different model family to judge it. A model grading its own output collapses into self-preference.
- Hold out 5–10% as an evaluation set before training. Never tune on it, never train on it.
10. Mimicking a writing style
This is the most personal use case: a model that writes like you. The good news is that style is one of the cheapest things to transfer. It saturates at around 100 examples, and 2026 work (InMyStyle) showed per-user LoRA adapters working on models as small as 0.5B, with judges rating the outputs over 20% less "AI-sounding." You do not need a giant model to clone a voice.
The pipeline
- Collect 200–500 samples of your writing in the target register to start; iterate toward 1–2,000 for polish. Emails if you want email voice; long-form if you want essay voice. One register per adapter; mixing registers blurs the result. Consistency of voice matters more than volume.
- Pair each sample with a plausible prompt. A raw essay is not training data. Write the instruction that would have produced it: "Write a memo to the team about the Q3 delay" → your actual memo. For emails, the prompt can be "Reply to this thread: ..." with the thread summarized. A strong model can draft these prompts for you; review them.
- Scrub. Remove anything you do not want the model to reproduce: names, private details, company internals, and your bad habits. The model will faithfully learn your typos and your tics. Curate like an editor.
- Train LoRA on the instruct checkpoint, not the base. The instruction-tuned conversational priors are worth preserving; full fine-tuning erodes them (researchers observed random topic-switching after full FT on personal corpora). Start at rank 8–16; rank 4 suffices for light tone shifts.
- Keep a short style system prompt too. The adapter and the prompt stack: the prompt gets you most of the way for free, the adapter makes it consistent across thousands of calls without eating context.
- Evaluate blind. Generate passages from prompts the model never saw. Mix them with your real writing. Ask someone who knows your voice to sort them, or run pairwise A/B judging ("which sounds more like X"). If they cannot tell reliably, you are done. Track a general benchmark alongside to catch capability regression.
Full fine-tuning rewrites the model's general language habits along with everything else, which often washes out the crisp edges of a personal voice. LoRA's small updates steer the existing capability instead of replacing it, so your phrasing habits come through while the model's grammar and reasoning stay intact.
What does not work
- Too little data, too many epochs. Fifty examples trained for twenty epochs gives you a parrot that repeats your exact sentences, not your voice. More data, fewer epochs.
- Mixed registers. Training on tweets, whitepapers, and love letters at once produces an average of all three, which is none of them.
- Expecting facts to transfer. The model will write like you about things you never wrote about, drawing on its own knowledge. It will not know your private opinions on new topics. Style transfers; knowledge does not.
11. Compute and cost
The memory math
Three numbers dominate GPU memory during training: the weights, the gradients, and the optimizer state.
- Weights: parameters × bytes per parameter. 7B in 16-bit = ~14 GB. In 4-bit = ~3.5 GB.
- Gradients: only for trainable parameters. LoRA makes this tiny.
- Optimizer (Adam): ~8 bytes per trainable parameter. Full SFT on 7B needs ~84 GB for optimizer state alone. LoRA on 7B with rank 16 trains ~40M parameters: ~0.3 GB.
- Activations: grow with batch size and sequence length. Gradient checkpointing trades compute for memory here.
| Setup | Typical VRAM | Hardware |
|---|---|---|
| QLoRA, 7–8B model | 10–12 GB | Single RTX 4090 (24 GB) |
| QLoRA, 13B model | ~16 GB | RTX 4090, or single 48 GB card |
| QLoRA, 32–33B model | ~24 GB | RTX 5090 (32 GB) / A100 40 GB |
| QLoRA, 70B model | ~48 GB | Single A100/H100 80 GB |
| Full SFT, 7B model | ~112 GB | 2× 80 GB or multi-GPU with sharding |
| Full SFT, 70B model | ~1.1 TB | Multi-node cluster; not an individual project |
Renting GPUs
You do not need to own hardware. Spot and on-demand GPU rentals put a card in your hands in minutes. Indicative September 2026 rates on RunPod: RTX 4090 around $0.74/hr ($0.34 community/spot), A100 80GB around $1.59/hr, H100 around $3.49/hr; Vast.ai runs cheaper but host reliability varies. A typical 7B QLoRA run takes about 45 minutes on a 4090-class card: a few dollars per experiment. The expensive part is iterating ten times, so budget for iteration, not for one run. Prices move constantly; check the provider's live pricing before you commit.
Free options
Google Colab (free tier) gives you a modest GPU with time limits, enough for small LoRA experiments on 1–3B models or short 7B runs. Kaggle offers weekly free GPU hours. Both are fine for learning the workflow before you spend money.
Your first training run will not be your last. Plan for 5–15 experiments: data fixes, hyperparameter sweeps, and evaluation iterations. A realistic personal project budget is $50–$300 in compute for a 7B LoRA project, most of it spent on experiments 2 through 10.
12. The tooling stack
| Tool | What it is | Use it when |
|---|---|---|
| Hugging Face TRL | The reference trainers: SFT, DPO, GRPO, KTO | You want standard, maintained implementations |
| Unsloth | Hand-optimized kernels; 2–5× faster and ~70% less VRAM than the Hugging Face baseline | Single consumer GPU. The individual-developer default for LoRA/QLoRA |
| Axolotl | YAML-config training for many methods, multi-GPU with FSDP2 | Reproducible configs, multi-GPU, long context |
| LLaMA-Factory | Web UI + CLI for SFT/DPO/RL | You prefer a GUI over scripts |
| torchtune / LitGPT | Clean PyTorch-native training stacks | You want readable code you can modify |
| OpenRLHF / veRL | Full RLHF/RL pipelines | You are doing PPO or GRPO at scale |
| vLLM / SGLang | High-throughput inference servers | vLLM: production GPU serving with runtime LoRA-adapter swapping (one server, many adapters). SGLang: wins on agent/RAG workloads with repeated prefixes and heavy structured output |
| Ollama / llama.cpp | Local quantized inference | Running the model on your own machine |
| Hugging Face Hub | Model and dataset hosting, private repos | Storing adapters, sharing, versioning |
No-code and API options
Several providers offer fine-tuning as a service: OpenAI (about $0.25 per million training tokens), Together AI, Fireworks AI (from about $0.50 per million training tokens), and others let you upload a dataset and get back a tuned model behind an API. You trade control and data privacy for zero infrastructure. A million-token SFT job costs pocket change in training tokens; per-token inference pricing dominates the long-term bill. Read the terms on data retention before uploading anything sensitive, and never train on a provider's outputs against their terms. For style mimicry on non-sensitive text, these are genuinely the fastest path.
# The shortest real training loop (Unsloth + TRL), conceptually:
# 1. pip install unsloth trl datasets
# 2. Load a 4-bit base, attach LoRA adapters (rank 16, all linear layers)
# 3. Train on your chat-formatted JSONL for 1-3 epochs, lr 2e-4
# 4. Save the adapter (a few hundred MB), merge or serve with vLLM
13. Evaluation
Training without evaluation is guessing. Decide before training what "better" means and how you will measure it.
Three layers
- Task evals you build. 100–500 held-out examples with clear expected outputs. Score them automatically where possible (exact match, regex, unit tests) and by rubric where not. This is the eval that matters most.
- General benchmarks. MMLU, IFEval, MT-Bench, Arena-Hard. Note that MMLU is saturated above 88% by current models, so treat it as a regression canary (did I break anything?) rather than a differentiator. AlpacaEval's length-controlled win rate corrects for verbosity bias.
- Human or LLM-as-judge. For style and quality, have a strong model compare outputs pairwise against the base model, blind. Judges reach roughly 80% agreement with humans at about a cent per judgment: noisy but cheap, and reasonable on style tasks.
Checkpoint discipline
- Save checkpoints every 10–25% of training. The final checkpoint is often not the best.
- Pick the checkpoint with the best held-out score, not the lowest training loss. Training loss always goes down. That tells you nothing.
- Watch for the overfitting signature: training loss falling while eval scores stall or drop. Stop early when you see it.
For voice mimicry, automated metrics fail. The test is blind human sorting: can a reader who knows your writing distinguish your text from the model's? Run it on fresh prompts, not training examples. Fifty passages and one honest reader beats any BLEU score.
14. Serving the finished model
A trained adapter is not a product until it runs somewhere cheap and fast. The standard move is quantization: shrinking the weights so inference needs less memory and runs faster, with minimal quality loss.
| Format | Use case |
|---|---|
| GGUF | Local and edge inference via llama.cpp and Ollama. The universal CPU/laptop format. |
| AWQ / GPTQ | GPU inference with 4-bit weights. Pairs with vLLM for serving. AWQ is slightly better at 4-bit; both need calibration data. |
| bitsandbytes 8-bit | Quick local experiments without a conversion step. |
- Merge first, quantize second. Merge LoRA adapters into the base weights before quantizing. Quantizing first and merging after compounds the errors.
- Merge or not? Merged adapters become one file with zero inference overhead. Kept separate, adapters hot-swap at runtime in vLLM: one server can serve many voices or tasks. Merge for a single-purpose deployment; keep adapters separate for multi-tenant setups.
- vLLM is the default server for GPU deployment: continuous batching, OpenAI-compatible API, LoRA adapter hot-swapping.
- Ollama is the default for "it just runs on my machine": one command, GGUF under the hood.
- Cost reality: a quantized 8B model serves comfortably on a single 24 GB GPU, or on CPU for low traffic. A fine-tuned API from a provider costs per token with no ops burden. Do the per-token math at your expected volume before choosing.
15. Pitfalls and honest caveats
- Decision order: prompt, then RAG, then fine-tune. Each step costs roughly 10× the engineering of the last. Prompt engineering is free and instant, and it is the mandatory baseline: you cannot claim fine-tuning helped if you never tuned the prompt.
- RAG vs fine-tuning, again. If the problem is "the model doesn't know my documents," retrieval wins. Fine-tuning teaches behavior. Mixing them up wastes weeks.
- Fine-tuning does not fix reasoning. A model that cannot do the task at all with a good prompt will not learn it from 500 examples. It will memorize the examples instead. DPO only re-ranks existing behavior; it builds no new capability.
- Fine-tuning pays off economically at volume. Baked-in behavior replaces long few-shot prompts, cutting per-request token cost 30–70%. The rough break-even is above ten thousand requests.
- Data licensing is real. Training on text you do not own creates legal exposure, especially if you serve the result commercially. Provider terms from OpenAI, Anthropic, and Google prohibit training competing models on their API outputs; distill from open-weight teachers instead. Your own writing is clean. Everything else needs a license check.
- Privacy runs one way. Anything in the training data can leak out of the model. Never train on secrets, credentials, or private data about other people. Scrub first, keep style corpora in private repos.
- Small data lies. Great eval scores on 50 examples mean nothing. Overfit models ace tiny evals and fail in the wild. Make the eval bigger than feels necessary.
- Base model updates obsolete adapters. A LoRA trained for one base does not transfer to the next release. Budget for retraining when you switch bases.
- Safety tuning is thin on open models. Your fine-tune can weaken refusal behavior. If you serve the model to others, test adversarial prompts and consider keeping the provider's safety layers.
16. Three playbooks
Goal: a model that drafts in your voice.
1. Collect 200–500 samples of your writing in one register to start; iterate toward 1–2,000 for polish. 2. Write a plausible prompt for each (a strong model can draft these; you review). 3. Scrub names, secrets, and tics you dislike. 4. Format as chat JSONL, hold out 10%. 5. QLoRA on a 7–8B instruct model: rank 8–16, lr 2e-4, 2–3 epochs. 6. Generate on fresh prompts; blind-sort against your real writing. 7. If it is close but loose, add 200–500 DPO pairs (your edit preferred over the model's draft). 8. Quantize to GGUF, run in Ollama.
Budget: one evening of curation, $10–$50 in GPU rental.
Goal: a model that executes your multi-step workflow reliably, e.g. triaging tickets into your taxonomy with your fields.
1. Define the task as input→output with a strict schema. 2. Build 2,000–10,000 examples: distill from a strong model on your real inputs, then hand-fix the failures. 3. Mix in 10% general instruction data. 4. QLoRA rank 32–64, lr 1e-4, 1–2 epochs. 5. Eval on 300+ held-out cases with programmatic checks (schema valid, fields correct). 6. Mine the failures, add targeted examples, retrain. 7. If outputs are checkable (tests, schemas), add a GRPO pass with your verifier as the reward. 8. Serve merged + quantized behind vLLM.
Budget: a week of iteration, $100–$300 in compute.
Goal: a small model that reasons carefully about your kind of problem.
1. Start from a reasoning-distilled base (DeepSeek distills or similar). 2. SFT on teacher traces for your problem type so the format is right. 3. Build a verifier: unit tests, a checker script, a rubric with teeth. 4. GRPO with 4–8 rollouts per prompt, a few thousand prompts. 5. Watch for reward hacking from day one; hold out verification cases. 6. Keep the best checkpoint by held-out verifier score.
Budget: the verifier is the real cost (days of engineering); compute $200–$500.
Sources and further reading
Core methods trace to these papers: LoRA (Hu et al., 2021), QLoRA (Dettmers et al., 2023), DPO (Rafailov et al., 2023), RLHF (Ouyang et al., 2022), LIMA (Zhou et al., 2023), DeepSeek-R1 (DeepSeek-AI, 2025), IMPersona (arXiv 2504.04332), InMyStyle (arXiv 2607.29238). Current practice (hyperparameters, VRAM figures, rental pricing, tool comparisons) was gathered from web research in October 2026 at index level, not live-verified. Model releases and prices move fast: re-check leaderboards (LMArena, Open LLM Leaderboard) before choosing a base, and re-check provider price cards before budgeting compute.
- Hugging Face TRL documentation — SFTTrainer, DPOTrainer, GRPOTrainer references.
- Unsloth documentation — memory-efficient LoRA guides and VRAM tables.
- Axolotl and LLaMA-Factory documentation — config-driven training recipes.
- vLLM and SGLang documentation — serving, quantization support, adapter swapping.
- mergekit documentation — SLERP, TIES, DARE model merging.
The full research brief behind this report (method-by-method numbers, per-topic source notes) is kept with the author's notes.