Post-Training Guide / research report

Research report · October 2026

Post-Training Your Own Models

How to take an open model and teach it a custom task, a workflow, or your own writing voice. Methods, data, compute, tooling, evaluation, and playbooks you can run this week.

Scope: SFT · LoRA/QLoRA · DPO/RLHF · distillation · RL post-training Audience: technical, acting alone or on a small team Budget range: $0–$500 per project

1. Do you even need to fine-tune?

Fine-tuning is the heavy answer. Most problems have lighter ones. Run through this list before spending a dollar on training.

Rule of thumb

If you can describe the desired behavior in under a page of instructions and the model follows it, stop there. Fine-tune when you need the behavior baked in: always-on, no prompt to maintain, cheaper to serve, and robust to prompt injection drift.

2. The method map

Post-training means changing a model's weights after pretraining. The methods below stack. A common pipeline is: base model → SFT on demonstrations → preference alignment (DPO) → quantize and serve.

MethodWhat it doesWhen to use itCompute
Full SFTUpdates every weight on input→output examplesLarge domain shift; you have serious data and GPUsHigh
LoRA / QLoRATrains small adapter matrices; base weights stay frozenDefault choice for tasks and style; single-GPU friendlyLow
DPO / ORPO / KTOAligns on chosen vs rejected response pairs, no reward modelPolishing tone, format, and preference after SFTLow–medium
RLHF (PPO)Trains a reward model, then optimizes against itLarge labs; rarely worth it for individuals nowHigh
GRPO / RLVRReinforcement learning with verifiable rewards (math, code, tools)Reasoning tasks with checkable answersMedium–high
DistillationA small student learns from a large teacher's outputsCompressing capability into a cheap servable modelMedium
Model mergingCombines weights of two tuned models, no trainingBlending skills (e.g. style adapter + task adapter)Near zero

3. Supervised fine-tuning (SFT)

SFT is the foundation. You show the model thousands of input→output examples and train it to predict the outputs. Everything else in this guide either replaces SFT's full-weight update with something cheaper or polishes the result afterward.

How it works, in one paragraph

Each training example is a prompt plus the desired response. The model reads the prompt, predicts the response token by token, and the loss measures how far its predictions were from the actual tokens. Gradient descent nudges the weights to close the gap. Repeat for 1–3 passes over the data (epochs). That is the whole trick.

Full SFT vs adapters

Full SFT updates every parameter. For a 7B model in 16-bit precision that means ~14 GB just for the weights, plus gradients and optimizer state. Adam keeps a master copy plus two running statistics per parameter, so full fine-tuning needs roughly 16 bytes per parameter at 16-bit: about 112 GB for 7B. That is why individuals rarely do full SFT on anything above 3B without serious hardware. Section 4 covers the standard workaround.

Hyperparameters that actually matter

The one failure mode to respect

Catastrophic forgetting. SFT on narrow data erodes the model's general ability. A model trained only on your support tickets will get worse at everything else. Mix in 10–30% general instruction data to keep the base capabilities intact, and check both directions: your task metrics and a general benchmark.

4. LoRA and QLoRA: the default for everyone else

LoRA (Low-Rank Adaptation) freezes the base model and trains two small matrices per layer whose product approximates the weight change. Instead of updating a 4096×4096 weight matrix (16.8M numbers), you train a 4096×16 and a 16×4096 pair (131K numbers). Same architecture, 0.1–1% of the parameters.

Why LoRA usually wins for individuals

QLoRA: the 4-bit variant

QLoRA combines LoRA with a 4-bit quantized base (usually NF4) plus paged optimizers. It is the reason a single GPU can fine-tune 7B–13B models at all. The finding that launched it: 4-bit quantization barely dents fine-tuning quality. Use QLoRA unless you have a reason not to.

Settings that work

When overfitting hits

Overfitting is the number one LoRA failure mode: training loss falls while held-out scores stall. Fix in this order: reduce rank, raise dropout, cut epochs, add data. Do not reach for a bigger rank first.

Variants worth knowing

DoRA splits magnitude and direction of the update and slightly improves quality at small extra cost. rsLoRA fixes a scaling quirk so higher ranks actually help. LoRA+ uses different learning rates for the two adapter matrices. None of these change the workflow; they are drop-in upgrades in most libraries.

5. Preference alignment: RLHF, DPO, and friends

SFT teaches the model what to say. Preference methods teach it which answer is better. You give pairs: prompt, chosen response, rejected response. The model learns to prefer the chosen one.

DPO killed the RLHF pipeline for most people

Classic RLHF trains a separate reward model on human preferences, then optimizes the policy against it with PPO. Three models, finicky hyperparameters, frequent reward hacking. DPO (Direct Preference Optimization) skips all of it: a single loss function that pushes probability toward the chosen response and away from the rejected one. Same data format, one training run, stable optimization. For style polish and format compliance, DPO after SFT is the standard stack.

The family

What DPO cannot do

DPO is a gentle nudge, not a capability builder. It re-ranks behavior the model already has; it does not teach new skills. If the base model cannot do the task at all, fix that with SFT or distillation first.

Length bias

Preference data is full of an accidental signal: chosen responses tend to be longer. Models learn to ramble. Counter it by keeping chosen and rejected responses similar in length, or use a length-normalized method like SimPO.

Where do preference pairs come from? Generate two responses per prompt from your SFT model, then rank them: by hand for a few hundred (gold standard), with a stronger model as judge (cheap, slightly noisy), or from implicit signals like user edits and thumbs votes. A few hundred to a few thousand pairs is enough for a noticeable polish pass.

6. RL post-training: GRPO and RLVR

The biggest post-training story of the last two years is reinforcement learning with verifiable rewards. Instead of a learned reward model guessing what humans like, you use rewards you can check: the math answer is right, the code passes its tests, the tool call returned the right schema. DeepSeek-R1-Zero showed what this unlocks: pure RL on a base model took AIME math scores from 15.6% to 77.9%, with no human reasoning examples at all.

Why it matters

RLVR (reinforcement learning with verifiable rewards) is what made small models reason. DeepSeek-R1 showed that RL on checkable problems produces long chains of self-correction: trying, failing, backtracking. SFT can imitate reasoning traces, but RL discovers them. If your custom task has right answers you can verify programmatically, RL post-training beats more SFT.

GRPO in one paragraph

GRPO (Group Relative Policy Optimization) samples a group of responses to the same prompt, scores each with your verifier, and pushes the policy toward the better ones relative to the group average. No value network, no reward model, which makes it far cheaper than PPO. Libraries (TRL, OpenRLHF, veRL) now ship GRPO trainers you can point at your own reward function.

When to use it

Reward hacking is the tax

The model will find the cheapest way to score, not the way you meant. A code reward based on passing tests produces code that games the tests. Keep verifiers strict, hold out test cases the model never sees during training, and spot-check rollouts by hand. If your reward can be gamed, assume it will be.

7. Knowledge distillation

Distillation trains a small student on the outputs of a large teacher. The teacher generates answers (or reasoning traces, or full token probability distributions), and the student learns to match them. You get much of the teacher's skill in a model you can actually serve.

Three flavors

The honest version of the deal

Distillation transfers behavior, not knowledge the teacher does not reliably have. A student distilled from a teacher's shaky domain knowledge inherits the shakiness. Distill reasoning patterns and style freely; distill facts cautiously.

8. Choosing a base model

The base model sets your ceiling. Pick wrong and no amount of tuning recovers it. Three criteria dominate: capability at your size, license, and context length.

The current open-weight landscape

The 2026 default base is Qwen3: strongest performance per parameter in the open field, Apache 2.0 license, long context, and a large fine-tuning community. Llama 3.3 remains the safe generalist if you want the Meta ecosystem and tooling.

FamilySizesLicenseNotes
Qwen30.5B–72B+Apache 2.0Default fine-tune base; coding, agents, multilingual
Llama 3.31B–70BLlama Community LicenseExcellent ecosystem; license has use restrictions, read before productizing
DeepSeek1.5B–70B+ (distills)Often MIT for distillsReasoning distills are the best cheap starting point for RL work
Gemma1B–27BGoogle terms; check the model cardStrong at small sizes; licensing is inconsistently reported, verify per checkpoint
Mistral7B–24B+Apache 2.0 / MNPLEfficient architectures; check license per release
gpt-oss (OpenAI)20B / 120BApache 2.0OpenAI's open entry; surprisingly strong
Phi3B–14BMITMicrosoft's small models; punchy for their size

How to choose

Recency

Model releases move monthly. Before committing, check the current leaderboards (LMArena, Open LLM Leaderboard) at your size class. A model from six months ago is often strictly worse than this month's release at the same size.

9. Building the dataset

Data quality decides the outcome more than any hyperparameter. The repeated finding across the literature: a few thousand excellent examples beat a hundred thousand mediocre ones. The LIMA result is the landmark: 1,000 meticulously curated pairs on a 65B model beat 52,000 lower-quality examples. The refinement since: style saturates at around 100 examples, while new capabilities keep improving with more data.

Format

Most trainers expect instruction pairs or conversations. The modern standard is the chat template: a list of messages with roles.

{"messages": [
  {"role": "user", "content": "Summarize this quarter's churn drivers."},
  {"role": "assistant", "content": "Three drivers explain 80% of churn..."}
]}

Match the base model's own chat template exactly. Every major library applies it automatically; the mistake is hand-rolling a format the model never saw.

How much data

Where the data comes from

Cleaning checklist

10. Mimicking a writing style

This is the most personal use case: a model that writes like you. The good news is that style is one of the cheapest things to transfer. It saturates at around 100 examples, and 2026 work (InMyStyle) showed per-user LoRA adapters working on models as small as 0.5B, with judges rating the outputs over 20% less "AI-sounding." You do not need a giant model to clone a voice.

The pipeline

  1. Collect 200–500 samples of your writing in the target register to start; iterate toward 1–2,000 for polish. Emails if you want email voice; long-form if you want essay voice. One register per adapter; mixing registers blurs the result. Consistency of voice matters more than volume.
  2. Pair each sample with a plausible prompt. A raw essay is not training data. Write the instruction that would have produced it: "Write a memo to the team about the Q3 delay" → your actual memo. For emails, the prompt can be "Reply to this thread: ..." with the thread summarized. A strong model can draft these prompts for you; review them.
  3. Scrub. Remove anything you do not want the model to reproduce: names, private details, company internals, and your bad habits. The model will faithfully learn your typos and your tics. Curate like an editor.
  4. Train LoRA on the instruct checkpoint, not the base. The instruction-tuned conversational priors are worth preserving; full fine-tuning erodes them (researchers observed random topic-switching after full FT on personal corpora). Start at rank 8–16; rank 4 suffices for light tone shifts.
  5. Keep a short style system prompt too. The adapter and the prompt stack: the prompt gets you most of the way for free, the adapter makes it consistent across thousands of calls without eating context.
  6. Evaluate blind. Generate passages from prompts the model never saw. Mix them with your real writing. Ask someone who knows your voice to sort them, or run pairwise A/B judging ("which sounds more like X"). If they cannot tell reliably, you are done. Track a general benchmark alongside to catch capability regression.
Why LoRA preserves style better

Full fine-tuning rewrites the model's general language habits along with everything else, which often washes out the crisp edges of a personal voice. LoRA's small updates steer the existing capability instead of replacing it, so your phrasing habits come through while the model's grammar and reasoning stay intact.

What does not work

11. Compute and cost

The memory math

Three numbers dominate GPU memory during training: the weights, the gradients, and the optimizer state.

SetupTypical VRAMHardware
QLoRA, 7–8B model10–12 GBSingle RTX 4090 (24 GB)
QLoRA, 13B model~16 GBRTX 4090, or single 48 GB card
QLoRA, 32–33B model~24 GBRTX 5090 (32 GB) / A100 40 GB
QLoRA, 70B model~48 GBSingle A100/H100 80 GB
Full SFT, 7B model~112 GB2× 80 GB or multi-GPU with sharding
Full SFT, 70B model~1.1 TBMulti-node cluster; not an individual project

Renting GPUs

You do not need to own hardware. Spot and on-demand GPU rentals put a card in your hands in minutes. Indicative September 2026 rates on RunPod: RTX 4090 around $0.74/hr ($0.34 community/spot), A100 80GB around $1.59/hr, H100 around $3.49/hr; Vast.ai runs cheaper but host reliability varies. A typical 7B QLoRA run takes about 45 minutes on a 4090-class card: a few dollars per experiment. The expensive part is iterating ten times, so budget for iteration, not for one run. Prices move constantly; check the provider's live pricing before you commit.

Free options

Google Colab (free tier) gives you a modest GPU with time limits, enough for small LoRA experiments on 1–3B models or short 7B runs. Kaggle offers weekly free GPU hours. Both are fine for learning the workflow before you spend money.

Budget for the loop, not the run

Your first training run will not be your last. Plan for 5–15 experiments: data fixes, hyperparameter sweeps, and evaluation iterations. A realistic personal project budget is $50–$300 in compute for a 7B LoRA project, most of it spent on experiments 2 through 10.

12. The tooling stack

ToolWhat it isUse it when
Hugging Face TRLThe reference trainers: SFT, DPO, GRPO, KTOYou want standard, maintained implementations
UnslothHand-optimized kernels; 2–5× faster and ~70% less VRAM than the Hugging Face baselineSingle consumer GPU. The individual-developer default for LoRA/QLoRA
AxolotlYAML-config training for many methods, multi-GPU with FSDP2Reproducible configs, multi-GPU, long context
LLaMA-FactoryWeb UI + CLI for SFT/DPO/RLYou prefer a GUI over scripts
torchtune / LitGPTClean PyTorch-native training stacksYou want readable code you can modify
OpenRLHF / veRLFull RLHF/RL pipelinesYou are doing PPO or GRPO at scale
vLLM / SGLangHigh-throughput inference serversvLLM: production GPU serving with runtime LoRA-adapter swapping (one server, many adapters). SGLang: wins on agent/RAG workloads with repeated prefixes and heavy structured output
Ollama / llama.cppLocal quantized inferenceRunning the model on your own machine
Hugging Face HubModel and dataset hosting, private reposStoring adapters, sharing, versioning

No-code and API options

Several providers offer fine-tuning as a service: OpenAI (about $0.25 per million training tokens), Together AI, Fireworks AI (from about $0.50 per million training tokens), and others let you upload a dataset and get back a tuned model behind an API. You trade control and data privacy for zero infrastructure. A million-token SFT job costs pocket change in training tokens; per-token inference pricing dominates the long-term bill. Read the terms on data retention before uploading anything sensitive, and never train on a provider's outputs against their terms. For style mimicry on non-sensitive text, these are genuinely the fastest path.

# The shortest real training loop (Unsloth + TRL), conceptually:
# 1. pip install unsloth trl datasets
# 2. Load a 4-bit base, attach LoRA adapters (rank 16, all linear layers)
# 3. Train on your chat-formatted JSONL for 1-3 epochs, lr 2e-4
# 4. Save the adapter (a few hundred MB), merge or serve with vLLM

13. Evaluation

Training without evaluation is guessing. Decide before training what "better" means and how you will measure it.

Three layers

Checkpoint discipline

The style eval that works

For voice mimicry, automated metrics fail. The test is blind human sorting: can a reader who knows your writing distinguish your text from the model's? Run it on fresh prompts, not training examples. Fifty passages and one honest reader beats any BLEU score.

14. Serving the finished model

A trained adapter is not a product until it runs somewhere cheap and fast. The standard move is quantization: shrinking the weights so inference needs less memory and runs faster, with minimal quality loss.

FormatUse case
GGUFLocal and edge inference via llama.cpp and Ollama. The universal CPU/laptop format.
AWQ / GPTQGPU inference with 4-bit weights. Pairs with vLLM for serving. AWQ is slightly better at 4-bit; both need calibration data.
bitsandbytes 8-bitQuick local experiments without a conversion step.

15. Pitfalls and honest caveats

16. Three playbooks

Playbook A · Write like me (style mimicry)

Goal: a model that drafts in your voice.

1. Collect 200–500 samples of your writing in one register to start; iterate toward 1–2,000 for polish. 2. Write a plausible prompt for each (a strong model can draft these; you review). 3. Scrub names, secrets, and tics you dislike. 4. Format as chat JSONL, hold out 10%. 5. QLoRA on a 7–8B instruct model: rank 8–16, lr 2e-4, 2–3 epochs. 6. Generate on fresh prompts; blind-sort against your real writing. 7. If it is close but loose, add 200–500 DPO pairs (your edit preferred over the model's draft). 8. Quantize to GGUF, run in Ollama.

Budget: one evening of curation, $10–$50 in GPU rental.

Playbook B · A custom workflow (task specialist)

Goal: a model that executes your multi-step workflow reliably, e.g. triaging tickets into your taxonomy with your fields.

1. Define the task as input→output with a strict schema. 2. Build 2,000–10,000 examples: distill from a strong model on your real inputs, then hand-fix the failures. 3. Mix in 10% general instruction data. 4. QLoRA rank 32–64, lr 1e-4, 1–2 epochs. 5. Eval on 300+ held-out cases with programmatic checks (schema valid, fields correct). 6. Mine the failures, add targeted examples, retrain. 7. If outputs are checkable (tests, schemas), add a GRPO pass with your verifier as the reward. 8. Serve merged + quantized behind vLLM.

Budget: a week of iteration, $100–$300 in compute.

Playbook C · Reasoning on your domain

Goal: a small model that reasons carefully about your kind of problem.

1. Start from a reasoning-distilled base (DeepSeek distills or similar). 2. SFT on teacher traces for your problem type so the format is right. 3. Build a verifier: unit tests, a checker script, a rubric with teeth. 4. GRPO with 4–8 rollouts per prompt, a few thousand prompts. 5. Watch for reward hacking from day one; hold out verification cases. 6. Keep the best checkpoint by held-out verifier score.

Budget: the verifier is the real cost (days of engineering); compute $200–$500.

Sources and further reading

Core methods trace to these papers: LoRA (Hu et al., 2021), QLoRA (Dettmers et al., 2023), DPO (Rafailov et al., 2023), RLHF (Ouyang et al., 2022), LIMA (Zhou et al., 2023), DeepSeek-R1 (DeepSeek-AI, 2025), IMPersona (arXiv 2504.04332), InMyStyle (arXiv 2607.29238). Current practice (hyperparameters, VRAM figures, rental pricing, tool comparisons) was gathered from web research in October 2026 at index level, not live-verified. Model releases and prices move fast: re-check leaderboards (LMArena, Open LLM Leaderboard) before choosing a base, and re-check provider price cards before budgeting compute.

The full research brief behind this report (method-by-method numbers, per-topic source notes) is kept with the author's notes.