LoRA Pipeline
Fine-tuning is the easy part. Proving it helped is the project.
A QLoRA fine-tune that teaches a frozen 4-bit Qwen2.5-1.5B-Instruct to turn a messy booking request into strict, schema-valid JSON:
"book me a table for 4 at Nobu next friday at 7pm"
→ {"intent":"book","party_size":4,"venue":"Nobu","date":"2025-07-11","time":"19:00"}
The base model is never touched. What comes out the other end is a 10–30 MB adapter, an eval report comparing it against baselines, and a rank sweep that justifies the rank actually shipped.
Those last two artifacts are the reason the repo exists.
The thing most fine-tuning projects skip
Fine-tuning a small model with LoRA is, mechanically, not hard anymore. The libraries are good, the recipes are everywhere, and you can have a training run going in an afternoon. Then the loss curve goes down, and the project ends, and the write-up says the fine-tune worked.
A falling loss curve proves the optimizer did its job. It says nothing about whether the model got better at the task, and it certainly says nothing about whether the fine-tune was necessary — the base model might already have been able to do this, in which case you've spent a GPU session and produced an adapter that adds nothing.
The only honest answer is a metric with baselines next to it. So eval scores the base model and the fine-tuned model on the same held-out set and writes both to eval.json, and sweep runs the whole thing across ranks 4, 8, 16, and 32 so the production rank is a measured choice rather than a number copied from a tutorial. The config carries a default of r=16, and the sweep is explicitly allowed to overrule it.
That's the deliverable. The adapter is almost a byproduct.
Why this task, specifically
Structured extraction is a deliberately good choice for demonstrating a fine-tune, because success is mechanically checkable. Either the output parses against the schema or it doesn't; either party_size is 4 or it isn't. No LLM-as-judge, no rubric, no subjective grading, no argument about whether the tone improved.
Pick a style or personality fine-tune instead and you have a much more interesting demo and no trustworthy way to say whether it worked. I wanted the measurement to be the load-bearing part, which meant choosing a task where measurement is unambiguous.
Splitting the code by where it can run
The architectural decision I'd point at first has nothing to do with machine learning. The codebase is split along a hard line: everything that doesn't need a GPU is separated from everything that does, and the GPU half imports torch and bitsandbytes lazily.
The consequence is that 71 unit tests — schema, metrics, config resolution, the data generator, chat formatting, the loss mask — run on a laptop with no CUDA, no GPU, and torch not even installed.
This exists because of a real constraint: training happens on a free Colab T4, and GPU session time is the genuinely scarce resource in this project. Sessions disconnect, quotas run out, and every trivial bug discovered inside a session — a config typo, a malformed prompt template, a metrics function that divides by zero on an empty prediction — costs a chunk of the budget that was supposed to go to training. Pushing all of that onto the laptop, where iteration is free and instant, is what makes working under that constraint tolerable.
It also produces a much better codebase as a side effect. Forcing the pure logic to be importable without torch means it can't be entangled with the model, which means the schema, the generator, and the formatting are all independently testable — and they're exactly the components where a silent bug does the most damage.
The test that earns its keep
The highest-leverage test in the suite checks that the loss mask is completion-only and that the EOS token is left unmasked.
Both of those are the classic silent failures of instruction fine-tuning, and they share a nasty property: neither one throws an error. Training completes. Loss goes down. You get an adapter.
If loss is computed over the prompt tokens as well as the completion, the model spends its capacity learning to generate booking requests alongside JSON, diluting the signal you actually wanted. If EOS is masked out, the model never learns to stop — it emits perfectly valid JSON and then keeps going, which breaks strict parsing on every single example and looks, from the loss curve, like a successful run.
You find both of these after burning a session, by staring at outputs and wondering why the numbers are mediocre. Or you find them in half a second on your laptop before Colab is ever opened. The test asserts on the mask array directly, which is the only place the bug is visible before it becomes an expensive mystery.
Reproducibility choices
The repo carries the generator, not the dataset. Training data is produced by a deterministic synthetic generator from a seed, and the JSONL files are git-ignored. Anyone cloning the repo regenerates a byte-identical dataset rather than downloading one, and there's no large binary drifting out of sync with the code that made it.
The sweep is resumable. Free Colab sessions drop, and a four-rank sweep is long enough that this will happen. Re-running the sweep cell skips ranks already recorded in sweep.json rather than starting over. Small feature; unmistakable sign of code that was actually run rather than merely written.
The notebook contains no logic. It's a thin driver that checks the GPU, clones, installs pinned dependencies, mounts Drive, and calls the CLI. All the real work lives in version-controlled, tested modules. Colab notebooks are where reproducibility usually goes to die — cells run out of order, state that only exists in one session, logic that can't be tested. Keeping the notebook as a launcher sidesteps all of it.
One config file. Every tunable in configs/default.yaml, with precision auto-detected (fp16 on a T4) rather than hardcoded. Alpha is held at 2r across the entire sweep, so the sweep varies one thing and the comparison means something.
What I'd challenge about it
Synthetic training data means the eval is graded on home turf. The generator produces the training set and the held-out set, so both come from the same distribution. However good the reported numbers are, they measure performance on the generator's idea of a messy booking request — not on how people actually type. Real input has typos, ambiguity, mid-sentence corrections, multiple bookings in one message, and phrasings no generator template anticipated. A hand-labelled set of a few hundred real utterances would be the single most valuable addition to this project, and it would probably lower the reported score, which is exactly why it's worth doing.
Schema-valid and correct are two different things, and only one is easy to measure. "next friday at 7pm" has to resolve to an actual date relative to some reference day. A 1.5B model doing relative date arithmetic will sometimes produce output that parses perfectly against the schema and is off by a week. Validity is a floor, not the metric — the eval needs to separate parses from every field is right, because a system that scores 100% on the former can be badly wrong in production.
Constrained decoding is the baseline that has to be beaten. Grammar-constrained or structured-output decoding guarantees schema-valid JSON from an un-finetuned model, for free, with no training at all. That means schema validity is not where the fine-tune's value can come from — it has to show up in field-extraction accuracy, and specifically in the hard cases: implied party size, relative dates, venue names that look like ordinary words. Adding constrained-decoding to the baseline column would make the eval considerably more honest and considerably more convincing.
There's no abstention path. Feed it something that isn't a booking request and it will confidently emit JSON anyway. Greedy decoding gives no usable confidence signal, and nothing in the pipeline decides not to answer. Any real deployment needs that, and it's a harder problem than the extraction itself.
Stack
QLoRA (nf4 with double quantisation) over a frozen Qwen2.5-1.5B-Instruct, LoRA applied to all seven linear projections at r=16/alpha=32, lr 2e-4, 3 epochs, effective batch 16 via accumulation, 256-token sequences, greedy decoding at eval. Trained on a free Colab T4. CLI with gen-data / train / eval / sweep / infer, 71 CPU-only unit tests, and an optional 30-line FastAPI serving demo with its dependencies deliberately kept out of the Colab install.