Research Notes
Small updates break first
I fine-tuned small language models, quantized them the way people ship them, and measured how much of the gain survived. In all ten pre-registered tests at MLX 3-bit and GGUF Q3_K_M, the gentle LoRA kept less of its gain than the strong one, and gentle is mlx-lm’s default learning rate.
Abstract
Small open models are often fine-tuned and then shipped at 3 to 6 bits per weight, as MLX or GGUF files, on the assumption that the fine-tuning gain survives. I tested that assumption in five pre-registered rounds on one laptop: Qwen3 at 0.6B, 1.7B and 4B and OLMo-2 1B, fine-tuned on banking77 intent classification and MASSIVE slot annotation, and evaluated on their full test sets at bf16 and up to ten quantized formats. At MLX 3-bit and GGUF Q3_K_M, a LoRA trained at learning rate 1e-5, mlx-lm’s default, kept less of its gain than the same LoRA at 1e-4 in all 10 pre-registered tests, by 0.06 to 0.63 in gain retention. In exploratory comparisons over the eight settings where both were trained, the 1e-5 LoRA was the more accurate model at full precision and kept a smaller share of that accuracy at every one of the five formats compared, 40 of 40. Which one to ship depends on the file, also exploratory: the 1e-5 LoRA was more accurate at Q4_K_M in all eight settings, the 1e-4 LoRA at MLX 3-bit and Q2_K in all eight. In MLX the loss does not come from rounding the update. Keeping it exact on OLMo-2 1B at MLX 3-bit recovered 5% and 12% of the lost gain for two seeds, while training the 1e-5 LoRA on the 4-bit Qwen3-0.6B base delivered 96% of the bf16 gain, against 84%. A predictor built from the weights met its pre-registered error bar, but a post hoc model of the update’s size and the base damage did better.
Fine-tune a small model, then ship it as a 4-bit MLX model for Apple silicon or as a GGUF file for llama.cpp and the tools built on it. The usual assumption is that the gain you trained for comes along. I wanted to know when it does not, and which fine-tunes lose it.
The short version: the fine-tunes that lost the most were the gentle ones, with small learning rates and small weight updates. In the first round, the most accurate model at full precision kept 6% of its gain at MLX 3-bit.
Four base models, none larger than 4B: Qwen3 at 0.6B, 1.7B and 4B, and OLMo-2 1B. Two tasks with short answers scored by exact match. LoRA at rank 8 and scale 20, plus full fine-tuning on Qwen3-0.6B, for 500 steps each. MLX affine quantization at group size 64 and GGUF k-quants without an importance matrix, so no GPTQ or AWQ. Everything ran on one Apple M1 Pro laptop with 16 GB of memory. Every number below comes from the paper in the ftquant repository, where a script recomputes each one from the committed per-item predictions, run logs and result files. Larger models, long-form generation and calibrated formats may behave differently.
Two ways to count what survives
The measure I pre-registered is gain retention. Take the fine-tune’s accuracy minus its base model’s accuracy, with both quantized to the same format, and divide by the same difference at full precision. A value of 1 means the fine-tune is as far ahead of an equally quantized base as it was in bf16.
Gain retention can go above 1 when quantization costs the base more than the fine-tune, and at low bit widths that happened for every base that started with real zero-shot skill.1 So from round 2 on I also report accuracy retention: the fine-tune’s quantized accuracy as a share of its own bf16 accuracy, which does not involve the base at all.
I call a LoRA trained at learning rate 1e-5 gentle and one trained at 1e-4 strong. The gentle setting is mlx-lm’s default learning rate at its default LoRA scale, and in every pair the strong LoRA’s update was 8.7 to 12.9 times larger.2
The most accurate model kept 6%
The first round’s exploratory runs, on Qwen3-0.6B and banking77, showed the pattern everything after them tests. Of three fine-tunes, the gentle LoRA was the most accurate at full precision: 78.5%, against 76.6% for full fine-tuning at the same learning rate and 73.5% for the strong LoRA. At MLX 3-bit it kept 6% of its gain, and full fine-tuning kept 14%. The strong LoRA kept 95%, and 101% at Q3_K_M.
(opens in a new tab)
At MLX 3-bit the Qwen3-0.6B base scores 0% on its own, and the strong LoRA still keeps 95% of its gain. A large update can carry a broken base.
The bits per weight on that axis are measured from the files, and they are worth knowing on their own. GGUF’s k-quants spend more than their names suggest: Q3_K_M comes out at 4.2 to 4.4 bits per weight, closer in size to MLX 4-bit (4.5) than to MLX 3-bit (3.5). Across engines, compare formats by measured size, not by the number in the name.
Lower learning rates lost more
Round 2 swept the learning rate on Qwen3-0.6B. Within each kind of fine-tuning, gain retention rose with the learning rate in all eight combinations of fine-tuning kind and format I tracked at 3 and 4 bits, apart from one tie. Each has only three or four learning rates, so none is strong evidence alone. At MLX 3-bit the LoRA went from 0% at 3e-6 to 6%, 63% and 95% as the learning rate rose to 1e-4. Full fine-tuning went from 0% to 14% and 78%.
(opens in a new tab)
Learning rate is the knob I turned, and within LoRA it moves everything at once: the size of the update and whatever else a learning rate changes about the solution. So “small updates” in the title is shorthand for that bundle, not a claim that size alone is the cause. Across kinds of fine-tuning, the update’s size does not order retention. Full fine-tuning at 3e-5 kept 0.78 of its gain at MLX 3-bit and 0.89 at Q2_K with an update of norm 18.6, while the LoRA at 3e-5, with a larger update of 25.2, kept 0.63 and 0.01.
A second seed moved gain retention at MLX 3-bit by at most 0.061, while bf16 accuracy moved by 8.4 and 3.5 points. On Qwen3-1.7B the gentle LoRA kept 36% of its gain at MLX 3-bit, and in accuracy retention, which its collapsing base does not inflate, 0.32 against 0.93 for the strong one.
Ten out of ten
Rounds 3 and 4 turned that observation into confirmatory tests. Each asks whether the strong LoRA keeps more of its gain than the gentle one, by a paired bootstrap over test items at MLX 3-bit and at Q3_K_M, and passes only if the 95% interval lies above zero. Round 3 moved to a new model family, OLMo-2 1B on banking77, and to a new task, Qwen3-0.6B on MASSIVE slot annotation. Round 4 repeated both with a second seed and added Qwen3-4B. Here are all ten, as gain retention for the strong and the gentle LoRA:
| setting | format | strong | gentle | difference [95% CI] |
|---|---|---|---|---|
| OLMo-2 1B, banking77 | MLX 3-bit | 1.092 | 0.826 | +0.265 [+0.236, +0.294] |
| OLMo-2 1B, banking77 | Q3_K_M | 1.076 | 1.018 | +0.058 [+0.037, +0.078] |
| Qwen3-0.6B, MASSIVE | MLX 3-bit | 0.621 | 0.319 | +0.301 [+0.271, +0.332] |
| Qwen3-0.6B, MASSIVE | Q3_K_M | 0.946 | 0.843 | +0.102 [+0.078, +0.127] |
| OLMo-2 1B, banking77, seed 1 | MLX 3-bit | 1.105 | 0.813 | +0.292 [+0.263, +0.322] |
| OLMo-2 1B, banking77, seed 1 | Q3_K_M | 1.064 | 1.002 | +0.061 [+0.042, +0.079] |
| Qwen3-0.6B, MASSIVE, seed 1 | MLX 3-bit | 0.685 | 0.052 | +0.633 [+0.604, +0.659] |
| Qwen3-0.6B, MASSIVE, seed 1 | Q3_K_M | 0.916 | 0.777 | +0.138 [+0.111, +0.168] |
| Qwen3-4B, banking77 | MLX 3-bit | 1.464 | 0.925 | +0.539 [+0.451, +0.642] |
| Qwen3-4B, banking77 | Q3_K_M | 1.075 | 0.871 | +0.204 [+0.153, +0.256] |
All ten passed, and all ten intervals also clear zero at the Bonferroni-corrected level of 99.5%. They are five fine-tune pairs at two correlated formats, and each interval describes one trained pair.
The sizes are not comparable across settings. Gain retention divides by each fine-tune’s own gain, and a base that degrades inflates the difference in the strong LoRA’s favor: for Qwen3-4B at MLX 3-bit, 0.155 of the 0.539 comes from the base’s own drop. So, as an exploratory check, I repeated each comparison on a quantity that leaves the base out, how much more accuracy the gentle LoRA lost than the strong one. It was positive in all ten, by 3.6 to 41.8 points.
(opens in a new tab)
These comparisons are exploratory. In all eight of those settings the gentle LoRA had the higher bf16 accuracy and kept a smaller share of it at each of the five formats compared, MLX 4-bit and 3-bit and GGUF Q4_K_M, Q3_K_M and Q2_K: 40 comparisons, all in the same direction. Two cautions apply. The eight settings are not independent, since three are second seeds of others. And at Q4_K_M the gaps are small: six of the eight are under 2 points, and for two of them the interval includes zero.
Two checks argue against reading the gentle LoRAs as simply undertrained. Their final validation loss was no higher than the strong LoRAs’ in any of the eight pairs. And the loss is not confined to borderline items: for Qwen3-4B at MLX 3-bit, among the test items each model answered correctly at bf16 with confidence of at least 0.9, the gentle LoRA still answered 89.6% correctly after quantization and the strong one 99.8%.
Qwen3-4B differs in degree. Its base does not break at MLX 3-bit, and the gap between its two LoRAs is smaller there, 0.80 against 0.94 in accuracy retention, where Qwen3-0.6B showed 0.06 against 0.85. At Q2_K, where its base does break, the gentle 4B LoRA still lost 40% of its accuracy. With one pair per size, that is not a trend with size.
Seeds matter too. On OLMo the second seed reproduced closely, moving gain retention by at most 0.016 at MLX 3-bit and Q3_K_M. On MASSIVE the direction held, but the gentle LoRA’s gain retention at MLX 3-bit went from 0.32 with seed 0 to 0.05 with seed 1. The accuracy and retention intervals in this post cover test-set sampling only, not variation between training runs, and that pair is the reminder.
Which one to ship depends on the file
Retention is not what you ship. Accuracy is. Ranked by accuracy after quantization, in another exploratory comparison, the answer depends on the format:
| format | settings where the strong LoRA is more accurate |
|---|---|
| MLX 3-bit | 8 of 8 |
| GGUF Q2_K | 8 of 8 |
| GGUF Q3_K_M | 7 of 8 |
| MLX 4-bit | 5 of 8 |
| GGUF Q4_K_M | 0 of 8 |
At Q4_K_M the gentle LoRA’s higher starting accuracy outweighs its larger loss in every setting. At MLX 3-bit and Q2_K the strong LoRA wins every time, and at Q3_K_M in seven of eight. Its price was 1.5 to 9.9 points of accuracy at bf16. On Qwen3-0.6B and banking77 the two ends look like this: at Q4_K_M, 74.6% for the gentle LoRA against 73.0% for the strong one; at MLX 3-bit, 4.3% against 62.5%.
So the practical rule is short. Choose a fine-tune by evaluating the quantized file you will actually ship, not the bf16 checkpoint it came from.
In MLX, it is not the rounding
My first account of the mechanism was about rounding. Quantization snaps weights to a grid, and an update that is small next to the grid’s step either rounds away or drowns in the rounding noise. The weight measurements fit that story. For each fine-tune and format I measured how much of the update survives quantization along its own direction and how much noise lands on top of it, and in rounds 1 and 2, when the base was intact and most of the update survived, all 44 pairs with noise below three times the update kept at least 90% of their gain. I chose those bands after seeing the data, so they describe it and test nothing.
But round 1 already held evidence against it, in runs my first draft did not report. The gentle Qwen3-0.6B LoRA had also been evaluated unfused, with its update kept exactly in bf16 on top of the quantized base, and it lost about as much as the fused model: gain retention 0.83 unfused against 0.84 fused at MLX 4-bit. An update that is never quantized cannot be rounded away. A review of the draft pointed that out, so I ran a fifth round to test it, under a plan pushed to GitHub before anything ran.
| test | format | trained in bf16, fused | alternative | difference [95% CI] |
|---|---|---|---|---|
| OLMo-2 1B gentle, unfused | MLX 3-bit | 0.826 | 0.835 | +0.008 [-0.010, +0.027] |
| OLMo-2 1B gentle, seed 1, unfused | MLX 3-bit | 0.813 | 0.836 | +0.023 [+0.002, +0.042] |
| Qwen3-0.6B gentle, trained on the 4-bit base | MLX 4-bit | 0.841 | 0.957 | +0.116 [+0.097, +0.135] |
| Qwen3-0.6B gentle, trained on the 3-bit base | MLX 3-bit | 0.061 | 0.951 | +0.890 [+0.865, +0.916] |
| Qwen3-0.6B strong, trained on the 4-bit base | MLX 4-bit | 0.940 | 0.781 | -0.158 [-0.181, -0.136] |
The first three rows are the pre-registered tests; the last two are exploratory.
The first test kept the update exact. On OLMo-2 1B at MLX 3-bit, where the base stays usable and the fused gentle LoRAs kept only 0.83 and 0.81 of their gain, loading the adapter unfused recovered 5% and 12% of the lost gain for the two seeds, with upper 95% bounds of 16% and 22%. The plan counted that as no rescue if each seed’s upper bound stayed below +0.08 in gain retention, and both did.
The second test trained against the weights that ship. A gentle LoRA trained directly on the 4-bit Qwen3-0.6B base, the QLoRA setup, delivered 96% of the bf16 gain at 4 bits, against 84% for the same LoRA trained in bf16, fused and quantized. Its accuracy at 4 bits, 77.6%, was within a point of what the bf16-trained LoRA reached at bf16, 78.5%. On the 3-bit base, which scores 0% on its own, the same recipe reached 67.5% and delivered 95% of the bf16 gain, against 6% for fuse-then-quantize.
The remedy has a cost. At the strong learning rate, training on the quantized base made things worse at 4 bits: that LoRA trained less well there and delivered 78% of the bf16 gain, against 94% when trained in bf16, fused and quantized. The most accurate 4-bit and 3-bit models in this comparison were the gentle LoRAs trained on the quantized base.
So in these MLX tests the loss followed the base rather than the update. A gentle update trained against the bf16 weights did not survive those weights moving onto the quantized grid, even when the update itself was kept exact. For gentle LoRAs, the remedy is the one QLoRA builds in and LoftQ prepares for at initialization: train against the weights that will ship. GGUF has no such training path in my pipeline, so none of this was tested there.
Measuring it before you quantize
The tool that came out of the study, ftquant check, compares a fine-tune’s weights with its base and runs both through the exact quantizer you plan to ship, with numpy on a CPU. It reports how much of the update survives, how much noise lands on top of it, how much the format damages the base model itself, for the three bases it ships measurements for, and, where it has been tested, how much of the gain it expects you to keep. This is its real output for a Qwen3-0.6B LoRA trained on MASSIVE at learning rate 1e-5, a fine-tune its predictor had never seen:
| format | noise/update | update kept | base damage | predicted gain kept | reading |
|---|---|---|---|---|---|
| MLX 8-bit (group 64) | 0.47 | 100% | 0.004 | n/a | low noise |
| MLX 6-bit (group 64) | 1.59 | 100% | 0.021 | 100% | low noise |
| MLX 4-bit (group 64) | 3.80 | 100% | 0.255 | 97% | moderate noise |
| MLX 3-bit (group 64) | 5.52 | 99% | 1.207 | 41% | high noise, base model breaks |
Measured on 2,974 test items, it kept 93% of its gain at MLX 4-bit and 32% at 3-bit.
The predicted column comes from a three-parameter model of those measurements, fitted on rounds 1 and 2 and frozen before it was tested. The first version missed its pre-registered error bar by 0.007, so, as that plan required, the tool fell back to descriptive output. The second met its bar on both held-out rounds, with mean absolute errors of 0.059 and 0.084 against a bar of 0.10, though in round 4 the interval reached 0.146.
After the pre-registered tests, I fitted a model that only knows the size of the update and the base damage, post hoc and by the same procedure, and it did better: 0.049 in both rounds. The weight measurements describe what happens, but they have not been shown to add anything beyond how big the update is. The tool prints a prediction only for the formats and base models it was tested on, and not for Qwen3-4B, where the predictor’s error was 0.115.
Confidence moves before accuracy
One more result, exploratory but worth knowing if you gate outputs on confidence. Fit a threshold on the full-precision model so that the predictions it accepts automatically are at most 5% wrong, then quantize and keep the threshold. At low bit widths, coverage falls much faster than accuracy. For the round-1 Qwen3-1.7B LoRA, MLX 3-bit kept 93% of the accuracy but cut the share of accepted items from 38.1% to 27.8%, and Q2_K kept 85% of the accuracy but cut it from 36.4% to 10.1%.
The gentle LoRA’s coverage fell further than the strong one’s in 36 of 40 cases, and what it still accepted was less reliable: on Qwen3-1.7B at MLX 3-bit, its error among accepted items rose from 5.4% to 23.2%. A system that sends low-confidence outputs to a person would see that person’s workload change at formats where accuracy still looks fine. The one pre-registered test of this, at 4 bits, was not supported, because Q4_K_M moved coverage by only 1.2 points.
How the study was run
Every round’s confirmatory runs followed a written plan: the configurations, hypotheses, statistical tests and verdict rules, and from round 2 on, what each key outcome would change. Each plan was hashed and timestamped before its confirmatory runs, and the plans for rounds 3 to 5 were also pushed to GitHub before any of their fine-tunes had been evaluated. Results that already existed when a plan was written were labeled exploratory, which in round 1 meant all of Qwen3-0.6B.
Every pre-registered result is reported, including the round-1 hypotheses that were not supported and the bar the first predictor missed, and the paper lists six deviations from the plans. Every fine-tune of rounds 1 to 4 was evaluated on its complete test set in both engines.3
Bottom line
- Pick the fine-tune on the file you ship. The gentle LoRA was more accurate at Q4_K_M in all eight settings, the strong one at MLX 3-bit and Q2_K in all eight (exploratory).
- The LoRA at mlx-lm’s default learning rate kept less at MLX 3-bit and Q3_K_M. At MLX 3-bit and Q3_K_M, the LoRA at 1e-5 kept less of its gain than the one at 1e-4 in all 10 pre-registered tests.
- Unfusing does not save a gentle LoRA in MLX. Training on the quantized base does. It delivered 96% of the bf16 gain at 4 bits instead of 84%, but the same move cost the strong LoRA at 4 bits.
- Read GGUF names as labels, not sizes. Q3_K_M measured 4.2 to 4.4 bits per weight, closer to MLX 4-bit than to MLX 3-bit.
- Check coverage, not only accuracy. At a confidence threshold fixed at bf16, quantization at low bit widths cut the share of accepted items much faster than it cut accuracy (exploratory).
- A post hoc model of update size plus base damage predicted retention at least as well. The weight measurements describe the loss; they have not been shown to predict it better than the size of the update.
Sources
- Code, pre-registrations, per-item predictions, the paper and the
ftquant checktool: JayeshSuryavanshi/ftquant. The three charts are drawn by its plotting code from its committed result files. - Closest prior results: Li, Huang, Wang, Gotei, Scale-QLoRA: Code-Invariant Adapter Merging for Native 4-bit Microscaling LLMs, arXiv 2609.04526 (2026), on merging a LoRA into a native 4-bit checkpoint; Zhang et al., Catastrophic Failure of LLM Unlearning via Quantization, ICLR 2025, where quantization restores what small-learning-rate unlearning removed.
- Training on the quantized base: Dettmers et al., QLoRA, NeurIPS 2023. Quantization-aware initialization and training: Li et al., LoftQ, arXiv 2310.08659; Xu et al., QA-LoRA, arXiv 2309.14717.
- Data: banking77 (Casanueva et al., 2020), from PolyAI’s original files; MASSIVE en-US (FitzGerald et al., ACL 2023). Models: Qwen3 (Yang et al., 2025) and OLMo-2-0425-1B-Instruct (Team OLMo, 2025).
- Companion piece: The wrong opponent, which measured a small encoder on banking77.