Training decision models on a Mac
Results
Typed Decisions, full 400-case test split: 2,000 decisions, five questions per case, four workflows. Scored with the benchmark's own metrics, averaged over the 20 workflow/question buckets. Latency is per case on an M3 Max.
| Model | Accuracy | KL from gold ↓ | Brier ↓ | ECE ↓ | p50 per case |
|---|---|---|---|---|---|
| gemma-4-e2b-it-OptiQ-4bit-decision | 0.802 | 0.072 | 0.039 | 0.181 | 0.99 s |
| Qwen3.5-0.8B-decision | 0.772 | 0.097 | 0.054 | 0.177 | 0.34 s |
| clef-OptiQ-4bit (27B, trained by Cloudflare on other data) | 0.719 | 0.189 | 0.102 | 0.118 | 9.7 s |
| clef-flash-OptiQ-4bit (9B) | 0.704 | 0.201 | 0.106 | 0.118 | 2.7 s |
| Always answer the train-split average | 0.489 | 0.327 | 0.181 | – | – |
The two new models were trained on the benchmark's own train split and nothing else: 1,080 records. One seed each. The test cases are disjoint from the training records but come from the same four workflows.
What a decision model does
A decision model answers a question by returning a probability for each option instead of generating text. Typed Decisions has three question types: yes/no, named choices, and ordered score levels. Cloudflare's Clef and Jev are decision models. OptiQ already ships Clef as MLX quants and serves decision models at /v1/systemone through optiq serve.
optiq lora train --decision turns an ordinary language model into one: a LoRA on the backbone plus a new schema head that scores all the questions of a case in one forward pass. Training has three phases. The head warms up for 100 steps with the backbone frozen, then the LoRA and head train together, and finally one temperature per question type is fitted on held-out records.
The two runs
Gemma-4 E2B. Base: mlx-community/gemma-4-e2b-it-OptiQ-4bit. LoRA rank 16 over 35 blocks (24.2M parameters), head 28.9M parameters, 3 epochs, 405 steps with gradient accumulation 8. About 5 hours on an M3 Max, peak memory 13.9 GB.
Qwen3.5-0.8B. bf16 base, LoRA over 24 blocks (10.2M parameters), head 27.3M parameters, peak memory about 2.3 GB.
Both are on Hugging Face under mlx-community.
Compared with Unsloth
Unsloth's docs describe training a decision model and report typed-decisions numbers: Qwen3.5-0.8B from 36% to 73%, and Qwen3.5-2B from 33% to 78%. Their run used LoRA rank 64 on a GPU and mixed typed-decisions with 12 other public datasets (ag_news, arc, banking77, boolq, clinc150, commonsense_qa, mmlu, mnli, prompt_injections, snli, sst5, wanli).
The same 0.8B model reaches 77.1% here, and Gemma-4 E2B reaches 80.0%. Both were trained on one dataset's 1,080 records with a LoRA on a Mac. Unsloth does not publish KL, Brier or ECE, so accuracy is the only number to compare. Other models on the benchmark's leaderboard score higher than either of ours.
The bug we shipped
The first scores were 0.512 for Qwen and 0.511 for Gemma, barely above always predicting the train-split average (0.489). Training itself had reported 0.81 to 0.83 validation accuracy. That gap was the clue.
The trainer saved the LoRA tensors without the backbone's language_model. name prefix. Qwen3.5 and Gemma-4 both load as multimodal model classes, and mlx-lm skips adapter tensors whose names it cannot place. There was no error and no warning. Every loaded adapter ran with untrained LoRA weights.
We had already published the Qwen model with the 0.512 number. mlx-optiq 0.5.21 fixes it: adapters are saved under their full names, older adapters are bound correctly on load, and an adapter tensor that matches nothing is now an error. The Hugging Face repo has the fixed adapter and a dated correction. A new test loads a saved adapter into a fresh model and checks that the weights arrive.
The silent kills
The Gemma run died five times with no traceback. The Mac had 1.9 GB of disk free, so macOS could not grow swap, and its memory-pressure killer took the largest process. Freeing disk fixed it. Two trainer fixes came out of the hunt anyway. The attention router was forking sysctl on every attention call, about 250,000 processes in one run, and now caches the answer. The decision and DPO trainers now clear MLX's buffer cache every step.
Try it
optiq serve --model mlx-community/gemma-4-e2b-it-OptiQ-4bit-decision
optiq lora train <model> --decision --data <dir with train.jsonl and valid.jsonl> --output <adapter dir>
The decision models docs cover serving, the request format and training options.