Decision models on Apple Silicon
A decision model reads a piece of state and a set of typed questions, and answers all of them in one forward pass. For every option of every question it returns a probability. There is no generated text and nothing to parse. Cloudflare Clef is the open example, and the request and response match the Jev / SystemOne API, so existing clients work unchanged.
OptiQ serves these models at POST /v1/systemone and trains them with optiq lora train. Both are text-only for now.
The three question types
| Type | You send | You get back |
|---|---|---|
noul | A yes/no instruction | The probability that the answer is yes |
choice | criteria: option names with a description each | A probability per option, and the most likely one |
score | criteria: an ordered list of levels | A probability per level, plus the expected level |
Serve one
Point optiq serve at a decision model, meaning a model directory or repo that carries a joint_head.safetensors next to its weights. The endpoint turns on by itself.
pip install mlx-optiq optiq serve --model mlx-community/clef-OptiQ-4bit
curl http://127.0.0.1:8080/v1/systemone -d '{
"model": "clef",
"state": "Checkout has been failing for every customer for the last hour.",
"questions": {
"urgent": {"type": "noul", "instructions": "Is this urgent?"},
"team": {"type": "choice", "criteria": {"billing": "Payments", "technical": "Outages and errors"}
}
}'
{
"model": "clef",
"answers": {
"urgent": {"type": "noul", "noul": 0.9878},
"team": {"type": "choice", "choice": "billing", "confidence": 0.6225,
"probabilities": {"billing": 0.6225, "technical": 0.3775}
},
"usage": {"input_tokens": 214, "output_tokens": 0, "latency_ms": ...}
}
Errors come back as JSON. A malformed body or a missing field is a 400 with the reason. A request that does not fit the context window is a 413. By default a long state is cut to fit; send "truncate": false to get the 413 instead. The default window is 16,384 tokens.
The chat endpoints still answer on the same model, but a decision model is not a chat model and what it says there is not meaningful.
Use it from Python
from optiq.decision import DecisionModel
model = DecisionModel.load("mlx-community/clef-OptiQ-4bit")
print(model.predict({
"state": "Checkout has been failing for every customer for the last hour.",
"questions": {"urgent": {"type": "noul", "instructions": "Is this urgent?"},
})) # {'urgent': {'true': 0.99, 'false': 0.01}
Train one
Training data is JSONL, one record per line: a state, the questions, and the answers.
{"state": "Card declined twice, now locked out.",
"questions": {"team": {"type": "choice", "criteria": {"billing": "Payments", "technical": "Bugs"}},
"answers": {"team": "billing"}
To train on a distribution instead of one answer, use "gold": {"team": {"probabilities": {"billing": 0.7, "technical": 0.3}}. Records with several annotators, or a vote that split, train the same way. A noul answer is true or false; a score answer is the level index.
optiq lora train mlx-community/clef-flash-OptiQ-4bit --data ./data -o ./adapter optiq lora export mlx-community/clef-flash-OptiQ-4bit --adapter ./adapter -o ./my-decider optiq serve --model ./my-decider
--data is a folder with train.jsonl and, optionally, valid.jsonl. Training puts LoRA adapters on the backbone and trains the schema head in full, using cross-entropy against your answers. It keeps about 5% of the records back to fit one temperature per question type afterwards, which corrects over- or under-confident probabilities. The export carries the trained head and the temperatures along.
Start from any Qwen model
A model with no head can become a decision model. Add --decision and OptiQ gives it a new, untrained head in Clef's design, sized to the backbone.
optiq lora train mlx-community/Qwen3.5-2B-OptiQ-4bit --data ./data -o ./adapter \ --decision --head-warmup-steps 100
The head starts from random weights, so it needs more data and more steps than fine-tuning a model that already has one. --head-warmup-steps trains the head alone for that many steps first, with the backbone frozen, and then LoRA and head train together. --head-lr sets the head's learning rate; it defaults to 3e-4, higher than the LoRA's because the head starts from nothing.
Scores
Both quants were scored on the full 400-case test split of Typed Decisions with the benchmark's own metrics, one request per case, on an Apple M3 Max.
| Model | On disk | Accuracy | KL from gold | Brier | ECE |
|---|---|---|---|---|---|
| clef-OptiQ-4bit | 19 GB | 0.719 | 0.189 | 0.102 | 0.118 |
| clef-flash-OptiQ-4bit | 7.7 GB | 0.704 | 0.201 | 0.106 | 0.118 |
Accuracy is higher-is-better; the other three are lower-is-better. The gold labels are an average of teacher samples, so a score measures agreement with that teacher. See the benchmark card for what each number means.