mlx-optiq
Guide · Decision models

Decision models on Apple Silicon

A decision model reads a piece of state and a set of typed questions, and answers all of them in one forward pass. For every option of every question it returns a probability. There is no generated text and nothing to parse. Cloudflare Clef is the open example, and the request and response match the Jev / SystemOne API, so existing clients work unchanged.

OptiQ serves these models at POST /v1/systemone and trains them with optiq lora train. Both are text-only for now.

Images and video A request that carries images or video is refused with a 400 and a message, not answered as if they were absent. Send the state as text.

The three question types

TypeYou sendYou get back
noulA yes/no instructionThe probability that the answer is yes
choicecriteria: option names with a description eachA probability per option, and the most likely one
scorecriteria: an ordered list of levelsA probability per level, plus the expected level

Serve one

Point optiq serve at a decision model, meaning a model directory or repo that carries a joint_head.safetensors next to its weights. The endpoint turns on by itself.

terminalbash
pip install mlx-optiq
optiq serve --model mlx-community/clef-OptiQ-4bit
requestbash
curl http://127.0.0.1:8080/v1/systemone -d '{
  "model": "clef",
  "state": "Checkout has been failing for every customer for the last hour.",
  "questions": {
    "urgent": {"type": "noul", "instructions": "Is this urgent?"},
    "team": {"type": "choice", "criteria": {"billing": "Payments", "technical": "Outages and errors"}
  }
}'
responsejson
{
  "model": "clef",
  "answers": {
    "urgent": {"type": "noul", "noul": 0.9878},
    "team": {"type": "choice", "choice": "billing", "confidence": 0.6225,
             "probabilities": {"billing": 0.6225, "technical": 0.3775}
  },
  "usage": {"input_tokens": 214, "output_tokens": 0, "latency_ms": ...}
}

Errors come back as JSON. A malformed body or a missing field is a 400 with the reason. A request that does not fit the context window is a 413. By default a long state is cut to fit; send "truncate": false to get the 413 instead. The default window is 16,384 tokens.

The chat endpoints still answer on the same model, but a decision model is not a chat model and what it says there is not meaningful.

Use it from Python

decide.pypython
from optiq.decision import DecisionModel

model = DecisionModel.load("mlx-community/clef-OptiQ-4bit")
print(model.predict({
    "state": "Checkout has been failing for every customer for the last hour.",
    "questions": {"urgent": {"type": "noul", "instructions": "Is this urgent?"},
}))   # {'urgent': {'true': 0.99, 'false': 0.01}

Train one

Training data is JSONL, one record per line: a state, the questions, and the answers.

train.jsonljson
{"state": "Card declined twice, now locked out.",
 "questions": {"team": {"type": "choice", "criteria": {"billing": "Payments", "technical": "Bugs"}},
 "answers": {"team": "billing"}

To train on a distribution instead of one answer, use "gold": {"team": {"probabilities": {"billing": 0.7, "technical": 0.3}}. Records with several annotators, or a vote that split, train the same way. A noul answer is true or false; a score answer is the level index.

terminalbash
optiq lora train mlx-community/clef-flash-OptiQ-4bit --data ./data -o ./adapter
optiq lora export mlx-community/clef-flash-OptiQ-4bit --adapter ./adapter -o ./my-decider
optiq serve --model ./my-decider

--data is a folder with train.jsonl and, optionally, valid.jsonl. Training puts LoRA adapters on the backbone and trains the schema head in full, using cross-entropy against your answers. It keeps about 5% of the records back to fit one temperature per question type afterwards, which corrects over- or under-confident probabilities. The export carries the trained head and the temperatures along.

Start from any Qwen model

A model with no head can become a decision model. Add --decision and OptiQ gives it a new, untrained head in Clef's design, sized to the backbone.

terminalbash
optiq lora train mlx-community/Qwen3.5-2B-OptiQ-4bit --data ./data -o ./adapter \
  --decision --head-warmup-steps 100

The head starts from random weights, so it needs more data and more steps than fine-tuning a model that already has one. --head-warmup-steps trains the head alone for that many steps first, with the backbone frozen, and then LoRA and head train together. --head-lr sets the head's learning rate; it defaults to 3e-4, higher than the LoRA's because the head starts from nothing.

Other model families The prompt layout the released heads were trained on is Qwen's chat format, and Qwen models get exactly that. Other families use their own chat template around the same content. We have tested Qwen3.5 only.

Scores

Both quants were scored on the full 400-case test split of Typed Decisions with the benchmark's own metrics, one request per case, on an Apple M3 Max.

ModelOn diskAccuracyKL from goldBrierECE
clef-OptiQ-4bit19 GB0.7190.1890.1020.118
clef-flash-OptiQ-4bit7.7 GB0.7040.2010.1060.118

Accuracy is higher-is-better; the other three are lower-is-better. The gold labels are an average of teacher samples, so a score measures agreement with that teacher. See the benchmark card for what each number means.