Skip to content

Choose a model

Every model answers the same typed questions with the same API; they differ in speed, memory and what they are best at. All are Apache-2.0 and on Hugging Face.

Which one?

If you need… Use
many decisions per second, on CPU or any GPU opendecider-nano
the best accuracy on decisions unlike the training data, on a 16 GB Mac or one GPU opendecider-small
business workflows (triage, invoices, security alerts, agent traces) on a 16 GB Mac or one GPU opendecider-small-td
the most accurate model you can host yourself (NVIDIA, several GPUs) opendecider-medium-td
probabilities you can threshold on: the best calibration and agreement with people (NVIDIA, several GPUs) opendecider-large-td
a Mac with MLX opendecider-small-mlx-8bit (or -4bit with little memory)
LM Studio or Ollama opendecider-small-GGUF / opendecider-small-td-GGUF

All checkpoints

model backbone params memory use it for
opendecider-nano Ettin-encoder-400m ~400M 2.0 GiB speed: 17–18 ms per question, ~9 ms batched; typed business decisions
opendecider-small Qwen3-4B-Instruct-2507 + LoRA 4B 8.9 GiB, tested on a 16 GB Mac mini accuracy and calibration on decisions it has never seen
opendecider-small-td Qwen3-4B-Instruct-2507 + LoRA 4B 8.9 GiB business workflows like typed-decisions' (triage, invoices, security alerts, agent traces): 0.792
opendecider-medium-td Qwen3-30B-A3B-Instruct-2507 + LoRA 30B (3B active) 61 GB bf16, across several GPUs (tested on 4× L40S) the most accurate self-hostable model on general decisions (0.765); NVIDIA only
opendecider-large-td Qwen3-Next-80B-A3B-Instruct + LoRA 80B (3B active) 160 GB bf16, across several GPUs (tested on 4× L40S) the best calibration and agreement with people; NVIDIA only
opendecider-small-mlx-8bit opendecider-small, MLX 8-bit 4B 4.5 GB Macs: same answers as full precision on 1,955 of 2,000 typed-decisions questions, 66 ms per question
opendecider-small-mlx-4bit opendecider-small, MLX 4-bit 4B 2.6 GB Macs with little memory; about 2 points lower on typed-decisions (0.651)
opendecider-small-GGUF opendecider-small, GGUF Q8_0 / Q4_K_M 4B 4.3 / 2.5 GB LM Studio and Ollama: typed-decisions 0.669 at Q8_0 (full precision 0.671)
opendecider-small-td-GGUF opendecider-small-td, GGUF Q8_0 / Q4_K_M 4B 4.3 / 2.5 GB LM Studio and Ollama, business workflows: 0.794 at Q8_0 (full precision 0.792)

Load any of them by name: load("manjunathshiva/opendecider-small-td"). The Qwen-based models need pip install "opendecider[small]", the MLX builds pip install "opendecider[mlx]", and the GGUF builds run in an app (see LM Studio, Ollama and vLLM).

Hardware

hardware nano small / small-td medium-td large-td
CPU only ✅ 0.1–0.7 s per question slow (~17 GB RAM in fp32) – –
16 GB Mac (M4) ✅ 28 ms ✅ 280 ms (8.9 GiB of the 11.8 GiB GPU budget) – –
64 GB Mac (M4 Max) ✅ 18 ms ✅ 141 ms; MLX 8-bit 66 ms – –
one NVIDIA GPU ✅ 16 ms on an L40S ✅ 38 ms on an L40S (12 GB or more) – –
several NVIDIA GPUs ✅ 214 ms on 4× L40S (~64 GB in total) ✅ 440 ms on 4× L40S (~170 GB in total)

Latency is one question, median. Answers are identical to four decimals across the tested Macs and NVIDIA machines. Models larger than one GPU are spread across all visible GPUs automatically. large-td needs transformers 4.57 or newer; pip install flash-linear-attention speeds it up.

How they work

  • opendecider-nano: Ettin-encoder-400m (bidirectional, fully fine-tuned) reads question: …, [MASK] option 1, [MASK] option 2, …, input: <state>. The hidden state at each [MASK] goes through a small MLP to one logit, then a softmax across that question's options. The answer space is defined at request time, so new schemas need no retraining, and a 78-option question still costs one forward pass.
  • opendecider-small and small-td: Qwen3-4B-Instruct-2507 with a LoRA adapter (r = 16, all linear projections). The options are lettered, and one forward pass gives the probability of each letter as the next token. Above 26 options it scores each option name's log-probability after the shared prompt.
  • opendecider-medium-td and large-td: the same design on Qwen3-30B-A3B-Instruct-2507 and Qwen3-Next-80B-A3B-Instruct (mixtures of experts, 3B active). The adapter covers the attention projections; the experts are frozen.

Training. Two openly licensed teachers, Qwen3-235B-A22B-Instruct-2507 (Apache-2.0) and DeepSeek V4.1 Flash (MIT), scored every training question through token log-probabilities. Each teacher was temperature-scaled on held-out gold labels before the two were averaged, so the students learn calibrated distributions, not hard labels. nano, small-td, medium-td and large-td then had a short fine-tune on the typed-decisions train split. No benchmark dataset, or its family, is in the training data, and no outputs of Claude or GPT models were used. Training-data licences are in NOTICE.