LM Studio, Ollama and vLLM¶
The 4B models (small and small-td) can run in an app instead of PyTorch. The app runs the model; the opendecider
package builds the prompt the model was trained on and reads the option probabilities from the server's token
log-probabilities. No PyTorch is needed on the client.
| engine | build | agreement with the PyTorch model on typed-decisions |
|---|---|---|
| LM Studio, Ollama (llama.cpp) | small-GGUF, small-td-GGUF, Q8_0 | same top answer on 1,975 and 1,972 of 2,000; accuracy 0.669 vs 0.671 and 0.794 vs 0.792 |
| LM Studio, Ollama | Q4_K_M | about 93%, for machines with little memory |
| vLLM (NVIDIA) | the base model with the LoRA adapter, no merge | same top answer on 1,963 and 1,969 of 2,000; 0.6735 vs 0.6715 and 0.7945 vs 0.792 |
LM Studio¶
lms get https://huggingface.co/manjunathshiva/opendecider-small-GGUF --select # choose Q8_0 (or search "opendecider" in the app)
lms load opendecider-small@q8_0 --identifier opendecider-small
lms server start
from opendecider import load
model = load("lmstudio:opendecider-small")
print(model.system_one("I was charged twice. Please refund the extra payment.",
{"team": {"type": "choice", "instructions": "Which team?",
"criteria": {"billing": "payments, refunds", "technical": "bugs"}}})["answers"])
Load the GGUF build in LM Studio on a Mac too: its MLX engine returns no log-probabilities. For MLX, use the
MLX builds with pip install "opendecider[mlx]" instead.
Ollama¶
Ollama's own /v1/systemone
Ollama 0.35 added its own decision endpoint (open to third-party models from 0.35.1). It builds a different
prompt from the one these builds were trained on, which costs them several points. Use them through the
opendecider package as above, or through opendecider serve (below). Builds trained on Ollama's prompt as well
are in preparation.
vLLM¶
Serve the base model with the LoRA adapter; no merge or GGUF needed.
hf download manjunathshiva/opendecider-small --local-dir opendecider-small
vllm serve Qwen/Qwen3-4B-Instruct-2507 --enable-lora --max-lora-rank 16 --max-logprobs 20 --max-model-len 4096 \
--lora-modules opendecider-small=./opendecider-small
Tested with vLLM 0.30 on an NVIDIA L4, with both adapters on one server (add opendecider-small-td=... to
--lora-modules).
Other servers¶
Any server with an OpenAI-compatible /v1/chat/completions that returns top_logprobs works the same way:
load("openai:<model>", base_url="http://host:port/v1").
Jev's API on top of any of them¶
opendecider serve --model lmstudio:opendecider-small
opendecider serve --model ollama:hf.co/manjunathshiva/opendecider-small-GGUF:Q8_0
For openai: models, set OPENDECIDER_REMOTE_URL to the server's URL. See Serve it.
Settings and limits¶
- Server URL:
base_url=..., elseOPENDECIDER_REMOTE_URL, else LM Studio's (http://127.0.0.1:1234/v1) or Ollama's (http://127.0.0.1:11434/v1) default address. - API key: set
OPENDECIDER_REMOTE_API_KEYif the server needs one. It is sent only to that server (never on a redirect), and over plain HTTP only to this machine; for another host use https, or setOPENDECIDER_REMOTE_ALLOW_HTTP=1if the network path is trusted. - Up to 26 options per question. OpenAI-compatible servers return at most 20 log-probabilities, so with 21 to 26 options the least likely letters get probabilities near zero.