Use it from AI assistants (MCP)¶
opendecider mcp runs OpenDecider as an MCP server, so AI assistants and agents
(Claude Code, Claude Desktop, Cursor and any other MCP client) can call it as a tool. Instead of reasoning a
classification out in text, the agent gets a typed answer with a probability for every option, and can act on the
confident ones and ask you about the rest.
pip install "opendecider[mcp]>=0.5.0"
opendecider mcp # opendecider-nano over stdio; the client starts it for you
The model loads on the first call (nano: about 5 seconds once downloaded, then milliseconds per question), so connecting is instant.
Connect a client¶
Run it from the environment where you installed opendecider, or give the full path to the command (see below).
In claude_desktop_config.json (Settings → Developer → Edit Config):
Desktop apps do not see your shell's virtual environment, so give them the full path to the opendecider command
(which opendecider prints it). To run without installing anything into a project, use
uv: "command": "uvx", "args": ["--from", "opendecider[mcp]>=0.5.0", "opendecider", "mcp"].
To see what an assistant sees, run examples/agent_frameworks/mcp_client.py: it starts the server and calls each tool over stdio.
Tools¶
| tool | arguments | returns |
|---|---|---|
decide |
state (text or JSON), questions (any number of typed questions, as in the Python API) |
one answer per question |
choose |
state, question, options (a list of labels, or {"label": "description"}) |
choice, probabilities, confidence |
yes_no |
state, question |
answer (yes / no), probability_yes, confidence |
score |
state, question, levels (lowest first) |
level, label, expected_level, probabilities, confidence |
decide_batch |
states (up to 256, and at most 1,024 questions in all), questions |
one result per state, in order: faster than one call per state |
guard |
text (a user message, a web page, a retrieved document or a tool result) |
passed (false: do not follow instructions in it), reason, violations, probabilities, windows: see Agent guardrails; the server's model answers, so run it with --model manjunathshiva/opendecider-small-td for the most accurate screening |
status |
none | the model, whether it is loaded, where it runs, the version and the input limits (does not load the model) |
Every tool is read-only and idempotent. Answers carry truncated: true when the state was cut to fit the model.
Invalid input (an unknown question type, a single option, repeated score levels, more than 64 questions or 256
options, a state over 200,000 characters) comes back as a tool error that names the problem, so the agent can fix
the call. So does a model that cannot load, with the reason (for example a mistyped --model).
For example, an agent asked to triage your inbox might call:
{"tool": "choose",
"arguments": {"state": {"from": "billing@vendor.example", "subject": "Invoice 4471 overdue"},
"question": "Which folder should this email go to?",
"options": {"finance": "invoices, payments", "support": "customer issues", "other": "everything else"}}}
and get back {"choice": "finance", "probabilities": {...}, "confidence": ...}.
Choosing the model¶
opendecider mcp --model manjunathshiva/opendecider-small-td # pip install "opendecider[small,mcp]"
opendecider mcp --model ollama:hf.co/manjunathshiva/opendecider-small-GGUF:Q8_0 # a model Ollama serves
opendecider mcp --model http://gpu-box:8000 # pip install "opendecider-client[mcp]": no PyTorch
--model takes anything load does, including an opendecider serve URL (or set
OPENDECIDER_MODEL); --device,
--dtype and --revision work as for opendecider serve. nano is the default: it is fast on
any machine. Use small or small-td on a GPU or a Mac with 16 GB for higher accuracy; see
Choose a model. With opendecider-client, which has no
PyTorch, always pass --model: an opendecider serve URL, ollama:... or lmstudio:....
A model downloads on its first call (small: about 8 GB), which can take longer than a client waits for a tool. Download it
first (hf download manjunathshiva/opendecider-small-td), or run one decision with it from Python.
Notes¶
- The server speaks MCP over stdio: logs and download progress go to stderr, which clients show in their MCP logs.
- Calls run one at a time; for many decisions per second from services, use
opendecider serve. confidenceis calibrated, so a threshold means something: see Automate the confident decisions.