Serve it: a drop-in for TypeSafe Jev's API¶
opendecider serve puts any OpenDecider model behind an HTTP server that speaks TypeSafe Jev's /v1/systemone
protocol, so existing Jev clients work by changing the base URL. It was tested with TypeSafe's own SDK
(@typesafe-ai/sdk 0.6.0) and no code change.
Start it¶
The images serve opendecider-nano by default; set OPENDECIDER_MODEL to serve another model, and mount a volume at
/models so downloaded weights survive restarts. Pin a release with a version tag, for example
ghcr.io/manjunathshiva/opendecider:0.2.1.
Call it¶
curl -s localhost:8000/v1/systemone -H 'content-type: application/json' -d '{
"state": "Hi, we were billed twice for March. Please refund the duplicate today.",
"questions": {"department": {"type": "choice", "instructions": "Which department?",
"criteria": {"billing": "payments, refunds", "technical": "bugs, outages"}}}}'
With TypeSafe's SDK, point it at the server:
A Python client with retries (standard library only) is in examples/serve_client.py.
| endpoint | what it does |
|---|---|
POST /v1/systemone |
one state, any number of typed questions (Jev's request and response shape) |
POST /v1/systemone/batch |
{"states": [...], "questions": {...}}: the same questions about many states in one call |
GET /v1/models |
the served model (as Jev's models list) |
GET /health, GET /ready |
liveness and readiness probes (model, device, version, queue depth) |
GET /metrics |
Prometheus text: requests by status, latency and batch-size histograms, in-flight, queue depth |
Request and response fields, limits and status codes: HTTP API. Every flag and environment variable: Command line.
Production behaviour¶
- Batching across requests: one inference thread merges the questions of all waiting requests into shared batches, so nano's batched speed holds under concurrent load.
- Bounded load: above
--max-in-flightrequests the server answers 503 withRetry-Afterat once instead of queueing without limit, and a request not answered within--request-timeout-sgets 504. - Validation: request size, questions per state, options per question and state length are all bounded, and invalid requests get a 4xx with a message naming the problem.
- Security: bearer-token auth with
OPENDECIDER_API_KEY(constant-time comparison); server errors never leak internals to the client. - Observability: an
x-request-idheader on every response, and Prometheus metrics at/metrics.
Performance under load¶
benchmarks/load_test.py: 100 concurrent users, 20 s ramp, 40 s hold, three questions per request, one server process
(raw results in benchmarks/results/load/):
| model and hardware | requests / s | p50 | p95 | p99 | errors |
|---|---|---|---|---|---|
nano, CPU only (8 cores, AWS c7i.4xlarge), --dtype bfloat16 |
24 | 4.4 s | 4.9 s | 5.2 s | 0 |
| nano, CPU only (same machine), default fp32 | 9 | 16.6 s | 17.2 s | 19.0 s | 0 |
| nano, 1× NVIDIA L4 | 50 | 2.1 s | 2.3 s | 2.3 s | 0 |
small (4B), 1× NVIDIA L4, --small-batch 16 |
20 | 5.8 s | 6.5 s | 6.6 s | 0 |
| small (4B), 1× NVIDIA L4, default | 9 | 14.4 s | 14.7 s | 16.1 s | 0 |

With 100 users each waiting for their answer before sending the next, latency is roughly users ÷ throughput, so a single process is at its limit here. To serve more traffic, run more replicas behind a load balancer (one process per GPU; on CPU, one process per 8 or so physical cores).
Recommended settings¶
- CPU with bf16 units (Intel Sapphire Rapids and newer, i.e. AMX):
--dtype bfloat16, 2.8× the throughput. On typed-decisions it has the same accuracy as fp32 (0.796) and the same top answer on 1,992 of 2,000 questions. On CPUs without bf16 units keep the default. - Qwen-based models on a GPU (small, small-td, medium-td, large-td):
--small-batch 16, 2.2× the throughput. The top answer changes on about 1 question in 75 (bf16 arithmetic in padded batches), so leave it off where you need exactly the published answers. --threadsat most the number of physical cores; oversubscribing hyperthreads slows CPU inference.
Serve a model running in LM Studio, Ollama or vLLM¶
--model also takes a model served elsewhere; the server then answers /v1/systemone using that app as the engine:
opendecider serve --model lmstudio:opendecider-small
opendecider serve --model ollama:hf.co/manjunathshiva/opendecider-small-GGUF:Q8_0
OPENDECIDER_REMOTE_URL=http://localhost:8001/v1 opendecider serve --model openai:opendecider-small # e.g. vLLM
The last line assumes vLLM was started with --port 8001, since opendecider serve itself listens on 8000.