Benchmarks¶
Every model answered the same questions and was scored by the same code. TypeSafe Jev was measured directly through TypeSafe's own API, not taken from published figures. Full tables, per-task results and methodology: COMPARISON.md.
Reproduce every number: python benchmarks/report.py rebuilds all the tables from the committed results, and
python benchmarks/run.py --model <name> re-scores any model
(benchmarks/).

The test sets¶
- typed-decisions (dataset): 2,000 decisions across four business workflows, scored with the Jev-vs-Laya harness published by Kameshwara Pavan kumar Mantha and the Antz AI team. nano, small-td, medium-td, large-td and Laya's typed-decisions checkpoint were fine-tuned on its train split; the test split was never used for training or model selection. Zero-shot models are a reference there, not a head-to-head.
- 200 general decisions none of the models trained on: BANKING77 intent (78 options), BoolQ yes/no, Yelp 1–5 star rating and ChaosNLI (100 human labels per item).
- Laya's application battery: Laya's own 10-task benchmark, at a pinned commit.
OpenDecider vs TypeSafe Jev¶
| benchmark / metric | TypeSafe Jev 1.13 | nano | small | medium-td | large-td |
|---|---|---|---|---|---|
| typed-decisions, 2,000 decisions | 0.754 (zero-shot) | 0.796 | 0.672 (zero-shot) | 0.788 | 0.801 |
| 200 general decisions | 0.730 | 0.680 | 0.735 | 0.765 | 0.750 |
| Laya's application battery, 10 tasks | 0.774 | 0.656 | 0.702 | 0.725 | 0.718 |
| calibration error (ECE), general decisions ↓ | 0.164 | 0.092 | 0.087 | 0.110 | 0.083 |
| distance from the human label spread (ChaosNLI JSD) ↓ | 0.148 | 0.045 | 0.040 | 0.035 | 0.030 |
| median latency, 1 question | 404 ms (API) | 17 ms (L40S) | 40 ms (L40S) | 214 ms (4× L40S) | 440 ms (4× L40S) |
| weights | closed API | Apache-2.0 | Apache-2.0 | Apache-2.0 | Apache-2.0 |
Where Jev leads: Laya's application battery (0.774, and 0.803 on the five tasks Laya was not trained on), phishing (0.897 vs our 0.63–0.70), jailbreak detection (0.940), spam (0.985), model routing (0.975), BoolQ-style yes/no reading (0.94) and typed-decisions without fine-tuning (0.754 vs small's 0.672).
OpenDecider vs Laya¶
| benchmark | Laya | Laya typed-decisions | nano | small | medium-td | large-td |
|---|---|---|---|---|---|---|
| typed-decisions | 0.362 | 0.766 | 0.796 | 0.672 | 0.788 | 0.801 |
| 200 general decisions | 0.545 | 0.570 | 0.680 | 0.735 | 0.765 | 0.750 |
| Laya's battery, all 10 tasks | 0.695 | 0.702 | 0.656 | 0.702 | 0.725 | 0.718 |
| Laya's battery, the 5 tasks Laya was not trained on | 0.579 | 0.609 | 0.656 | 0.743 | 0.768 | 0.757 |
| calibration error (ECE), general decisions ↓ | 0.327 | 0.162 | 0.092 | 0.087 | 0.110 | 0.083 |
Like for like on typed-decisions (both fine-tuned on the train split), nano leads Laya's typed-decisions checkpoint by +0.030 (95% CI +0.014 to +0.044). Where Laya leads: the five datasets it was trained on (AG News 0.95, Enron spam 0.99, phishing 0.98), multilingual use (a 100+ language checkpoint; OpenDecider is evaluated in English only), and single-question speed on short inputs, where it is in the same range as nano.
Frontier LLMs (same 200 general decisions)¶
| model | accuracy | ECE ↓ | median latency | $ / 1,000 decisions |
|---|---|---|---|---|
| Claude Fable 5.1 | 0.840 | 0.064 | 4.27 s | $11.81 |
| GPT-6 Astra | 0.790 | 0.119 | 2.22 s | $6.96 |
| opendecider-medium-td | 0.765 | 0.110 | 214 ms | self-hosted |
| DeepSeek V4.1 Flash | 0.760 | 0.138 | 4.08 s | $0.158 |
| MiniMax M3 | 0.755 | 0.112 | 1.02 s | $0.149 |
| opendecider-large-td | 0.750 | 0.083 | 440 ms | self-hosted |
| Kimi K3 | 0.745 | 0.119 | 6.28 s | $3.64 |
| opendecider-small | 0.735 | 0.087 | 40 ms | self-hosted |
| TypeSafe Jev 1.13 | 0.730 | 0.164 | 404 ms | $0.025 |
| opendecider-nano | 0.680 | 0.092 | 17 ms | self-hosted |
| CLM-8B (Contrastive-LM) | 0.400 | 0.106 | ~35 ms | self-hosted |
Only Claude Fable 5.1 and GPT-6 Astra beat opendecider-medium-td here, at 10–20× its latency and with a per-call bill. The 200-item set is about ±3 points, so medium-td, DeepSeek V4.1 Flash and MiniMax M3 are close.
Speed¶
| questions per call | nano, NVIDIA L40S | nano, Apple M4 Max | small, NVIDIA L40S | small, Apple M4 Max |
|---|---|---|---|---|
| 1 | 16.1 ms | 18.1 ms | 37.6 ms | 141 ms |
| 5 | 24.4 ms (4.9 ms/q) | 54.3 ms (10.9 ms/q) | 190.1 ms (38.0 ms/q) | 680 ms (136 ms/q) |
| 10 | 42.9 ms (4.3 ms/q) | 98.1 ms (9.8 ms/q) | 388.2 ms (38.8 ms/q) | 1.37 s (137 ms/q) |
| 50 | 189.5 ms (3.8 ms/q) | 467 ms (9.3 ms/q) | 1.94 s (38.7 ms/q) | 6.86 s (137 ms/q) |
Throughput under concurrent load is in Serve it.