(Benchmarks)
Benchmarks
On decision tasks Hesperan 1 matches or beats Jev 1.13 on 8 of 12 public benchmarks. On knowledge-heavy tasks it is weaker — they are listed too.
| Benchmark | n | Hesperan 1 | Jev 1.13 | Δ vs Jev | Best other published |
|---|---|---|---|---|---|
| JevBench public | 231 | 88.7 | 86.6 | +2.1 | 87.0 Reflex-27B |
| JevBench public — hard | 111 | 77.5 | 73.0 | +4.5 | 75.7 Reflex-27B |
| jabr classifier v1 | 78 | 98.7 | 97.4 | +1.3 | 92.3 Von 1.0.1 |
| jabr classifier v2 | 866 | 96.2 | 96.7 | −0.5 | 68.9 GLiNER2 |
| jabr support tickets | 500 | 97.4 | 97.8 | −0.4 | 86.0 Von 1.0.1 |
| Nimble public (13 sets, macro)† | 3,880 | 76.1 | 76.0 | +0.1 | 77.4 Decider-35B |
| typed-decisions | 2,000 | 73.0 | 72.7 | +0.3 | 36.0 Laya |
| Kev transfer-v4 dev | 656 | 85.5 | 85.7 | −0.2 | 81.2 Kev-9B |
| Kev transfer-v9 dev | 1,046 | 79.9 | 85.4 | −5.5 | — |
| OOD support tickets (EN/KO) | 900 | 77.8 | 75.1 | +2.7 | — |
| Phishing (trifleen) | 100 | 83.0 | 82.0 | +1.0 | — |
| BoolQ validation† | 3,270 | 91.8 | 91.6 | +0.2 | — |
| Benchmark | n | Hesperan 1 | Jev 1.13 | Δ vs Jev | Best other published |
|---|---|---|---|---|---|
| OpenBookQA | 500 | 91.8 | 94.2 | −2.4 | — |
| CommonsenseQA† | 1,221 | 86.3 | 88.1 | −1.8 | — |
| HellaSwag | 2,000 | 84.2 | 86.1 | −1.9 | — |
| MMLU-Pro (1k sample) | 1,000 | 61.2 | 82.9 | −21.7 | — |
Hesperan 1 measured by us on 22 Sep 2026 on the public items of each benchmark, with the benchmark's own metric. All other numbers are the values published by the benchmark authors.
† The Hesperan 1 training mix contains the train split of this source (in-distribution).
No Jev outputs were used for training, distillation, labelling or model selection. Independent evaluation; not affiliated with or endorsed by TypeSafe AI. Jev and TypeSafe are trademarks of TypeSafe AI, Inc.