Fig. 3 — Measured runs
Benchmarks & Calibration
All runs on this page are self-reported measurements on the public half of JevBench v1.4 — 231 items: easy 48 / standard 72 / hard 111 — using the benchmark's own official scoring code, through the live server, end to end.
§ 1The shipped lines, side by side
| v1-meticulous | v1-intuit | v1.1c (multimodal) | |
|---|---|---|---|
| JevBench-style score (public half) | 65.71 | 63.18 | 63.72 |
| Intelligence | 60.3 | 61.1 | 68.33 |
| top-label ECE → Calibration | 0.0496 → 90.1 | 0.1196 → 76.1 | 0.0984 → 80.3 |
| public accuracy (easy / standard / hard) | 0.710 (1.000 / 0.889 / 0.469) | 0.714 (1.000 / 0.903 / 0.468) | 0.766 (1.000 / 0.931 / 0.559) |
| skills-slice accuracy (609 held-out items) | 0.659 | 0.814 | — * |
| skills-slice ECE | 0.063 | 0.043 | — * |
| p50 / p95 latency | 0.17 s / 1.2 s | same | 0.26 s / 10.6 s |
| tokens/decision | 597 | same | 562 |
Robust claims: -intuit ≫ -meticulous on the skills slice (0.814 vs 0.659 on 609 items, ~20σ); -meticulous ≫ -intuit on bench ECE (0.05 vs 0.12, beyond single-run resolution); v1.1c's composite is not separable from v1 (63.72 vs 65.71, inside the ±5.9pt CI) — what is separable: hard tier +9.0pt (0.469 → 0.559), accuracy +5.6pt, matching Intelligence (+8.0), and the image/Chinese capability. A measured pipeline noise floor (~±2–3 composite points across runs differing only by seed and serving constants) bounds everything finer-grained: improvements below it are not evidence.
The benchmark updated to v1.4.1/v1.4.2 during this work (roster now 93 systems, 89 ranked; at run time the newest clone is bb05a33 — harness code and the 231 public items unchanged). Verified: the 231 public items, the 308-item sealed set, and the composite formula are unchanged — the official arithmetic was reproduced exactly — so the runs of record remain valid as scored. The frozen v1.5 method (904 open / 720 sealed) is not self-runnable from this repo and goes through Benchmark Heaven's submission pipeline.
kapteeni-v1.1c — the pre-registered gates, all passed in its one training run
| Gate | Reading | Bar |
|---|---|---|
| English image decisions (synth3-val, n=590) | 0.9661 | > 0.8068 (the frozen backbone's own score) |
| Chinese image decisions (synth3zh-val, n=204) | 0.9412 | reported, no bar |
| English text — never trained on (MNLI-noul, n=150) | 0.8800 | ≥ 0.84 |
| Chinese text — never trained on (OCNLI-noul, n=150) | 0.8467 | ≥ 0.83 |
| Chinese rule skills (synth2zh-val, n=403) | 0.9132 | ≥ 0.90 |
| English rule skills (synth2-EN-val, n=609) | 0.9048 | ≥ 0.90 |
| calibration after the fit (noul / choice / score) | 0.024 / 0.018 / 0.069 | ≤ 0.10 each |
§ 2The four axes (official formulas)
| Axis | meticulous | intuit | v1.1c | How it is measured |
|---|---|---|---|---|
| Intelligence | 60.3 | 61.1 | 68.33 | chance-corrected tier accuracy, weights easy .14 / standard .28 / hard .30 (renormalized — the judge tier is sealed); from public accuracy 0.710 / 0.714 / 0.766, 95% CI ±5.9pt (n=231) |
| Calibration | 90.1 | 76.1 | 80.3 | from top-label ECE 0.0496 / 0.1196 / 0.0984 across all 231 answers (ECE half only; the benchmark also averages a private gold-distribution half) |
| Speed | 81.1 | 81.0 | 69.6 | server latency p50 0.17 s, p95 1.2 s on an AMD Strix Halo iGPU, ×2 self-hosted adjustment (adjusted p50 0.34 s); v1.1c p50 0.26 s / p95 10.6 s — hard items run deep multi-pass chains, and image questions re-attach the image on every pass |
| Cost | 42.4 | 42.3 | — | 597 input tokens/decision × hosted list price — verified 2026-09-26: $0.03/M (Novita, qwen3-4b-fp8) = $0.018/1k decisions; nearest official Alibaba tier (qwen-turbo, $0.05/M) = $0.030/1k; the earlier $0.14/M assumption is kept as the pessimistic bound. v1.1c: 562 tokens/decision — its Cost axis is not computed in the record |
§ 3Per-tier accuracy (231 items)
| tier | n | meticulous | intuit | v1.1c |
|---|---|---|---|---|
| easy | 48 | 1.000 | 1.000 | 1.000 |
| standard | 72 | 0.889 | 0.903 | 0.931 |
| hard | 111 | 0.469 | 0.468 | 0.559 |
| pooled | 231 | 0.710 | 0.714 | 0.766 |
§ 4Per-family accuracy (public items — v1-meticulous)
| family | n | acc | family | n | acc |
|---|---|---|---|---|---|
| fact | 12 | 1.000 | routing | 12 | 0.750 |
| tool_selection | 12 | 1.000 | ambiguous | 7 | 0.429 |
| ordinal | 12 | 1.000 | probability | 10 | 0.400 |
| extraction | 24 | 0.917 | long_policy | 19 | 0.316 |
| intent | 24 | 0.958 | multi_hop | 18 | 0.278 |
| policy | 12 | 0.917 | temporal_numeric | 15 | 0.200 |
| adequacy | 12 | 0.917 | tradeoff | 6 | 0.167 |
| adversarial | 6 | 1.000 | judge_hard | 17 | 0.706 |
| trap | 8 | 0.875 | routing_hard | 5 | 1.000 |
The -intuit variant was developed for exactly the weak families above; on its 609-item held-out skills slice it reaches temporal_numeric 0.811 / multi_hop 0.917 / long_policy 0.644. It was not run family-by-family on the public items; only its tier accuracies above are public-half measured. Ordinal expectation MAE 0.482 levels; paraphrase consistency 0.861 (meticulous).
v1.1c — families that moved (vs v1-meticulous; n unchanged)
| family | v1 | v1.1c | Δ |
|---|---|---|---|
| multi_hop (n=18) | 0.278 | 0.667 | +38.9pt |
| long_policy (n=19) | 0.316 | 0.421 | +10.5pt |
| extraction (n=24) | 0.917 | 1.000 | +8.3pt |
| trap (n=8) | 0.875 | 1.000 | +12.5pt |
| judge_hard (n=17) | 0.706 | 0.765 | +5.9pt |
| tradeoff (n=6) | 0.167 | 0.333 | +16.7pt |
| adversarial (n=6) | 1.000 | 0.833 | −16.7pt |
| routing (n=12) | 0.750 | 0.667 | −8.3pt |
| temporal_numeric (n=15) | 0.200 | 0.133 | −6.7pt |
§ 5Calibration — the shipped models
Top-label expected calibration error across all 231 public answers, scored by the benchmark's own code (ECE half only; the benchmark also averages a private gold-distribution half):
| Model | top-label ECE | → Calibration | Serving constants |
|---|---|---|---|
| kapteeni-v1-meticulous | 0.0496 | 90.1 | head temperatures + blend weights fitted on the mixed-domain val set, through the merged model |
| kapteeni-v1-intuit | 0.1196 | 76.1 | fitted on the deployment-diverse held-out set, per pre-registration |
| kapteeni-v1.1c | 0.0984 | 80.3 | per-primitive temperatures on combined held-out val — the only constants fitted anywhere in the line; fitted ECE noul 0.024 / choice 0.018 / score 0.069, each within the ≤ 0.10 bar (PREREG-KAPTEENI-V11C.md) |
§ 6Run history
-
01
kapteeni-v1-meticulous, the flagship (2026-09-25)
LoRA r=32 on all projections + retrained heads over the 8.7M-token mix; MNLI held out of training as the OOD gate (0.88 → 0.893). All serving constants refit on mixed-domain val through the merged model — zero public-half selection in the chain.
-
02
kapteeni-v1-intuit (2026-09-26)
The second text variant — same architecture and wire format; the weak-family skills added to training and serving constants fitted on the deployment-diverse held-out set, per pre-registration. Sharper where the flagship is conservative, with calibration traded accordingly; the guidance of when to prefer it lives in the table above.
-
03
kapteeni-v1.1c, the multimodal line (2026-10-02)
The same phase pipeline ported to Qwen3.5-4B (images + Chinese); every pre-registered gate passed on the first reading. Run of record on the released pack: ECE 0.0984; not separable from v1 on the composite (inside the ±5.9pt CI). Published pinned at
17affb8; bench requests open for JevBench and Image JevBench (fstandhartinger/jevbench#178).
The records behind these three shipped models — every pre-registration, gate, and run (successful or not) — are kept in full in JEVBENCH.md and WORKLOG.md; a design proposal for a deployment-oriented benchmark alternative lives in DECISIONBENCH-DRAFT.md.
§ 7How to read every number above
- Sample sizes are small. All accuracies are on 231 items; one item is 0.43 accuracy points. The 95% binomial CI on public accuracy is ±5.9pt (0.710 / 0.714 / 0.766). Tier-level CIs are wider (hard: ±9.3pt at 0.469).
- Small composite gaps are ties. Differences of a few points between any two runs sit inside run-to-run variance plus axis assumptions (the Cost axis alone swings ±7 points across the stated price range).
- A measured pipeline noise floor. Runs differing only by training seed and val-fit constants scored ~2.6 composite points apart — treat ~±2–3 composite points as the single-seed noise floor.
- ECE at n=231 is biased and high-variance. Differences below ~0.02 between single runs are unresolvable; 0.05-vs-0.07 comparisons between runs should not be read as meaningful.
- Calibration is measured on the benchmark's distribution, which is OOD relative to any deployment: a model whose constants are tuned for the bench mix is not thereby calibrated for yours. The record documents this failure mode in the lab, not just in theory (JEVBENCH).
- Intelligence is renormalized over the three public tiers because the judge tier is sealed; any score that folds the sealed+judge weight in is not comparable to ours. Jev's accuracy on the same public items is still higher (0.866 vs 0.710).
- Composite gaps sit inside the CI. 63.72 (v1.1c) vs 65.71 (v1) is not separable at n=231 — the honest claim is "not separable on the composite with this sample size," not a win or a loss. The separable claims are recorded exactly: the hard-tier jump (+9.0pt), the accuracy gain (+5.6pt), and the image/Chinese capability.
§ 8Reproducibility
# server with the trained bundle
python3 -m kapteeni.serve --bundle model_cache/kapteeni_v1.pt \
--lora model_cache/kapteeni_p2/adapter --served-as kapteeni-v1-meticulous --port 8000
# multimodal line through the released pack
python3 -m kapteeni.serve_v11c --hf TriusAI/kapteeni-v1.1c --port 8002
# bench (from jevbench/)
TYPESAFE_ENDPOINT=http://127.0.0.1:8000 TYPESAFE_API_KEY=local \
TYPESAFE_PRICE_INPUT_PER_M=0 TYPESAFE_PRICE_OUTPUT_PER_M=0 \
python3 -m jevbench.cli run \
--tasks datasets/public/easy.jsonl,datasets/public/original.jsonl,datasets/public/hard.jsonl \
--adapter typesafe --results <out>.jsonl --raw-dir <raw>
python3 ../score_bench.py <out>.jsonl
Raw per-item outputs live in docs/bench/: one raw record per shipped model (231 rows each, all schema-valid; the v1.1c run went through the released pack — the pack itself is the measured artifact — with constants exactly as packed, one attempt per item). Pre-registration for the multimodal line: PREREG-KAPTEENI-V11C.md; The posted bench request is preserved as BENCH-REQUEST.md. The full experiment history is JEVBENCH.md.