KAPTEENI a system one decision model Operating manual — MMXXVI

Fig. 3 — Measured runs

Benchmarks & Calibration

All runs on this page are self-reported measurements on the public half of JevBench v1.4 — 231 items: easy 48 / standard 72 / hard 111 — using the benchmark's own official scoring code, through the live server, end to end.

§ 1The shipped lines, side by side

v1-meticulousv1-intuitv1.1c (multimodal)
JevBench-style score (public half)65.7163.1863.72
Intelligence60.361.168.33
top-label ECE → Calibration0.0496 → 90.10.1196 → 76.10.0984 → 80.3
public accuracy (easy / standard / hard)0.710 (1.000 / 0.889 / 0.469)0.714 (1.000 / 0.903 / 0.468)0.766 (1.000 / 0.931 / 0.559)
skills-slice accuracy (609 held-out items)0.6590.814— *
skills-slice ECE0.0630.043— *
p50 / p95 latency0.17 s / 1.2 ssame0.26 s / 10.6 s
tokens/decision597same562
* v1.1c was not run on the 609-item skills slice; its pre-registered gate on synth2-EN-val (n=609, English rule skills) reads 0.9048 against the ≥ 0.90 bar. Latency: measured on the training box (AMD Strix Halo iGPU, ×2 self-hosted adjustment for the adjusted p50 0.34 s); v1.1c's p95 is pushed by hard items' deep multi-pass chains. Backbone: Qwen3-4B-Instruct-2507 (text) vs Qwen3.5-4B (v1.1c) — text-only requests are unchanged in the wire.

Robust claims: -intuit ≫ -meticulous on the skills slice (0.814 vs 0.659 on 609 items, ~20σ); -meticulous ≫ -intuit on bench ECE (0.05 vs 0.12, beyond single-run resolution); v1.1c's composite is not separable from v1 (63.72 vs 65.71, inside the ±5.9pt CI) — what is separable: hard tier +9.0pt (0.469 → 0.559), accuracy +5.6pt, matching Intelligence (+8.0), and the image/Chinese capability. A measured pipeline noise floor (~±2–3 composite points across runs differing only by seed and serving constants) bounds everything finer-grained: improvements below it are not evidence.

The benchmark updated to v1.4.1/v1.4.2 during this work (roster now 93 systems, 89 ranked; at run time the newest clone is bb05a33 — harness code and the 231 public items unchanged). Verified: the 231 public items, the 308-item sealed set, and the composite formula are unchanged — the official arithmetic was reproduced exactly — so the runs of record remain valid as scored. The frozen v1.5 method (904 open / 720 sealed) is not self-runnable from this repo and goes through Benchmark Heaven's submission pipeline.

kapteeni-v1.1c — the pre-registered gates, all passed in its one training run

These bars were deliberately hard; they had never been met by any earlier training configuration of this pipeline. The full history, including the runs that came close and failed, is in WORKLOG.md.
GateReadingBar
English image decisions (synth3-val, n=590)0.9661> 0.8068 (the frozen backbone's own score)
Chinese image decisions (synth3zh-val, n=204)0.9412reported, no bar
English text — never trained on (MNLI-noul, n=150)0.8800≥ 0.84
Chinese text — never trained on (OCNLI-noul, n=150)0.8467≥ 0.83
Chinese rule skills (synth2zh-val, n=403)0.9132≥ 0.90
English rule skills (synth2-EN-val, n=609)0.9048≥ 0.90
calibration after the fit (noul / choice / score)0.024 / 0.018 / 0.069≤ 0.10 each

§ 2The four axes (official formulas)

Composite = equal-weight harmonic mean of the four axes, × the (Intelligence/50)² gate for Intelligence < 50 — all three shipped models clear it. Jev-class eligibility for the two text variants (the benchmark's published definition: cost ≤ $0.080/1k decisions, adjusted median latency ≤ 1.30 s): both qualify at verified prices; the multimodal line's eligibility is not claimed in the record.
Axismeticulousintuitv1.1cHow it is measured
Intelligence60.361.168.33chance-corrected tier accuracy, weights easy .14 / standard .28 / hard .30 (renormalized — the judge tier is sealed); from public accuracy 0.710 / 0.714 / 0.766, 95% CI ±5.9pt (n=231)
Calibration90.176.180.3from top-label ECE 0.0496 / 0.1196 / 0.0984 across all 231 answers (ECE half only; the benchmark also averages a private gold-distribution half)
Speed81.181.069.6server latency p50 0.17 s, p95 1.2 s on an AMD Strix Halo iGPU, ×2 self-hosted adjustment (adjusted p50 0.34 s); v1.1c p50 0.26 s / p95 10.6 s — hard items run deep multi-pass chains, and image questions re-attach the image on every pass
Cost42.442.3—597 input tokens/decision × hosted list price — verified 2026-09-26: $0.03/M (Novita, qwen3-4b-fp8) = $0.018/1k decisions; nearest official Alibaba tier (qwen-turbo, $0.05/M) = $0.030/1k; the earlier $0.14/M assumption is kept as the pessimistic bound. v1.1c: 562 tokens/decision — its Cost axis is not computed in the record

§ 3Per-tier accuracy (231 items)

tiernmeticulousintuitv1.1c
easy481.0001.0001.000
standard720.8890.9030.931
hard1110.4690.4680.559
pooled2310.7100.7140.766

§ 4Per-family accuracy (public items — v1-meticulous)

familynaccfamilynacc
fact121.000routing120.750
tool_selection121.000ambiguous70.429
ordinal121.000probability100.400
extraction240.917long_policy190.316
intent240.958multi_hop180.278
policy120.917temporal_numeric150.200
adequacy120.917tradeoff60.167
adversarial61.000judge_hard170.706
trap80.875routing_hard51.000

The -intuit variant was developed for exactly the weak families above; on its 609-item held-out skills slice it reaches temporal_numeric 0.811 / multi_hop 0.917 / long_policy 0.644. It was not run family-by-family on the public items; only its tier accuracies above are public-half measured. Ordinal expectation MAE 0.482 levels; paraphrase consistency 0.861 (meticulous).

v1.1c — families that moved (vs v1-meticulous; n unchanged)

Declines are within noise at these n. temporal_numeric is the arc-level negative: the bench's narrow temporal templates do not accept the synthetic temporal training — it remains the frontier for every model this repository has trained.
familyv1v1.1cΔ
multi_hop (n=18)0.2780.667+38.9pt
long_policy (n=19)0.3160.421+10.5pt
extraction (n=24)0.9171.000+8.3pt
trap (n=8)0.8751.000+12.5pt
judge_hard (n=17)0.7060.765+5.9pt
tradeoff (n=6)0.1670.333+16.7pt
adversarial (n=6)1.0000.833−16.7pt
routing (n=12)0.7500.667−8.3pt
temporal_numeric (n=15)0.2000.133−6.7pt

§ 5Calibration — the shipped models

Top-label expected calibration error across all 231 public answers, scored by the benchmark's own code (ECE half only; the benchmark also averages a private gold-distribution half):

Read with the noise-floor note in § 1: at n=231, ECE differences below ~0.02 between single runs are unresolvable.
Modeltop-label ECE→ CalibrationServing constants
kapteeni-v1-meticulous0.049690.1head temperatures + blend weights fitted on the mixed-domain val set, through the merged model
kapteeni-v1-intuit0.119676.1fitted on the deployment-diverse held-out set, per pre-registration
kapteeni-v1.1c0.098480.3per-primitive temperatures on combined held-out val — the only constants fitted anywhere in the line; fitted ECE noul 0.024 / choice 0.018 / score 0.069, each within the ≤ 0.10 bar (PREREG-KAPTEENI-V11C.md)

§ 6Run history

  1. 01

    kapteeni-v1-meticulous, the flagship (2026-09-25)

    LoRA r=32 on all projections + retrained heads over the 8.7M-token mix; MNLI held out of training as the OOD gate (0.88 → 0.893). All serving constants refit on mixed-domain val through the merged model — zero public-half selection in the chain.

  2. 02

    kapteeni-v1-intuit (2026-09-26)

    The second text variant — same architecture and wire format; the weak-family skills added to training and serving constants fitted on the deployment-diverse held-out set, per pre-registration. Sharper where the flagship is conservative, with calibration traded accordingly; the guidance of when to prefer it lives in the table above.

  3. 03

    kapteeni-v1.1c, the multimodal line (2026-10-02)

    The same phase pipeline ported to Qwen3.5-4B (images + Chinese); every pre-registered gate passed on the first reading. Run of record on the released pack: ECE 0.0984; not separable from v1 on the composite (inside the ±5.9pt CI). Published pinned at 17affb8; bench requests open for JevBench and Image JevBench (fstandhartinger/jevbench#178).

The records behind these three shipped models — every pre-registration, gate, and run (successful or not) — are kept in full in JEVBENCH.md and WORKLOG.md; a design proposal for a deployment-oriented benchmark alternative lives in DECISIONBENCH-DRAFT.md.

§ 7How to read every number above

  • Sample sizes are small. All accuracies are on 231 items; one item is 0.43 accuracy points. The 95% binomial CI on public accuracy is ±5.9pt (0.710 / 0.714 / 0.766). Tier-level CIs are wider (hard: ±9.3pt at 0.469).
  • Small composite gaps are ties. Differences of a few points between any two runs sit inside run-to-run variance plus axis assumptions (the Cost axis alone swings ±7 points across the stated price range).
  • A measured pipeline noise floor. Runs differing only by training seed and val-fit constants scored ~2.6 composite points apart — treat ~±2–3 composite points as the single-seed noise floor.
  • ECE at n=231 is biased and high-variance. Differences below ~0.02 between single runs are unresolvable; 0.05-vs-0.07 comparisons between runs should not be read as meaningful.
  • Calibration is measured on the benchmark's distribution, which is OOD relative to any deployment: a model whose constants are tuned for the bench mix is not thereby calibrated for yours. The record documents this failure mode in the lab, not just in theory (JEVBENCH).
  • Intelligence is renormalized over the three public tiers because the judge tier is sealed; any score that folds the sealed+judge weight in is not comparable to ours. Jev's accuracy on the same public items is still higher (0.866 vs 0.710).
  • Composite gaps sit inside the CI. 63.72 (v1.1c) vs 65.71 (v1) is not separable at n=231 — the honest claim is "not separable on the composite with this sample size," not a win or a loss. The separable claims are recorded exactly: the hard-tier jump (+9.0pt), the accuracy gain (+5.6pt), and the image/Chinese capability.

§ 8Reproducibility

# server with the trained bundle
python3 -m kapteeni.serve --bundle model_cache/kapteeni_v1.pt \
    --lora model_cache/kapteeni_p2/adapter --served-as kapteeni-v1-meticulous --port 8000
# multimodal line through the released pack
python3 -m kapteeni.serve_v11c --hf TriusAI/kapteeni-v1.1c --port 8002

# bench (from jevbench/)
TYPESAFE_ENDPOINT=http://127.0.0.1:8000 TYPESAFE_API_KEY=local \
TYPESAFE_PRICE_INPUT_PER_M=0 TYPESAFE_PRICE_OUTPUT_PER_M=0 \
python3 -m jevbench.cli run \
  --tasks datasets/public/easy.jsonl,datasets/public/original.jsonl,datasets/public/hard.jsonl \
  --adapter typesafe --results <out>.jsonl --raw-dir <raw>
python3 ../score_bench.py <out>.jsonl

Raw per-item outputs live in docs/bench/: one raw record per shipped model (231 rows each, all schema-valid; the v1.1c run went through the released pack — the pack itself is the measured artifact — with constants exactly as packed, one attempt per item). Pre-registration for the multimodal line: PREREG-KAPTEENI-V11C.md; The posted bench request is preserved as BENCH-REQUEST.md. The full experiment history is JEVBENCH.md.