Fig. 1 — A Jev-compatible System One decision model
Kapteeni
Kapteeni reads a state and typed questions, and
returns calibrated probability distributions your code can branch
on. It generates no text. The reference behavior is the TypeSafe
System One API (Jev); Kapteeni
reimplements it end-to-end on local hardware — a frozen 4B backbone
with trained readout heads: Qwen3-4B-Instruct-2507 on the text
line, Qwen3.5-4B on the multimodal line (images and Chinese).
Status: two shipped lines. The text flagship — kapteeni-v1-meticulous (the default) and kapteeni-v1-intuit. The multimodal line — kapteeni-v1.1c, the only model here that can look at an image. All three are published, ready to serve, with every benchmark number measured on the public half of JevBench and the caveats stated next to them; see Benchmarks.
§ 1What it is
state + questions ──▶ serialize ──▶ 4B backbone ──▶ one forward pass per
question / option / level
──▶ readout heads (+ native readout) ──▶ temperature ──▶ contract shaping
──▶ {answers: typed + probabilities + confidence, usage}
That structure is the whole model. There is no generated text to
parse and no free-form reasoning step: every question becomes a
few forward passes through a frozen backbone, and readout heads
trained to produce calibrated numbers read the last hidden state
off each pass. The multimodal line keeps this exact structure —
it only adds the optional image field to the state,
re-attached on every pass so an image question is simply several
more forwards.
The behavior is the reference's, on purpose: its example requests run unchanged. Noul answers are absolute probabilities (P(A) + P(¬A) need not sum to 1 — that is how the reference behaves, and this reproduction keeps it). Choice distributions sum to exactly 1. Score is the expectation over independently judged levels, 0-based. Two documented divergences: confidence is derived from the distribution shape with our formula (1 − H(p)/ln K — the reference's exact formula is not public), and the model is fully deterministic where the reference shows small run-to-run jitter.
§ 2How to use it
Which model for your traffic
| Model | Reach for it when |
|---|---|
| kapteeni-v1-meticulous | the traffic is unknown, messy, or adversarial — and your code consumes the confidence values |
| kapteeni-v1-intuit | the traffic is well-formed (documents, policies, SLAs, forms) and needs numeric, temporal, or multi-step judgment |
| kapteeni-v1.1c | the state includes an image — screenshots, scanned forms, signs, photographs — in English or Chinese |
Serve it
cd kapteeni
python3 -m pytest # 170+ tests, mock model, no GPU needed
# text line — from Hugging Face, or a local distribution pack
python3 -m kapteeni.serve --hf TriusAI/kapteeni-v1-meticulous --port 8000
python3 -m kapteeni.serve --dist ./kapteeni-v1-meticulous-dist --port 8000
python3 -m kapteeni.serve --dist ./kapteeni-v1-intuit-dist --port 8000
# multimodal line — API takes state.image; demo site at /
python3 -m kapteeni.serve_v11c --hf TriusAI/kapteeni-v1.1c --port 8002
curl localhost:8000/v1/systemone -d '{"state":"...","model":"jev-latest","questions":{...}}'
Requests may address any model by name, by the legacy
kapteeni-v1, or by the alias jev-latest —
all of them route to the model the server was launched with. The
response's model field always reports the real name.
Request shapes, error contract, and measured image limits are on
the Wire API page.
See it work (demo site)
The multimodal server ships a self-contained demo page at its root — the fastest way to see what this does. Fifteen pre-configured cases: nine built fresh from the training generators (their gold answers exact by construction) and six real photographs, in English and Chinese. The photographs hand-checked at 11 of 12 answers correct, with the one miss documented in the record. Each case renders its gold annotation next to the model's live distribution, and the request JSON is editable before sending. Attach your own image: the page downscales it before sending and shows the same advisory notice the API returns.
Latency on the training box (AMD Strix Halo iGPU): ~0.2 s per text-only request, ~4–7 s per image question — the image rides on every pass. Images: at most 8 MiB decoded PNG/JPEG, quality validated at 640×640, anything larger gets an advisory notice on the response.
§ 3Where to get it
| Model | Download | Note |
|---|---|---|
| kapteeni-v1-meticulous | huggingface.co/TriusAI/kapteeni-v1-meticulous | conservative confidence |
| kapteeni-v1-intuit | huggingface.co/TriusAI/kapteeni-v1-intuit | sharper decisions |
| kapteeni-v1.1c | huggingface.co/TriusAI/kapteeni-v1.1c | frozen build 17affb8 |
§ 4How good it is
| v1-meticulous | v1-intuit | v1.1c | |
|---|---|---|---|
| Score | 65.71 | 63.18 | 63.72 |
| Intelligence | 60.3 | 61.1 | 68.33 |
| top-label ECE → Calibration | 0.0496 → 90.1 | 0.1196 → 76.1 | 0.0984 → 80.3 |
| public accuracy (easy / standard / hard) | 0.710 (1.000 / 0.889 / 0.469) | 0.714 (1.000 / 0.903 / 0.468) | 0.766 (1.000 / 0.931 / 0.559) |
Read honestly: at n=231 the composite gap between the multimodal line and the flagship (63.72 vs 65.71) is inside the run-to-run ±5.9pt interval — the claim is not separable, not better or worse. What is separately true: accuracy 0.766 (+5.6pt), hard tier 0.559 (+9.0pt), and v1.1c's image capability. Where the targeted skills moved: multi_hop 0.278 → 0.667, long_policy +10.5pt. The full story — the v1.1c gate table, per-family movements, calibration, the noise floor, and reproducibility — is on the Benchmarks page.
§ 5Design decisions
- Noul is trained absolute — a single sigmoid against soft targets, never a yes/no softmax pair. A complement-consistent Noul would be exactly the bug the reference's numbers (0.72 + 0.47 = 1.19) rule out; a test asserts non-identity.
- Choice softmax happens over the row's option group in-loss so gradients shape relative separation; at the API layer the distribution is rounded to 4 decimals with the argmax repaired so the decimal sum is exactly 1 (matching the reference's published examples).
- Score levels are judged independently (per-level BCE, "does the state match this level?"), normalized only at the API layer — this reproduces the reference's documented "numbers-only levels fail" asymmetry by construction.
- Soft labels from teacher k-sample agreement (mean of k judge calls, weight = 1 − std). Ambiguous rows are kept and down-weighted: the calibration band is learned from disagreement.
- Deterministic row splits by
sha256(row_id) % 10— the same rule in expansion, distillation, and training, so val rows never leak into train passes, and the distiller never labels val rows. - No chat template — passes are plain text; behavior is template-independent and fully deterministic.
- The image serving policy is settled by measurement — images above a 640 px longest edge are bounded to the validated 640×640 rendering distribution, and every bounded image triggers a plain-language notice on the response; no pasting or cropping — letterboxing measurably flipped a boundary case (WORKLOG, 2026-10-02).
§ 6Repository layout
- kapteeni/contract.pythe wire format as executable code — validation, answer shaping, confidence, exact-1 probability sums
- kapteeni/serialize.pypass serialization (state JSON text +
[noul]/[choice]/[score]suffixes; ids never serialized) - kapteeni/backbone.pyfrozen text-line backbone (bf16/ROCm), batched
h_lastprecompute with resume - kapteeni/heads.py · train.pyper-primitive readout heads, proper-scoring losses (BCE/CE, soft targets), per-head temperature fit
- kapteeni/build_data.py · criteria.py · distill.pydatasets → rows → passes; teacher-written rubrics; k-sample soft labels
- kapteeni/verbalizer.pythe backbone's native decision readout, blended with the trained heads
- kapteeni/model.py · serve.pytext-line runtime model + stdlib HTTP server
- kapteeni/model_v11c.py · serve_v11c.pymultimodal runtime + server:
state.image, same-origin demo site at/, one forward per question with the image re-attached - kapteeni/pack.py · pack_v11c.pybuild the self-contained distribution packs — merged model, heads, serving constants, model card (+ demo website for the multimodal line)
- kapteeni/committee.pyoutput-averaging of two served variants (
kapteeni-v1-committee) - kapteeni/p2_finalize.pyfits final temperatures and blend constants; emits a servable bundle
- kapteeni/demo/demo cases + images, committed;
scripts/make_demo.pyregenerates them deterministically from the generators - tests/parity, asymmetry, score-semantics, and schema/end-to-end suites; the tests run against the mock by default and against the trained model via
KAPTEENI_TEST_BUNDLE= - remaining modulesdata synthesis and experiment harnesses (synth, synth2(+zh), synth3(+families/zh/gfx), familyval, fit_diverse, evalx, report, metrics, mock, ollama) — see WORKLOG.md
§ 7License
- Code, docs, generators: Apache-2.0 (LICENSE) — matching the Apache-2.0 base models, Qwen3-4B-Instruct-2507 and Qwen3.5-4B.
- Trained weights: CC BY-SA 4.0, with the full training-data provenance and attribution guidance in WEIGHTS-LICENSE.md (the ShareAlike term comes from MultiNLI and the underlying FEVER annotations).