KAPTEENI a system one decision model Operating manual — MMXXVI

Fig. 1 — A Jev-compatible System One decision model

Kapteeni

Kapteeni reads a state and typed questions, and returns calibrated probability distributions your code can branch on. It generates no text. The reference behavior is the TypeSafe System One API (Jev); Kapteeni reimplements it end-to-end on local hardware — a frozen 4B backbone with trained readout heads: Qwen3-4B-Instruct-2507 on the text line, Qwen3.5-4B on the multimodal line (images and Chinese).

Status: two shipped lines. The text flagship — kapteeni-v1-meticulous (the default) and kapteeni-v1-intuit. The multimodal line — kapteeni-v1.1c, the only model here that can look at an image. All three are published, ready to serve, with every benchmark number measured on the public half of JevBench and the caveats stated next to them; see Benchmarks.

§ 1What it is

state + questions ──▶ serialize ──▶ 4B backbone ──▶ one forward pass per
                                                    question / option / level
                 ──▶ readout heads (+ native readout) ──▶ temperature ──▶ contract shaping
                 ──▶ {answers: typed + probabilities + confidence, usage}
Passes carry only the content, never the question ids — so adding or removing a question never changes another answer.

That structure is the whole model. There is no generated text to parse and no free-form reasoning step: every question becomes a few forward passes through a frozen backbone, and readout heads trained to produce calibrated numbers read the last hidden state off each pass. The multimodal line keeps this exact structure — it only adds the optional image field to the state, re-attached on every pass so an image question is simply several more forwards.

The behavior is the reference's, on purpose: its example requests run unchanged. Noul answers are absolute probabilities (P(A) + P(¬A) need not sum to 1 — that is how the reference behaves, and this reproduction keeps it). Choice distributions sum to exactly 1. Score is the expectation over independently judged levels, 0-based. Two documented divergences: confidence is derived from the distribution shape with our formula (1 − H(p)/ln K — the reference's exact formula is not public), and the model is fully deterministic where the reference shows small run-to-run jitter.

§ 2How to use it

Which model for your traffic

ModelReach for it when
kapteeni-v1-meticulousthe traffic is unknown, messy, or adversarial — and your code consumes the confidence values
kapteeni-v1-intuitthe traffic is well-formed (documents, policies, SLAs, forms) and needs numeric, temporal, or multi-step judgment
kapteeni-v1.1cthe state includes an image — screenshots, scanned forms, signs, photographs — in English or Chinese

Serve it

cd kapteeni
python3 -m pytest                          # 170+ tests, mock model, no GPU needed

# text line — from Hugging Face, or a local distribution pack
python3 -m kapteeni.serve --hf TriusAI/kapteeni-v1-meticulous --port 8000
python3 -m kapteeni.serve --dist ./kapteeni-v1-meticulous-dist --port 8000
python3 -m kapteeni.serve --dist ./kapteeni-v1-intuit-dist --port 8000

# multimodal line — API takes state.image; demo site at /
python3 -m kapteeni.serve_v11c --hf TriusAI/kapteeni-v1.1c --port 8002

curl localhost:8000/v1/systemone -d '{"state":"...","model":"jev-latest","questions":{...}}'

Requests may address any model by name, by the legacy kapteeni-v1, or by the alias jev-latest — all of them route to the model the server was launched with. The response's model field always reports the real name. Request shapes, error contract, and measured image limits are on the Wire API page.

See it work (demo site)

The multimodal server ships a self-contained demo page at its root — the fastest way to see what this does. Fifteen pre-configured cases: nine built fresh from the training generators (their gold answers exact by construction) and six real photographs, in English and Chinese. The photographs hand-checked at 11 of 12 answers correct, with the one miss documented in the record. Each case renders its gold annotation next to the model's live distribution, and the request JSON is editable before sending. Attach your own image: the page downscales it before sending and shows the same advisory notice the API returns.

Latency on the training box (AMD Strix Halo iGPU): ~0.2 s per text-only request, ~4–7 s per image question — the image rides on every pass. Images: at most 8 MiB decoded PNG/JPEG, quality validated at 640×640, anything larger gets an advisory notice on the response.

§ 3Where to get it

Every pack is self-contained — merged model, heads, serving constants, model card; the multimodal pack adds the serving package and the demo website. Nothing needs to be trained or merged locally. Weights CC BY-SA 4.0; code/generators Apache-2.0.
ModelDownloadNote
kapteeni-v1-meticuloushuggingface.co/TriusAI/kapteeni-v1-meticulousconservative confidence
kapteeni-v1-intuithuggingface.co/TriusAI/kapteeni-v1-intuitsharper decisions
kapteeni-v1.1chuggingface.co/TriusAI/kapteeni-v1.1cfrozen build 17affb8

§ 4How good it is

JevBench v1.4 public half — 231 items, official scoring code, through the live servers. Self-reported; no placement claims.
v1-meticulousv1-intuitv1.1c
Score65.7163.1863.72
Intelligence60.361.168.33
top-label ECE → Calibration0.0496 → 90.10.1196 → 76.10.0984 → 80.3
public accuracy (easy / standard / hard)0.710 (1.000 / 0.889 / 0.469)0.714 (1.000 / 0.903 / 0.468)0.766 (1.000 / 0.931 / 0.559)

Read honestly: at n=231 the composite gap between the multimodal line and the flagship (63.72 vs 65.71) is inside the run-to-run ±5.9pt interval — the claim is not separable, not better or worse. What is separately true: accuracy 0.766 (+5.6pt), hard tier 0.559 (+9.0pt), and v1.1c's image capability. Where the targeted skills moved: multi_hop 0.278 → 0.667, long_policy +10.5pt. The full story — the v1.1c gate table, per-family movements, calibration, the noise floor, and reproducibility — is on the Benchmarks page.

§ 5Design decisions

  1. Noul is trained absolute — a single sigmoid against soft targets, never a yes/no softmax pair. A complement-consistent Noul would be exactly the bug the reference's numbers (0.72 + 0.47 = 1.19) rule out; a test asserts non-identity.
  2. Choice softmax happens over the row's option group in-loss so gradients shape relative separation; at the API layer the distribution is rounded to 4 decimals with the argmax repaired so the decimal sum is exactly 1 (matching the reference's published examples).
  3. Score levels are judged independently (per-level BCE, "does the state match this level?"), normalized only at the API layer — this reproduces the reference's documented "numbers-only levels fail" asymmetry by construction.
  4. Soft labels from teacher k-sample agreement (mean of k judge calls, weight = 1 − std). Ambiguous rows are kept and down-weighted: the calibration band is learned from disagreement.
  5. Deterministic row splits by sha256(row_id) % 10 — the same rule in expansion, distillation, and training, so val rows never leak into train passes, and the distiller never labels val rows.
  6. No chat template — passes are plain text; behavior is template-independent and fully deterministic.
  7. The image serving policy is settled by measurement — images above a 640 px longest edge are bounded to the validated 640×640 rendering distribution, and every bounded image triggers a plain-language notice on the response; no pasting or cropping — letterboxing measurably flipped a boundary case (WORKLOG, 2026-10-02).

§ 6Repository layout

  • kapteeni/contract.pythe wire format as executable code — validation, answer shaping, confidence, exact-1 probability sums
  • kapteeni/serialize.pypass serialization (state JSON text + [noul]/[choice]/[score] suffixes; ids never serialized)
  • kapteeni/backbone.pyfrozen text-line backbone (bf16/ROCm), batched h_last precompute with resume
  • kapteeni/heads.py · train.pyper-primitive readout heads, proper-scoring losses (BCE/CE, soft targets), per-head temperature fit
  • kapteeni/build_data.py · criteria.py · distill.pydatasets → rows → passes; teacher-written rubrics; k-sample soft labels
  • kapteeni/verbalizer.pythe backbone's native decision readout, blended with the trained heads
  • kapteeni/model.py · serve.pytext-line runtime model + stdlib HTTP server
  • kapteeni/model_v11c.py · serve_v11c.pymultimodal runtime + server: state.image, same-origin demo site at /, one forward per question with the image re-attached
  • kapteeni/pack.py · pack_v11c.pybuild the self-contained distribution packs — merged model, heads, serving constants, model card (+ demo website for the multimodal line)
  • kapteeni/committee.pyoutput-averaging of two served variants (kapteeni-v1-committee)
  • kapteeni/p2_finalize.pyfits final temperatures and blend constants; emits a servable bundle
  • kapteeni/demo/demo cases + images, committed; scripts/make_demo.py regenerates them deterministically from the generators
  • tests/parity, asymmetry, score-semantics, and schema/end-to-end suites; the tests run against the mock by default and against the trained model via KAPTEENI_TEST_BUNDLE=
  • remaining modulesdata synthesis and experiment harnesses (synth, synth2(+zh), synth3(+families/zh/gfx), familyval, fit_diverse, evalx, report, metrics, mock, ollama) — see WORKLOG.md

§ 7License

  • Code, docs, generators: Apache-2.0 (LICENSE) — matching the Apache-2.0 base models, Qwen3-4B-Instruct-2507 and Qwen3.5-4B.
  • Trained weights: CC BY-SA 4.0, with the full training-data provenance and attribution guidance in WEIGHTS-LICENSE.md (the ShareAlike term comes from MultiNLI and the underlying FEVER annotations).