KAPTEENI a system one decision model Operating manual — MMXXVI

Fig. 2 — The wire, as executable code

Wire API

Kapteeni serves the TypeSafe System One wire format. Field names, bounds, and answer shapes mirror docs.typesafe.ai/api; the reference's example requests run unchanged against this server. The HTTP layer is Python stdlib only; the format itself lives in kapteeni/contract.py — deliberately framework-free and torch-free. The multimodal server (kapteeni.serve_v11c) serves the same contract with the optional state.image field added, and ships the demo website at its own root.

§ 1Endpoints & authentication

POST /v1/systemone Validate the request, evaluate every question against the state, return typed answers and usage.
GET /v1/models Model cards: the jev-latest alias plus a card for the served line.
GET / & GET /demo/* The multimodal server (serve_v11c) additionally serves its self-contained demo website from the same origin — no CORS setup; same protocol as the wire.

Authentication is optional: set KAPTEENI_API_KEY to require a Authorization: Bearer <key> header; without it the server accepts any or no key (local development).

StatusMeaning
200Evaluation succeeded; body carries model, answers, usage.
401Missing or invalid API key (only when KAPTEENI_API_KEY is set).
404No route for the path.
422Validation failure (including a body that is not valid JSON); the error names the field.
HTTP/1.1 422
{"error": {"message": "choice criteria must be a map of at least 2 options",
           "field": "questions.department.criteria"}}

§ 2Request body

state required string, object, or array. Structured states are addressed from instructions with dot-paths (e.g. ticket.messages[0].text).
state.image optional, on the multimodal server only: base64 PNG/JPEG, one image per request, at most 8 MiB decoded bytes (larger → HTTP 422). Measured limits below.
model required string. Accepted names: jev-latest (alias, wire compat), kapteeni-v1 (legacy name for -meticulous), kapteeni-v1-meticulous, kapteeni-v1-intuit, kapteeni-v1-committee. Any of them routes to the variant the server was launched with; the response's model field always reports the served variant's real name. On the multimodal server: jev-latest and kapteeni-v1.1c.
questions required map of question id to question; non-empty. Ids must be non-empty strings. Ids never reach the model — passes are content-only, so adding or removing a question never changes another answer.

state.image — measured limits (v1.1c, probed 2026-10-02)

LimitValue & behavior
Wire formatbase64 PNG/JPEG, one image per request; text-only requests are unchanged.
Size cap8 MiB decoded bytes; larger → HTTP 422.
Pixel budgetthe Qwen3.5 processor downscales above ~16.78 MP to its pixel budget — up to ~16,300 vision tokens per image (tokens ≈ pixels/1024; each covers a 32×32 px merged patch).
Validated distributionquality validated at 640×640 (400 tokens) — larger resolutions are accepted but out-of-distribution, and per-question latency scales roughly linearly with pixels (a 6-option choice on a 12 MP photo is ~70k vision tokens of forward).
Bounding noticeany image above 640 px longest edge is bounded, and the response carries a plain-language notice: small text and fine detail may become unreadable, and accuracy may differ from the full-resolution result — precision-critical callers should pre-resize or crop to the region of interest. The demo page carries the same notice next to its attach control.
Usage accountingusage.input_tokens counts the real image tokens (kapteeni.model_v11c.vision_tokens, anchored to the measured grids).

§ 3The three primitives

noul {type, noul}
instructions required string, object, or array.
criteria optional object with optional true / false descriptions (string, object, or array).

The absolute readout: noul is P(true) ∈ [0,1] with no renormalization and no confidence field. P(A) + P(¬A) ≠ 1 is reference behavior, kept by construction.

{"state": "Help! My payouts have been failing for 3 days.",
 "model": "jev-latest",
 "questions": {"is_urgent": {"type": "noul",
                              "instructions": "Does this convey urgency?"}}}
choice {type, choice, probabilities, confidence}
instructions required string, object, or array.
criteria required map of 2–255 options; keys are non-empty strings, values are string, object, array, or null.

The relative readout: per-option scores are softmaxed over the full option group. Probabilities are rounded to 4 decimals and the argmax entry is repaired so the decimal sum is exactly 1, matching the reference's published examples.

score {type, score, legend, probabilities, confidence}
instructions required string, object, or array.
criteria required ordered array of 2–10 level descriptions (string, object, or array).

Levels are judged independently (per-level BCE, "does the state match this level?") and normalized only at the API layer. score is the expectation Σ level_index × p_level over a 0-based index — the docs' example 1.05 = 0×0.0 + 1×0.95 + 2×0.05 confirms the base. legend and probabilities keys are string level indices.

One request, three primitives — response shape

{
  "model": "kapteeni-v1-meticulous",
  "answers": {
    "department": {
      "type": "choice", "choice": "technical",
      "probabilities": {"billing": 0.003, "technical": 0.995, "sales": 0.002},
      "confidence": 0.9683 },
    "is_urgent": { "type": "noul", "noul": 0.95 },
    "frustration": {
      "type": "score", "score": 1.095,
      "legend": {"0": "Calm, just stating facts",
                 "1": "Frustrated but civil",
                 "2": "Very angry, strong language"},
      "probabilities": {"0": 0.235, "1": 0.435, "2": 0.33},
      "confidence": 0.0275 }
  }
}

Values above are illustrative of the shape. Measured values for the docs' own example requests are in § 6. Each response also carries a usage object: input = state + questions, output = serialized answers; token counts are from our tokenizer (documented divergence in accounting, not semantics).

With an image — a real case (menu-en, from the shipped demo)

{
  "state": {"image": "<base64 PNG — at most 8 MiB decoded bytes>",
             "note": "A cafe menu."},
  "model": "kapteeni-v1.1c",
  "questions": {
    "priciest": {"type": "choice",
                  "instructions": "Which drink is the most expensive?",
                  "criteria": {"Americano": null, "Cappuccino": null,
                                "Chai Latte": null, "Flat White": null,
                                "Latte": null, "Macchiato": null}},
    "cheapest": {"type": "choice",
                  "instructions": "Which drink is the cheapest?",
                  "criteria": {"Americano": null, "Cappuccino": null,
                                "Chai Latte": null, "Flat White": null,
                                "Latte": null, "Macchiato": null}},
    "price_gate": {"type": "noul",
                    "instructions": "Does the Macchiato cost more than $4.00?",
                    "criteria": {"true": "the price is above the threshold",
                                  "false": "the price is not above it"}}
  }
}
Case menu-en from the shipped demo page (kapteeni/demo/cases.json): one image, three independent questions — two choice, one noul — each forwarded on its own with the image re-attached. The demo's gold annotations: priciest = Flat White, cheapest = Latte, price_gate = yes. Serve kapteeni-v1.1c and open the server root to run it and see the distributions rendered next to the gold.

§ 4Confidence & determinism

confidence appears on choice and score answers only. The reference's exact formula is not public; ours is 1 − H(p)/ln K, floored at 0 (a documented divergence, asserted by the test suite). Run-to-run behavior is fully deterministic — a documented improvement over the reference's std ≈ 0.01; ensemble noise is a future flag. The multimodal line is no different: each question is forwarded on its own — the image rides on every pass where present — so independence stays exact, and an image question costs ~4–7 s against ~0.2 s for a text-only pass.

§ 5Serving

# packaged distribution (merged model + heads + constants; no local artifacts)
python3 -m kapteeni.serve --dist ./kapteeni-v1-meticulous-dist --port 8000
python3 -m kapteeni.serve --hf TriusAI/kapteeni-v1-intuit --port 8000

# multimodal line: API takes state.image; self-contained demo site at /
python3 -m kapteeni.serve_v11c --dist ./kapteeni-v1.1c-dist --port 8002
python3 -m kapteeni.serve_v11c --hf TriusAI/kapteeni-v1.1c --port 8002

# from local training artifacts
python3 -m kapteeni.serve --bundle model_cache/kapteeni_v1.pt \
    --lora model_cache/kapteeni_p2/adapter --served-as kapteeni-v1-meticulous --port 8000

# two-variant committee (output-averaged answers)
python3 -m kapteeni.serve --dist ./kapteeni-v1-meticulous-dist \
    --dist2 ./kapteeni-v1-intuit-dist --committee --port 8000

# mock model: no GPU, full wire contract (text server; add --mock for the v1.1c dry-run)
python3 -m kapteeni.serve --port 8000
FlagEffect
--hf / --distload a packaged distribution from Hugging Face or a local directory
--bundle / --lora / --fitserve from training artifacts (head bundle, LoRA adapter, blend-constants JSON)
--modelmerged model directory for --bundle serving (default: the base backbone)
--readoutblend (head + verbalizer, default), head, or verb
--committeeensemble with a second model (--bundle2/--lora2/--fit2 or --dist2) by output-averaging
--served-asname reported in responses and /v1/models (default: the dist's config, else kapteeni-v1)
--host / --portbind address (default 127.0.0.1:8000)
--mock(serve_v11c) interface + demo dry-run without a GPU
--warmup / --no-warmup(serve_v11c, default on) pre-run one forward per demo image size at startup (~9 s), amortizing the per-shape searches out of first requests

§ 6Live demo — the docs' own examples, measured

Run against the served model by demo_requests.py; every request below is taken unchanged from docs.typesafe.ai.

RequestAnswer
"Help! My payouts have been failing for 3 days." — urgency (noul)0.94
same state — department (choice)technical 0.99, confidence 0.96
"API returning 500s, orders blocked" — is_urgent0.95
same — frustration (score, 3 levels)1.10 (0.24/0.43/0.33), confidence 0.02 — honest uncertainty on a genuinely ambiguous case
State-page refund workflow (dot-paths)refund_requested 0.99, policy_supports 0.98
held-out fact check (FEVER-style)0.96