Fig. 2 — The wire, as executable code
Wire API
Kapteeni serves the TypeSafe System One wire format.
Field names, bounds, and answer shapes mirror
docs.typesafe.ai/api; the reference's example requests run unchanged
against this server. The HTTP layer is Python stdlib
only; the format itself lives in kapteeni/contract.py —
deliberately framework-free and torch-free. The multimodal server
(kapteeni.serve_v11c) serves the same contract with the
optional state.image field added, and ships the demo
website at its own root.
§ 1Endpoints & authentication
jev-latest alias
plus a card for the served line.
serve_v11c)
additionally serves its self-contained demo website from the
same origin — no CORS setup; same protocol as the wire.
Authentication is optional: set KAPTEENI_API_KEY to require
a Authorization: Bearer <key> header; without it the
server accepts any or no key (local development).
| Status | Meaning |
|---|---|
| 200 | Evaluation succeeded; body carries model, answers, usage. |
| 401 | Missing or invalid API key (only when KAPTEENI_API_KEY is set). |
| 404 | No route for the path. |
| 422 | Validation failure (including a body that is not valid JSON); the error names the field. |
HTTP/1.1 422
{"error": {"message": "choice criteria must be a map of at least 2 options",
"field": "questions.department.criteria"}}
§ 2Request body
ticket.messages[0].text).
jev-latest
(alias, wire compat), kapteeni-v1 (legacy name for
-meticulous), kapteeni-v1-meticulous,
kapteeni-v1-intuit, kapteeni-v1-committee.
Any of them routes to the variant the server was launched with;
the response's model field always reports the served
variant's real name. On the multimodal server:
jev-latest and kapteeni-v1.1c.
state.image — measured limits (v1.1c, probed 2026-10-02)
| Limit | Value & behavior |
|---|---|
| Wire format | base64 PNG/JPEG, one image per request; text-only requests are unchanged. |
| Size cap | 8 MiB decoded bytes; larger → HTTP 422. |
| Pixel budget | the Qwen3.5 processor downscales above ~16.78 MP to its pixel budget — up to ~16,300 vision tokens per image (tokens ≈ pixels/1024; each covers a 32×32 px merged patch). |
| Validated distribution | quality validated at 640×640 (400 tokens) — larger resolutions are accepted but out-of-distribution, and per-question latency scales roughly linearly with pixels (a 6-option choice on a 12 MP photo is ~70k vision tokens of forward). |
| Bounding notice | any image above 640 px longest edge is bounded, and the response carries a plain-language notice: small text and fine detail may become unreadable, and accuracy may differ from the full-resolution result — precision-critical callers should pre-resize or crop to the region of interest. The demo page carries the same notice next to its attach control. |
| Usage accounting | usage.input_tokens counts the real image tokens (kapteeni.model_v11c.vision_tokens, anchored to the measured grids). |
§ 3The three primitives
true / false descriptions (string,
object, or array).
The absolute readout: noul is P(true) ∈ [0,1] with
no renormalization and no confidence field. P(A) + P(¬A) ≠ 1
is reference behavior, kept by construction.
{"state": "Help! My payouts have been failing for 3 days.",
"model": "jev-latest",
"questions": {"is_urgent": {"type": "noul",
"instructions": "Does this convey urgency?"}}}
The relative readout: per-option scores are softmaxed over the full option group. Probabilities are rounded to 4 decimals and the argmax entry is repaired so the decimal sum is exactly 1, matching the reference's published examples.
Levels are judged independently (per-level BCE, "does the state
match this level?") and normalized only at the API layer.
score is the expectation
Σ level_index × p_level over a 0-based index — the docs'
example 1.05 = 0×0.0 + 1×0.95 + 2×0.05 confirms the base.
legend and probabilities keys are string
level indices.
One request, three primitives — response shape
{
"model": "kapteeni-v1-meticulous",
"answers": {
"department": {
"type": "choice", "choice": "technical",
"probabilities": {"billing": 0.003, "technical": 0.995, "sales": 0.002},
"confidence": 0.9683 },
"is_urgent": { "type": "noul", "noul": 0.95 },
"frustration": {
"type": "score", "score": 1.095,
"legend": {"0": "Calm, just stating facts",
"1": "Frustrated but civil",
"2": "Very angry, strong language"},
"probabilities": {"0": 0.235, "1": 0.435, "2": 0.33},
"confidence": 0.0275 }
}
}
Values above are illustrative of the shape. Measured
values for the docs' own example requests are in § 6. Each
response also carries a usage object: input = state +
questions, output = serialized answers; token counts are from our
tokenizer (documented divergence in accounting, not semantics).
With an image — a real case (menu-en, from the shipped demo)
{
"state": {"image": "<base64 PNG — at most 8 MiB decoded bytes>",
"note": "A cafe menu."},
"model": "kapteeni-v1.1c",
"questions": {
"priciest": {"type": "choice",
"instructions": "Which drink is the most expensive?",
"criteria": {"Americano": null, "Cappuccino": null,
"Chai Latte": null, "Flat White": null,
"Latte": null, "Macchiato": null}},
"cheapest": {"type": "choice",
"instructions": "Which drink is the cheapest?",
"criteria": {"Americano": null, "Cappuccino": null,
"Chai Latte": null, "Flat White": null,
"Latte": null, "Macchiato": null}},
"price_gate": {"type": "noul",
"instructions": "Does the Macchiato cost more than $4.00?",
"criteria": {"true": "the price is above the threshold",
"false": "the price is not above it"}}
}
}
menu-en from the shipped demo page
(kapteeni/demo/cases.json): one image, three
independent questions — two choice, one noul — each forwarded on
its own with the image re-attached. The demo's gold annotations:
priciest = Flat White, cheapest = Latte,
price_gate = yes. Serve kapteeni-v1.1c and open the
server root to run it and see the distributions rendered next to
the gold.§ 4Confidence & determinism
confidence appears on choice and score answers only. The
reference's exact formula is not public; ours is
1 − H(p)/ln K, floored at 0 (a documented divergence, asserted by
the test suite). Run-to-run behavior is fully deterministic — a
documented improvement over the reference's std ≈ 0.01; ensemble
noise is a future flag. The multimodal line is no different: each
question is forwarded on its own — the image rides on every pass
where present — so independence stays exact, and an image
question costs ~4–7 s against ~0.2 s for a text-only
pass.
§ 5Serving
# packaged distribution (merged model + heads + constants; no local artifacts)
python3 -m kapteeni.serve --dist ./kapteeni-v1-meticulous-dist --port 8000
python3 -m kapteeni.serve --hf TriusAI/kapteeni-v1-intuit --port 8000
# multimodal line: API takes state.image; self-contained demo site at /
python3 -m kapteeni.serve_v11c --dist ./kapteeni-v1.1c-dist --port 8002
python3 -m kapteeni.serve_v11c --hf TriusAI/kapteeni-v1.1c --port 8002
# from local training artifacts
python3 -m kapteeni.serve --bundle model_cache/kapteeni_v1.pt \
--lora model_cache/kapteeni_p2/adapter --served-as kapteeni-v1-meticulous --port 8000
# two-variant committee (output-averaged answers)
python3 -m kapteeni.serve --dist ./kapteeni-v1-meticulous-dist \
--dist2 ./kapteeni-v1-intuit-dist --committee --port 8000
# mock model: no GPU, full wire contract (text server; add --mock for the v1.1c dry-run)
python3 -m kapteeni.serve --port 8000
| Flag | Effect |
|---|---|
--hf / --dist | load a packaged distribution from Hugging Face or a local directory |
--bundle / --lora / --fit | serve from training artifacts (head bundle, LoRA adapter, blend-constants JSON) |
--model | merged model directory for --bundle serving (default: the base backbone) |
--readout | blend (head + verbalizer, default), head, or verb |
--committee | ensemble with a second model (--bundle2/--lora2/--fit2 or --dist2) by output-averaging |
--served-as | name reported in responses and /v1/models (default: the dist's config, else kapteeni-v1) |
--host / --port | bind address (default 127.0.0.1:8000) |
--mock | (serve_v11c) interface + demo dry-run without a GPU |
--warmup / --no-warmup | (serve_v11c, default on) pre-run one forward per demo image size at startup (~9 s), amortizing the per-shape searches out of first requests |
§ 6Live demo — the docs' own examples, measured
Run against the served model by demo_requests.py; every
request below is taken unchanged from docs.typesafe.ai.
| Request | Answer |
|---|---|
| "Help! My payouts have been failing for 3 days." — urgency (noul) | 0.94 |
| same state — department (choice) | technical 0.99, confidence 0.96 |
| "API returning 500s, orders blocked" — is_urgent | 0.95 |
| same — frustration (score, 3 levels) | 1.10 (0.24/0.43/0.33), confidence 0.02 — honest uncertainty on a genuinely ambiguous case |
| State-page refund workflow (dot-paths) | refund_requested 0.99, policy_supports 0.98 |
| held-out fact check (FEVER-style) | 0.96 |