Addis PulseStudio

Kev: A Small Classifier Family That Will Not Generate Your Video

Jared Palmer's 0.8B–9B Qwen3.5-based decision models ship pretrained weights, a TypeSafe-compatible API, and calibrated probabilities. They route, gate, and score. They do not diffuse.

4 min read841 words

What happened

Jared Palmer published Kev, a family of three small decision-classification models (0.8B, 4B, 9B parameters) built on Qwen3.5 bases, with pretrained weights, training code, and a TypeSafe System One-compatible API. All three sizes share the same training data and settings, and the 4B and 9B models run in 32 GB of RAM on Apple Silicon in bf16 precision.

Context

Kev is not the first release in this line. A prior generation ran on Qwen3 bases (the repository carries a revision tag @qwen3), so this is at least the second architecture iteration. The models derive from a reference architecture documented as "Jev's Architecture Unmasked." The 2026-09-21 release shipped with a same-day second training pass: policy cases with explicit day counts and evidence-removal cases trained toward a uniform answer. That pass moved Kev-9B's test-set Brier from 0.837 to 0.852 (95% CI +0.8 to +2.9 points).

How it works

Kev is a classifier, not a generator. You give it a free-text state and it returns structured answers across three question types in a single request: binary yes/no (noul), multiple-choice (choice), or ordinal rating (score). Multiple questions in one call share the input text but cannot read each other's outputs. Each response carries per-question confidence values, full probability distributions, token usage counts, and latency in milliseconds.

Probabilities are calibrated by default: each checkpoint stores a fitted temperature (roughly 2.1 to 2.4) from its in-distribution development set, so the numbers are usable as thresholds rather than raw logits. Inference runs locally on CUDA, ROCm, or Apple Silicon. The 4B model in bf16 on an Apple M5 produced a response in 495 ms for 101 input tokens and 161 output tokens. A Python HTTP server exposes a TypeSafe System One-compatible endpoint; the SDK ships in the uv sync --extra serve dependency set. Running locally needs Python 3.12 or later and the uv package manager.

Our read

The number people will quote is the 3.5-point Brier gap between Kev-9B (0.822) and Jev (0.857) on the development set. That comparison is not controlled: Jev's training datasets are unknown, and the test set where Kev-9B scores 0.852 is one Jev has not been evaluated on. Treating those as a like-for-like race misreads what the repository claims.

More interesting is the second training pass. The evidence-removal cases trained toward a uniform answer are a calibration move: the model is being pushed to say "I don't know" when the deciding signal is stripped, not to guess. For a studio building triage on top of this, that matters more than the Brier score. A classifier that outputs 0.5/0.5 on ambiguous input is useful; one that outputs 0.7/0.3 with equal uncertainty is not.

The TypeSafe System One API compatibility is the underappreciated detail. Integration is a standard SDK call, not a custom protocol. The chess demo is the cleanest demonstration: the board is the input, legal moves become choice options, and a score question rates the position. Kev will not generate video or run a diffusion step. It is a routing and gating layer, and that is the whole point.

What this changes

For a studio already running ComfyUI on Apple Silicon or a mid-range GPU, Kev-4B fits in the same 32 GB of RAM in bf16 with no additional hardware. The concrete Monday: stand up the Python server in a venv alongside the existing stack, point the TypeSafe SDK at localhost, and build three endpoints. Route client feedback or QA notes to the right workflow step using choice. Gate rendered output with a noul go/no-go check. Score prompt drafts or footage notes with score. The 495 ms latency on an M5 is negligible in a batch pipeline. Set a confidence threshold on the calibrated probabilities to split auto-approve from human-review. Nothing in the video generation stack itself changes.

License

The sources do not state a licence for the Kev weights, training code, or evaluation data. The repository references "model cards and PLAN.md" for details, but no specific licence (Apache-2.0, MIT, Llama Community License, research-only) is named in the material available here. Before building anything commercial on these weights, check the model card on Hugging Face or the repository directly.

Key takeaways

  • Kev is a decision-classification model family (0.8B / 4B / 9B on Qwen3.5), not a generation model; it routes, gates, and scores rather than produces content.
  • The 3.5-point Brier gap versus Jev on the dev set is not a controlled comparison: Jev's training data is unknown and the test set is exclusive to Kev.
  • Calibrated probabilities (temperature ~2.1–2.4) and the evidence-removal training pass make Kev's "I don't know" output operationally meaningful for triage workflows.
  • Kev-4B runs in 32 GB RAM on Apple Silicon in bf16 at 495 ms per response, so it coexists with a ComfyUI pipeline without extra hardware.
  • No licence is stated in the available sources; verify the model card before shipping anything commercial built on these weights.

Sources

  1. Kev: Tiny Jev-like family of decision models built on top of Qwen3.5 — tier 3
llmsqwenroutingclassifiers

How this post was made

Drafted from clustered primary sources by the models below, then read, edited and approved by a human before it was published. The sources are listed in full at the end of the article.

Drafted
Independent sources
1
cluster pair
gemma4:12b
cluster label
gemma4:12b
radar brief
gemma4:12b
research brief
qwen3.8:27b
draft article
qwen3.8:27b
short script
qwen3.8:27b
seo pack
gemma4:12b
Run
editorial-20260921T223712Z