Jared Palmer, VP of Engineering at the AI lab Cognition, has released four free “decision models” built to slot into a competitor’s product without code changes. Palmer announced the family, called Kev, on his personal blog on 1 October. The post is unusually candid about what its numbers do and do not prove.

First, what a decision model is. It never writes a reply. You hand it a document, a question and a list of possible answers, and it returns a probability for each one. Your own software then decides what to do with, say, a 74 percent. Kev gets this behavior from a “pointer head,” a small added layer that scores the options on top of an ordinary language model. The document is read once, and each further question reuses that work.

The compatibility is the commercial point. Kev copies the interface of Jev, a decision model sold by TypeSafe, so a team with an app built on TypeSafe’s SDK only has to repoint two settings: the endpoint address and the model name. Palmer says he liked Jev’s interface but wanted open weights that he could host and fine-tune himself. He thanks TypeSafe for building “an API worth being compatible with.” An executive at one lab has therefore built a drop-in alternative to another company’s service, which is a pattern familiar from open-source database and storage clones.

The smallest of the four sizes has 0.8 billion parameters and the largest has 27 billion. All ship under the Apache 2.0 license on Hugging Face, with a server on GitHub that runs on GPUs or on Apple Silicon. The two largest models carry the headline figures. In Palmer’s own test harness, Kev-27B reaches 91.8 percent on questions about applying written policies and rules, and 90.8 percent when assigning consumer complaints to categories. Across a panel of 14 public datasets built to test generalisation, it manages 75.7 percent. In the post’s words, “all results are from the author’s evaluation harness, not an independent leaderboard.”

He lists further caveats himself. Model selection “was not fully blind,” because he knew test results for earlier candidates before choosing the released 27B. The base models come from Qwen, whose training data is unknown, so leaving a task out of Kev’s fine-tuning does not rule out that Qwen saw it. Palmer also says the size comparison “doesn’t isolate the effect of size or establish an advantage over a prompted chat model.”

The most credible passage is the regression. On short inputs, the new 27B performs about the same as its predecessor. On long legal contracts it does worse, missing more answers while sounding surer of itself, and Palmer tells readers to test the older checkpoint on their own documents if that is their workload. Few vendors volunteer a worse result in their launch post.

Overconfidence matters because decision models exist to be trusted at a threshold. Palmer reports Kev-27B’s calibration error (how far its stated probabilities drift from how often it is actually right) at 0.019 when all the data is pooled. Score each dataset separately and average, and the figure climbs to 0.073. He also tuned a confidence cutoff on development data to target 5 percent error. At that setting the 27B resolved 54 percent of the test questions, with errors landing at 5.5 percent, while the smallest model resolved only 15 percent. He warns those rates may not hold on other data. Estimated serving cost on Modal runs from $3.54 per million requests for the smallest model to $44.09 for the largest.

Training provenance is the other thing worth reading closely. The corpus is about 146,000 examples, and the full 27B training set is not public. Palmer states flatly, “No Jev outputs were used.” That is a claim about his own data that outsiders cannot check. It lands the day after OpenAI accused people associated with Moonshot AI of copying model reasoning, when training-data lineage is under scrutiny.

The limits are plain. The model’s job ends at scoring the answers it is handed. It writes no explanations and fetches no missing facts, so callers have to bring the relevant policy or evidence themselves, and also decide what happens when none of the options fits. Fine-tuning gains varied: one pass over 5,219 labelled complaints lifted Kev-4B from 80 to 90 percent, but a separate support workload trained on about 1,000 synthetic examples moved accuracy only from 68 to 74 percent.

For teams routing tickets, screening documents or gating agent actions, the useful question is cost per correct automated decision, and Kev gives a free baseline to check TypeSafe’s pricing against. Run it on your own labelled data first, because the only benchmark so far belongs to its author.

Reported from Jared Palmer’s personal blog, post “Introducing Kev,” published 1 October 2026.