AutoTrust Releases JEV-27B for Self-Hosted AI Agent Decisions
The open-weights model pairs fast, structured choices with text generation. Its close benchmark result against a hosted rival comes from AutoTrust’s own tests.
Loading page…
The open-weights model pairs fast, structured choices with text generation. Its close benchmark result against a hosted rival comes from AutoTrust’s own tests.
Listen to this story
AutoTrust AI has released JEV-27B under Apache-2.0 for developers who want agent decisions to run on their own infrastructure. Its detachable 108.9-million-parameter block sits atop frozen Qwen3.8-27B weights, routing bounded choices separately from open-ended generation; AutoTrust says one B200 serves both paths. Company-run tests show a near tie with hosted Jev 1.13 (84.07% vs. 83.85%), not an independently validated win. The release enables local experimentation, but AutoTrust flags weaknesses and advises that J
The decision block adds about 0.4% to the combined model; AutoTrust says training it took roughly 9.2 hours on one B200.
On a 16-option independently scored test, JEV-27B reached 96% of Jev 1.13’s accuracy but did not surpass it.
AutoTrust reports 137 ms median decision time and about 130 decisions per second on a B200; hosted API timings are not a controlled comparison.
An AI agent can now use AutoTrust AI’s newly released JEV-27B to choose a route or score an option without sending that decision to a hosted service. The open-weights model also retains a path for writing and reasoning. That combination gives developers a self-hosted option, though the strongest performance comparisons come from AutoTrust’s own tests.
AutoTrust released JEV-27B under the Apache-2.0 license, with model weights, a decision adapter, training and serving code, vLLM support, and evaluation reports. Customers can run it on their own infrastructure; AutoTrust says one NVIDIA B200 can serve both its decision and generation paths. The release follows JEV-9B, the company’s earlier model combining those two functions.
The new path handles bounded questions: yes or no, a choice from supplied options, or a rating from zero to five. AutoTrust says it answers in one processing pass and returns a probability for each option. An agent could use those answers to select a next step, while leaving open-ended writing to the generation path.
JEV-27B starts with a Qwen3.8-27B model whose existing weights stay frozen. AutoTrust trained a detachable decision block containing 108.9 million parameters, about 0.4% of the combined model. A router sends requests to that block or to the original generation path. The company says training the addition took about 9.2 hours on one B200.
AutoTrust tested whether adding the block changed the base model’s coding output. With the decision block switched off, all 164 completions in its HumanEval test matched the base model byte for byte. That is evidence for the tested configuration, not a demonstration that every kind of generation remains identical.
The model was trained on a public corpus of outputs from TypeSafe AI’s Jev 1.13. AutoTrust says JEV-27B shares no weights or code with TypeSafe and is not affiliated with it. That origin helps explain why AutoTrust compares the two decision models, but it does not make the resulting performance figures independent.
Across six public text-decision benchmark groups, AutoTrust reported an equal-weight average of 84.07% for JEV-27B and 83.85% for the hosted Jev 1.13 API. JEV-27B scored higher on four groups and lower on two. The close averages cannot establish a winner: AutoTrust ran the comparison itself.
AutoTrust’s equal-weight average for JEV-27B across six text-decision benchmark groups.
AutoTrust’s measurement of TypeSafe Jev 1.13 on the same groups; the comparison was not independently validated.
Other numbers answer different questions. AutoTrust reproduced published scores for four external models without rerunning them, so those figures are not a fresh same-environment comparison. It also reported that JEV-27B reached 96% of Jev 1.13’s accuracy on an independently scored test with 16 answer options. In that test, JEV-27B approached Jev but did not surpass it.
AutoTrust measured a median decision time of 137 milliseconds and about 130 decisions per second on one B200. Its release also cites slower third-party measurements for Jev’s hosted API. Those timings are not a controlled speed contest: the API results include network time, and the hardware and serving conditions differ.
A demonstration reel shows the self-hosted model making decisions for tasks including a flight search, a simulated drone course and a sample authorization-code review. Those examples show the intended range, not how reliably it handles unattended work. AutoTrust identifies weaknesses in multi-step reasoning, arithmetic, dates and adversarial inputs, and says the training data is English-centric.
The company says JEV-27B is not meant for high-stakes decisions and recommends using its confidence scores to decide when to accept an answer. AutoTrust plans to add the decision block to future models in its Guru family. For now, the release gives developers the weights and tools to test whether local control is worth using on their own, lower-stakes decisions.
Loading discussion...
Join the conversation
What would make you comfortable using them?
Be the first to share a perspective or an experience.
Reader comments
Newest comments first. Replies stay oldest first.