CISION PR Newswire - ซิชั่น พีอาร์ นิวส์ไวร์
![]() |
Trained in about 9.2 hours on one NVIDIA B200, JEV-27B adds fast, calibrated System 1 decisions to a frozen Qwen3.8-27B backbone while preserving its System 2 generation path.
SINGAPORE, Sept. 29, 2026 /PRNewswire/ -- AutoTrust AI today released JEV-27B, an Apache-2.0 open-weights model designed to handle frequent, structured decisions inside AI agent workflows while retaining the underlying model's full generation and reasoning path.
JEV-27B answers yes/no, multiple-choice and 0–5 rating questions in a single forward pass and returns a calibrated probability for every option. It runs on one NVIDIA B200 inside a customer's own infrastructure and serves both fast System 1 decisions and deliberate System 2 generation from a single set of weights.
The model trains a 108.9-million-parameter decision block—about 0.4% of the full model—on top of a frozen Qwen3.8-27B backbone. AutoTrust reports that training took approximately 9.2 B200-hours. With the decision block switched off, all 164 HumanEval completions were byte-identical to those produced by the base model.
"Jev proved there is real demand for models that decide rather than write," said Daniel Tang, AutoTrust AI's chief executive and co-founder. "JEV-27B shows that this capability can run on one GPU inside a customer's own infrastructure, next to a reasoning model. For companies that cannot send every decision to a third-party API, that changes both the cost and the risk."
Evaluation and evidence
AutoTrust AI evaluated JEV-27B across six public text-decision benchmark groups. It reported scores of 88.70% on JevBench, 83.75% on Kev, 73.89% on OpenJev text, 92.91% on Nimble, 77.46% on VitaminC and 87.71% on MASSIVE-en, for an equal-weight six-group mean of 84.07%.
For additional context, AutoTrust AI also ran the hosted TypeSafe Jev 1.13 API on the same benchmark groups and reported a six-group mean of 83.85%, with JEV-27B scoring higher on four groups and lower on two. Because AutoTrust conducted this comparison itself, the figures should be read as internal comparative evidence—not as an independent third-party validation or a claim of across-the-board superiority.
For public baseline context, AutoTrust reproduced the scores published by TokenRhythm for NeoHorse-Jev, Open-Jev, Kev and Laya English. AutoTrust did not rerun those four external baselines. The pinned source table is available at https://huggingface.co/TokenRhythm/NeoHorse-Jev-4B/blob/b50e043e22e0e41e7fc0c244e4daa707b8124930/README.md. These figures are included as published benchmark context rather than as a new same-environment comparison by AutoTrust.
JEV-27B was also evaluated for fidelity to its distillation target. On 25,376 held-out questions labeled with Jev 1.13 probability distributions, JEV-27B reported a mean KL divergence of 0.017, where zero means identical distributions. On decision-models-under-pressure, an independent benchmark scored against human labels, JEV-27B reached 96% of Jev 1.13's accuracy with 16 answer options. In that independent test, JEV-27B approached—but did not exceed—Jev 1.13.
AutoTrust AI measured a median decision latency of 137 milliseconds and sustained throughput of about 130 decisions per second on one B200. The model card also cites third-party measurements of 238 to 301 milliseconds and 23 decisions per second for Jev's hosted API. These are not controlled, like-for-like results: the hosted API measurements include network time, while AutoTrust AI's local measurements do not, and the hardware, serving and concurrency conditions differ.
How it works
AutoTrust AI uses the terms System 1 for fast, typed decisions and System 2 for deliberate generation and reasoning. JEV-27B serves both from one set of weights. It follows JEV-9B as the company's second integrated System 1 and System 2 open model.
The model is built with AutoTrust AI's Blocks of Experts recipe. A strong pretrained model, Alibaba's open-weights Qwen3.8-27B, stays frozen as one expert block. A small, detachable block is trained for a single skill, and a router sends each request either to the fast decision block or to the deliberate generation block.
The decision block holds 108.9 million trained parameters, 0.4% of the model, and took about 9.2 hours to train on one NVIDIA B200. The reasoning path was left untouched. With the decision block switched off, JEV-27B scores 78.0% on the HumanEval coding test, and all 164 of its completions are byte-identical to the base model's.
In a demonstration reel released with the model, a self-hosted JEV-27B served as the decision engine for 10 tasks. It played Doom, making 64 decisions in a target-practice scenario, and steered a simulated drone through a MuJoCo obstacle course. It ran a live Google Flights search from Zurich to London and verified 21 results, navigated Wikipedia to Gödel's incompleteness theorems, flagged four regression risks in a sample change to authorization code and routed a billing-refund ticket to support.
"Every AI agent is really a long chain of small decisions—which button to press, which file to open, which queue a ticket belongs in," said Josh Liu, AutoTrust AI's chairman and co-founder. "Make each one fast, private and cheap, and you change the economics of the whole chain."
AutoTrust AI said JEV-27B inherits Jev 1.13's blind spots, including multi-hop reasoning, arithmetic, dates and adversarial inputs, and that its training data is English-centric. The company says the model is not meant for high-stakes decisions and recommends gating its answers on confidence.
Benchmark context and source notes

Scores in percent. JEV-27B and TypeSafe Jev 1.13 measured by AutoTrust AI; NeoHorse-Jev, Open-Jev, Kev and Laya English are published baselines from TokenRhythm. Six-benchmark averages: JEV-27B 84.07%, TypeSafe Jev 1.13 83.85%, NeoHorse-Jev 77.70%, Open-Jev 75.67%, Kev 74.25%, Laya English 58.24%.
Scores are in percent. JEV-27B and the hosted TypeSafe Jev 1.13 API were measured by AutoTrust AI; this is not third-party validation. NeoHorse-Jev, Open-Jev, Kev and Laya English are published baselines reproduced from TokenRhythm and were not rerun by AutoTrust. TokenRhythm source (pinned revision): https://huggingface.co/TokenRhythm/NeoHorse-Jev-4B/blob/b50e043e22e0e41e7fc0c244e4daa707b8124930/README.md. Full AutoTrust methodology and source notes: https://huggingface.co/autotrust/JEV-27B.
Availability
JEV-27B is available under the Apache-2.0 license at huggingface.co/autotrust/JEV-27B. The release includes the weights, decision adapter, training and serving code, vLLM support and full evaluation reports. A demonstration reel is available at https://huggingface.co/spaces/autotrust/JEV-27B-Demo.
AutoTrust AI plans to build the JEV decision block into future models in its Guru family, which powers the ScienceGuru research platform. ScienceGuru is available for Windows and macOS at https://scienceguru.ai/. AutoTrust AI also offers customized sovereign deployments for enterprises.
JEV-27B was trained on SargeDev/jev-distill-corpus-v3, a public, Apache-2.0-licensed corpus of Jev 1.13's outputs. It shares no weights or code with, and is not affiliated with or endorsed by, TypeSafe AI. Jev and TypeSafe are trademarks of their respective owners.
About AutoTrust AI
AutoTrust AI Pte. Ltd. is a Singapore-incorporated AI research company building the Guru family of foundation models and ScienceGuru, an AI research platform for scientists and research teams. Its Blocks of Experts architecture combines pretrained expert blocks with small trained adapters to build frontier-capable and sovereign models efficiently. Learn more at autotrust.ai.

Comment :