Hacker News Viewer

Show HN: JevBench, a reproducible benchmark for typed decision models

by florianstandhar on 9/22/2026, 1:01:03 PM

Hi HN! I built JevBench because Jev kicks ass, and the world deserves to know how the serious open source and fake lookalike projects <i>really</i> perform in comparison.<p>Jev-class models return bounded choices and probabilities instead of text, and are disruptively faster and cheaper than LLMs, while being similarly intelligent on the text input they operate on.<p>JevBench allows looking at accuracy, latency and price all at once, in a weighted way - you can even configure the weighting.<p>A full run asks 534 English decisions. The v1.3 score combines chance-corrected Intelligence, Calibration, Speed and Cost.<p>Leaderboard right now:<p><pre><code> #1 - Jev 74.4 #2 - SemIf 73.1 #3 - djev 73.0 #4 - Winnow-12B Q8 71.2 #5 reflex 4B 70.3. </code></pre> MIT harness, public items, frozen artifacts, scoring code and public per-task outcomes:<p><a href="https:&#x2F;&#x2F;github.com&#x2F;fstandhartinger&#x2F;jevbench" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;fstandhartinger&#x2F;jevbench</a><p>Two no-signup demos:<p><a href="https:&#x2F;&#x2F;who-is-right.app.mintapis.com" rel="nofollow">https:&#x2F;&#x2F;who-is-right.app.mintapis.com</a><p><a href="https:&#x2F;&#x2F;is-it-ai-slop.app.mintapis.com" rel="nofollow">https:&#x2F;&#x2F;is-it-ai-slop.app.mintapis.com</a><p>Limitations: English-only; latency from one German server; local&#x2F;demo latency gets a disclosed ×2 adjustment (+150 ms on my servers) which is an informed assumption; held-out prompts still reach evaluated services; ~1-point gaps can be noise.<p>Wdyt?

https://benchmarkheaven.com/jev-models

Comments

by: sean_pedersen

Good project but this one also exists <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;spaces&#x2F;multimodalart&#x2F;jev-decision-index" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;spaces&#x2F;multimodalart&#x2F;jev-decision-ind...</a> and the results do not seem to add up and also model sets are different... still needs time to mature likely

9/22/2026, 9:37:13 PM


by: swyx

jev ceo on why he eschewed benchmarking: <a href="https:&#x2F;&#x2F;www.latent.space&#x2F;i&#x2F;216783460&#x2F;privacy-benchmarking-and-trusting-intelligence" rel="nofollow">https:&#x2F;&#x2F;www.latent.space&#x2F;i&#x2F;216783460&#x2F;privacy-benchmarking-an...</a>

9/22/2026, 8:42:42 PM


by:

9/22/2026, 1:01:03 PM