Decisions is live on the PublicAI Index. It ranks 85 models on the small calls an application makes all day, it sits beside Coding, Reasoning and the rest, and it has two domains: Routing & classification and Calibration.

Three rows carry an Overall chip. Those three are the only systems here that any other board on the Index has measured.
The calls nobody logs
Which model takes this request. What label this ticket gets. Whether a human should see it. These run on every request that arrives, they are answered by something small and cheap, and when they are wrong nothing says so: the request goes to the weaker model, the answer comes back plausible, and no line in the log claims a better one would have done better. A wrong essay is visible. A wrong route is not.
Which is why the category has two domains and not one. Whether a system is right is Routing & classification. Whether it knows when it is not is Calibration, and that is the number that decides whether a fallback ever fires. They come apart further than you would expect. BAAI's bge-reranker-v2-m3 scores 5.0 on the first, near the bottom of the board, and 84.2 on the second, above most models that beat it everywhere else. It is almost never the one to ask, and it rarely pretends otherwise. DeepSeek V4.1 Flash leads Calibration at 95.5, ahead of GPT-6 Luna's 93.5, having lost to it on accuracy.
Eighty-two of the eighty-five are on no other board
That is the reason this column exists. These systems are small, specialised, and invisible to the general leaderboards, so until now a reader comparing them had nothing here to read. Twenty-six of them have Jev in the name: Open-Jev, SimpleJev, LitJev, OpenSourceJev, Open Alternative Jev, JevAct, JevOne, Djev. One product named a category, and the category filled with systems built to replace it.
The bar for adding a category has not moved: a recognised publisher has to put a table on a public page we can read and link back to. JevBench, from Benchmark Heaven, did. It is one board, plus one report ✱, and we would rather say that than let a badge strip imply more. So Decisions never enters the Overall index and never anchors an estimate; a model measured only here gets a column and no rank. A second publisher with a table would change that.
The order today
| # | model | Decisions |
|---|---|---|
| 1 | DeepSeek V4.1 Flash | 73.3 |
| 2 | GPT-6 Luna | 72.7 |
| 3 | GPT-5.6 Luna | 70.6 |
| 4 | Djev | 70.0 |
| 5 | Gemma 4 31B IT (Autoloops) | 62.7 |
| 6 | Instinct | 59.3 |
| 7 | Jev 1.13.0 | 59.3 |
General-purpose models hold the top three, and JevBench's own ranking does not, which is worth knowing before reading either. Its headline is four axes: Intelligence, Calibration, Speed and Cost. GPT-6 Luna has the highest Intelligence on the board at 97.4 and its two entries place 31st and 35th there, because by the board's measure it is slow and expensive. For a team paying per decision that may be the only thing that matters. It is not a capability question, so this column takes the two axes that describe the model and leaves price and latency to you, which is also why a large API model can lead a column about what you would run instead of one.
The half you cannot see
The board keeps a control, which is rare and to its credit: 534 public decisions, 308 sealed ones, and a column for the gap.
| board rank | public (n=534) | sealed (n=308) | gap | |
|---|---|---|---|---|
| decider-4b v2 | 1 | 83.5% | 34.7% | +48.8 pp |
| Jev 1.13.0 | 2 | 86.6% | 36.7% | +49.9 pp |
| GPT-6 Luna | 35 | 99.6% | 95.5% | +4.1 pp |
Thirty-five entries score 80% or better on the public half. Thirty of them fall under 50% on the sealed one. Sort by the sealed column and the top four are general-purpose models, with the best specialist twenty-nine points behind the fourth.
A gap is not proof of anything by itself: a sealed set can be harder, and a small model tuned for one shape of task can be brittle rather than fitted. But the two figures agree only loosely, so if you are choosing one of these systems, the sealed number is the one to ask for. It is the board's own column, not our inference.
LargitData's write-up of the same question is indexed here too, marked ✱, at a tenth of a board's weight, as every report is. And one more piece of credit: Benchmark Heaven runs a router of its own, and excludes it from its own ranking.
Every figure above links back to the page it was read from. The same rows are available as JSON, over MCP, and as an RSS feed.