‹ PublicAI Index
The LLM benchmark aggregator.
Zephyr 7B Beta
Hugging Face · 7B
Strongest in Fairness (#177 of 299, on 1 of its 2 boards), weakest in Human preference (#314 of 341). Above par in 1 of 8 scopes. Among the models it meets almost everywhere, it finishes behind Claude Opus 5.5 and Claude Opus 4.7 and ahead of Command R and Grok 3.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Safety45.5−16#283/3361/3
Fairness45.7−24.3#177/2991/2
Toxicity avoidance50.3−8.7#197/2721/1
Jailbreak resistance46.2−18.7#204/2721/1
Secure code42.4−24.1#205/2741/1
Harm refusal43.8−17.8#260/2991/2
Human preference31.1−36.4#314/3411/1
Human preference31.1−36.4#314/3411/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Zephyr 7B Beta, left for the other.
§ 3 · Sources
Where the numbers come from
2 publications, 7 figures. Every one links to the page it was read from.
LMArena Text 1131
Enkrypt · Jailbreak risk 16%Enkrypt · Harmful content risk 80%Enkrypt · CBRN risk 8.8%Enkrypt · Toxicity risk 6.1%Enkrypt · Bias risk 84.5%Enkrypt · Insecure code risk 48.9%
Badge
[](https://publicai.io/model-index/m/zephyr-7b-beta)