‹ PublicAI Index
The LLM benchmark aggregator.
Gemma 4 31B
Google · 31B · open weights
Strongest in Decisions (#5 of 85), weakest in Fairness (#269 of 299). Above par in 13 of 29 scopes. Among the models it meets almost everywhere, it finishes behind Claude Opus 5 and Claude Opus 5.5 and ahead of MiniMax M3 and GPT-4.1 Mini.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Decisions62.7−10.6#5/851/1
Routing & classification63.7−10.3#5/851/1
Calibration61.7−11#8/851/1
Core abilities53.7−13.7#56/2041/3
General intelligence55.3−14.8#58/2041/3
Long context49.8−14.8—/0✱✱
Human preference62.1−5.4#60/3411/1
Human preference62.1−5.4#60/3411/1
Professional47.8−13.8#106/1681/1
Finance50.6−12.3#85/1511/1
Legal43.4−23.7#115/1511/1
Agents44.6−23.5#187/2672/5
Knowledge work43−30.6#115/1772/2
Tool use45.8−28.2—/81✱0/1
Safety49.7−11.8#188/3362/3
Factual grounding56.5−14.1#34/1011/1
Secure code60.6−5.9#62/2741/1
Toxicity avoidance52.7−6.3#160/2721/1
Harm refusal48.5−13.1#179/2991/2
Jailbreak resistance42.2−22.7#212/2721/1
Fairness41.2−28.8#269/2991/2
Reasoning49.4−17.5—/178✱0/4
Science49.7−13.8—/122✱0/1
Expert reasoning46−14.9—/0✱✱
Coding50.8−18.3—/164✱0/5
Agentic coding50−18.7—/157✱0/4
Code generation54.2−15.8—/77✱0/2
Knowledge44.1−19.4—/138✱0/2
Factuality41.2−17.8—/0✱✱
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Gemma 4 31B, left for the other.
§ 3 · Sources
Where the numbers come from
10 publications, 25 figures. Every one links to the page it was read from.
LMArena Text 1453
GDPval-AA 606
AA-Briefcase 367
Kagi LLM Benchmark 63.5%
Vals · CaseLaw 52.63%Vals · MortgageTax 61.37%
JevBench · Intelligence 59.8%JevBench · Calibration 79.2%
Enkrypt · Jailbreak risk 19.4%Enkrypt · Harmful content risk 6.1%Enkrypt · CBRN risk 33.3%Enkrypt · Toxicity risk 4.4%Enkrypt · Bias risk 91.5%Enkrypt · Insecure code risk 12%
Vectara · Factual consistency 92.6%
tau3-Banking 14.8%Terminal-Bench 2.1 43.4%SciCode 43.4%Humanity's Last Exam (without tools) 23.6%GPQA Diamond 85.7%CritPt 1.4%AA-LCR 68.3%AA-Omniscience Accuracy 20%AA-Omniscience Non-Hallucination 15%
Decision correct 77%
Badge
[](https://publicai.io/model-index/m/gemma-4-31b)