‹ PublicAI Index
The LLM benchmark aggregator.
Llama 3.2 3B Instruct
Meta · 3.2B · open weights
Strongest in Tool use (#71 of 81), weakest in Human preference (#295 of 341). Above par in 3 of 10 scopes. Among the models it meets almost everywhere, it finishes behind Claude Sonnet 5 and Claude Sonnet 4.5 and ahead of DeepSeek R1 and DeepSeek V3.
§ 1 · Profile
What it is good at
Bars run from 50 — the average of the models each source lists — so right of the line is above par. The middle column is the gap to whoever leads that scope.
Agents43.6−24.5#204/2671/5
Tool use38.5−35.5#71/811/1
Safety49−12.5#209/3361/3
Harm refusal53.9−7.7#90/2991/2
Jailbreak resistance55.7−9.2#106/2721/1
Toxicity avoidance52.6−6.4#164/2721/1
Secure code32.6−33.9#241/2741/1
Fairness42.9−27.1#246/2991/2
Human preference34.6−32.9#295/3411/1
Human preference34.6−32.9#295/3411/1
§ 2 · Head to head
What it beats, and what beats it
The same models turn up scope after scope. Counted once: where both were placed, who finished higher. Bars run right for Llama 3.2 3B Instruct, left for the other.
§ 3 · Sources
Where the numbers come from
3 publications, 8 figures. Every one links to the page it was read from.
LMArena Text 1167
BFCL v4 21.95%
Enkrypt · Jailbreak risk 8%Enkrypt · Harmful content risk 15%Enkrypt · CBRN risk 14.3%Enkrypt · Toxicity risk 4.5%Enkrypt · Bias risk 88.9%Enkrypt · Insecure code risk 68.9%
Badge
[](https://publicai.io/model-index/m/llama-3-2-3b-instruct)