LLM Benchmark Leaderboard  •  HHRI-AI
FoxBrain EvalHub ↗ GitHub ↗

HHRI-AI LLM EvalBoard

Comprehensive LLM Benchmark Leaderboard — Traditional Chinese, Multilingual & General Capabilities

TMMLU+ TCEval-v2 BigBenchHard τ-bench & τ²-bench MRCR v2 AIEC
🔍

Each axis is a discipline. A model's polygon extends outward as performance improves. Scores are mean accuracy across all subjects in that discipline. Showing top 7 models by current sort.

Horizontal bars show mean accuracy per discipline for 7 models.

Model Name
Overall Acc
TMMLU+ Avg
Rank
Models to Compare