When a lab announces a model, it shows a table of benchmark scores. Each row is a fixed set of tasks with a fixed way of checking the answers, so a score is only as meaningful as the tasks behind it. This guide explains the benchmarks that appear most often: the six our leaderboard used to show, and the newer ones that have largely replaced them in announcements.
It doesn't list any model's scores. Vendors report different benchmarks, in different versions, under their own settings, and that is why the leaderboard now shows a single measure, LMArena's rating, for every model.
Before comparing two numbers
The same benchmark name can hide very different measurements. Before treating two scores as comparable, check:
- Which version or subset. Several benchmarks below exist in more than one version, and a score on one says little about another.
- The settings. How much reasoning the model was allowed ("high", "max"), whether it could use tools such as web search or code execution, and, for agent benchmarks, the harness: the program that runs the model, gives it tools and decides when it is done. The same model can score very differently in two harnesses.
- Who ran it. Vendor-reported scores use the vendor's own setup. Independent re-runs sometimes come out lower.
- Whether the questions could have leaked into training. A public test set that has been online for years may be partly memorised. Benchmarks that add new problems over time, or keep some hidden, guard against this.
- How close it is to the ceiling. When most strong models score near the maximum, small differences are mostly noise, and the benchmark no longer tells them apart.
- How many questions there are. On a 200-question test, one question is half a percentage point.
Knowledge and reasoning
MMLU-Pro
What it is. About 12,000 multiple-choice questions (12,032 in the released set) across 14 subject areas, from biology and law to engineering and psychology. It was introduced in 2024 as a harder version of MMLU, an older knowledge test that strong models had largely mastered.
How it differs from MMLU. Each question has up to ten options instead of four, so a lucky guess is right one time in ten. Trivial and ambiguous MMLU questions were removed, and new reasoning-heavy questions were added from STEM websites, TheoremQA and SciBench. Its authors found scores varied less with the wording of the prompt than on MMLU, and that working through the problem step by step helps, which suggests more of it needs reasoning rather than recall.
Watch out for. It is still multiple choice, so it measures picking an answer, not producing one.
GPQA Diamond
What it is. GPQA is a set of 448 multiple-choice questions in biology, physics and chemistry, written by people with or working towards PhDs in those fields. GPQA Diamond is its highest-quality subset: 198 questions that the expert validators answered correctly (or where an expert's mistake was clearly identifiable) and that most non-experts got wrong.
Why it's called "Google-proof". In the original study, domain experts reached 65% on the full set (74% after discounting mistakes they recognised afterwards), while skilled non-experts reached only 34%, despite spending over 30 minutes per question with unrestricted web access.
Watch out for. With four options, guessing alone scores about 25%. And with only 198 questions, a difference of one or two points between models can be chance.
Humanity's Last Exam (HLE)
What it is. 2,500 questions written by nearly 1,000 subject experts, across more than a hundred subjects. It was built in 2025 because popular benchmarks had become too easy to separate the strongest models. Each question has one verifiable answer that can't be found quickly by searching the internet.
How it's scored. About 24% of questions are multiple choice, and the rest need an exact short answer; around 14% need an image as well as text. A language model acting as a judge compares each answer with the correct one. Models also state how confident they are, so HLE reports calibration too: whether a model that says "90% sure" is right about 90% of the time. The organisers keep a private set of held-out questions to check for overfitting.
Watch out for. Scores "with tools" (search, code execution) and "without tools" are very different measurements, and some vendors report a text-only subset. The authors' own audit estimated that experts disagree with the official answer on about 15% of the public questions, so a perfect score isn't a meaningful target.
Maths
MATH-500
What it is. 500 competition maths problems taken from MATH, a 2021 dataset of 12,500 problems from high-school maths competitions such as the AMC 10, AMC 12 and AIME, each with a worked solution. OpenAI picked this subset as a test set in 2023, in the paper Let's Verify Step by Step, and it became a standard yardstick.
How it's scored. The model's final answer is checked against the known answer.
Watch out for. The problems have been public since 2021, and frontier models now score close to the maximum, so it no longer separates them. Announcements have moved to harder or newer maths tests.
Coding
LiveCodeBench
What it is. Competitive-programming problems collected continuously from LeetCode, AtCoder and Codeforces since May 2023. Every problem is tagged with the date it was published, so a model can be tested only on problems that appeared after its training data was collected. That is the benchmark's answer to leaked test questions.
How it's scored. The generated program must pass the problem's hidden tests, usually on the first attempt. The benchmark also has tasks for repairing code, predicting a program's output and predicting test results, but headline numbers are usually code generation.
Watch out for. Each release covers a different window of problems: the first had 400 (May 2023 to March 2024), the fifth 880 (May 2023 to January 2025). Scores on different windows aren't comparable. And contest puzzles are a different skill from working in a large codebase.
SWE-bench Verified
What it is. SWE-bench, introduced in 2023, has 2,294 tasks taken from real GitHub issues and the pull requests that fixed them, in 12 popular Python projects. The model gets the project's code and the issue text, and must edit the code to fix the issue. SWE-bench Verified is a 500-task subset that people checked by hand, so that every issue is clearly described and its tests are fair.
How it's scored. The project's own tests are run on the edited code: the tests that the real fix made pass must now pass, and the ones that already passed must still pass.
Watch out for. It covers only Python and only 12 projects, and the tasks have been public since 2023. Results depend heavily on the agent harness, so two scores for the same model can differ by several points.
SWE-bench Pro
What it is. A harder successor from Scale AI (2025): 1,865 tasks from 41 actively maintained projects, such as business applications and developer tools. They are long tasks that could take a professional engineer hours to days, often changing several files. The tasks are split into a public set (11 projects), a held-out set (12) and a commercial set from 18 private codebases of partner companies. The last two aren't published, which makes them hard to train on.
How it's scored. As with SWE-bench, hidden tests check that the change does what was asked without breaking anything else, run in a clean container.
Watch out for. The public set has changed. It started with 731 tasks; a second version (September 2026) dropped 89 tasks found to be invalid, leaving 642, and there is also a 51-task "hard" subset. Check which one a score refers to. A 2026 study also found ways for agents to game the original version, such as finding the reference solution, which can inflate scores; it proposed a cleaned-up "SWE-Bench Pro Verified".
Terminal-Bench
What it is. Hard tasks an agent completes in a computer terminal, inspired by problems from real workflows. In Terminal-Bench 2.0 (89 tasks), each task has its own environment, a reference solution written by a person, and tests that check the final result.
How it's scored. The share of tasks whose checks pass at the end of the agent's run.
Watch out for. Versions 2.0, 2.1, 3.0 and 4.0 are all quoted in recent announcements, and they are different task sets, so the version matters as much as the number. Scores depend a lot on the harness around the model. Research in 2026 found that some tasks could be passed by getting around the checks rather than doing the work, and that some tasks no model solved were broken rather than hard.
Images and diagrams
MMMU
What it is. About 11,500 questions from college exams, quizzes and textbooks that need an image to answer: charts, diagrams, maps, tables, sheet music, chemical structures and more. They cover six broad disciplines (art and design, business, science, health and medicine, humanities and social science, tech and engineering), split into 30 subjects.
How it's scored. Accuracy against the known answers.
Watch out for. Some questions can be answered from the text alone. MMMU-Pro (2024) removes those, adds more answer options, and adds a setting where the whole question is shown as an image. Recent announcements mostly report MMMU-Pro, and its scores can't be compared with MMMU scores.
How this differs from the Arena rating
LMArena's rating, the one our leaderboard shows, works differently from all of the above. People type in their own prompt, see answers from two anonymous models side by side, and vote for the better one. A statistical model turns these head-to-head votes, crowdsourced from many users, into one rating per model.
That makes it a measure of which answers people prefer, on the questions they choose to ask, rather than of correctness on a fixed test. There is no fixed answer key to memorise, and it covers everyday use, but style can sway votes too: length, formatting and tone, not just correctness. Read it alongside benchmark results, not instead of them.
Sources
- MMLU-Pro: Wang et al., 2024; dataset
- GPQA: Rein et al., 2023
- Humanity's Last Exam: Phan et al., 2025; lastexam.ai
- MATH: Hendrycks et al., 2021; MATH-500 from Lightman et al., 2023
- LiveCodeBench: Jain et al., 2024; dataset versions
- SWE-bench: Jimenez et al., 2023; SWE-bench Verified
- SWE-bench Pro: Scale AI, 2025; dataset and versions; SWE-Bench Pro Verified, 2026
- Terminal-Bench: Terminal-Bench 2.0, 2026; tbench.ai; on gamed checks, Hack-Verifiable Terminal Bench, 2026; on broken tasks, What Makes a Terminal-Bench Task Hard?, 2026
- MMMU: Yue et al., 2023; MMMU-Pro: Yue et al., 2024
- Chatbot Arena (LMArena): Chiang et al., 2024