Benchmark performance
Compare measured intelligence, cost, speed, and latency. Check the model variant and the units shown in each column.
Explore Artificial Analysis in a new tabTHE MODEL OBSERVATORY
Understand the scores. Compare the trade-offs.
Independent benchmarks, model specifications, and a practical cost comparison. Build your shortlist with the evidence in view.
Use the table’s filters and column headings to explore the results.
Opening the source view…
View blank or blocked? Open the leaderboard directly in a new tab.
Current source. Clear context. The time above records when the view opened, not when a model was evaluated. New results appear when the source publishes them. Refresh pauses in other views, offline, and while you interact inside the table. Turn it off to keep filters between refreshes; your preference is saved on this browser.
Check capabilities and provider pricing before you choose.
Opening model explorer…
CHOOSE YOUR EVIDENCE
Compare the task, model version, and test setup. These sources measure different things and should not be blended into one score.
Compare measured intelligence, cost, speed, and latency. Check the model variant and the units shown in each column.
Explore Artificial Analysis in a new tabExplore which responses people prefer in comparisons. Preferences can vary by task, language, and category.
Explore Arena in a new tabSee which models people use on OpenRouter. Popularity measures adoption; it is a different signal from benchmark performance.
Explore OpenRouter in a new tabExplore results on repository issues. Compare the benchmark split and agent setup as well as the underlying model; results measure the full system.
Explore SWE-bench in a new tabTHE DETAILS THAT MATTER