Scale Labs

[LEADERBOARDS]

testing the limits of AI.

Benchmarks for frontier, agentic, and safety capabilities

Benchmarks20+

Including benchmarks on agentic coding, frontier reasoning, and safety alignment.

Models evaluated100+

From leading AI labs including OpenAI, Anthropic, Google, Meta, and open-source contributors.

DrugDiscoveryBench

DrugDiscoveryBench: 82 expert-curated tasks evaluating how reliably frontier coding agents perform the computational work of early-stage drug discovery.

1

GPT 6 Astra (mini-SWE-agent) max

NEW

68.70±2.50

2

GPT 6 Astra (Codex) max

NEW

65.40±1.90

3

Muse-Spark 1.3 (mini-SWE-agent) high

NEW

62.20±2.10

View Full Ranking

SWE Atlas - Refactoring

Evaluating an agent's ability to restructure code while preserving behavior.

1

GPT 6 Astra (Codex) xHigh

NEW

59.05±6.43

1

Fable-5.1 (Claude Code) xHigh

NEW

56.67±6.52

1

Fable-5 (Claude Code) xHigh

54.76±6.76

View Full Ranking

SWE Atlas - Test Writing

Evaluating an agent’s ability to write production-grade tests

1

Fable-5.1 (Claude Code) xHigh*

NEW

67.04±5.33

1

Opus 5 (Claude Code) xHigh

62.22±5.58

1

Fable-5 (Claude Code) xHigh*

55.60±5.80

View Full Ranking

SWE Atlas - Codebase QnA

Evaluating deep code comprehension and reasoning

1

Opus 5 (Claude Code) xHigh

63.17±5.01

1

Fable-5.1 (Claude Code) xHigh*

NEW

59.95±4.85

1

GPT 6 Astra (Codex) xHigh*

NEW

59.14±4.88

View Full Ranking

HiL-Bench (Human-in-Loop Benchmark)

Evaluates whether agents recognize information gaps and ask targeted clarifying questions.

1

Claude Fable 5.1

NEW

61.50±6.47

1

Claude Opus 5

57.00±5.48

1

Claude Fable 5

56.33±5.50

View Full Ranking

MCP Atlas

Evaluating real-world tool use through the Model Context Protocol (MCP)

1

Muse Spark 1.1

88.10±1.95

1

Fable 5.1

NEW

87.20±2.05

2

claude-opus-5 (xhigh)

85.80±2.10

View Full Ranking

SWE-Bench Pro (Public Dataset)

Evaluating long-horizon software engineering tasks in public open source repositories

1

Muse Spark 1.1*

NEW

61.50±3.10

1

gpt-5.4 (xHigh)*

59.10±3.56

3

Muse Spark*

55.00±3.60

View Full Ranking

SWE-Bench Pro (Private Dataset)

Evaluating long-horizon software engineering tasks in commercial-grade private repositories

1

Muse Spark 1.1*

NEW

51.50±5.50

1

claude-opus-4-6 (thinking)*

47.10±6.07

3

Muse Spark*

44.70±6.05

View Full Ranking

SciPredict

Forecasting scientific experiment outcomes

1

gemini-3-pro-preview

25.27±1.92

1

claude-opus-4-5-20251101

23.05±0.51

1

claude-opus-4-1-20250805

22.22±1.48

View Full Ranking

Humanity's Last Exam

Challenging LLMs at the frontier of human knowledge

1

54.80±1.94

2

46.50±2.00

2

46.44±1.96

View Full Ranking

Humanity's Last Exam (Text Only)

Challenging LLMs at the frontier of human knowledge

1

54.17±2.09

2

47.31±2.11

2

46.80±2.10

View Full Ranking

AudioMultiChallenge

Evaluating spoken dialogue systems in multi-turn interaction

1

gemini-3.8-flash (high)

NEW

60.40±4.49

1

Inkling (Thinking)*

56.64±4.55

1

Inkling-small

54.87±4.57

View Full Ranking

AudioMultiChallenge - Audio Output

Evaluating spoken dialogue systems in multi-turn interaction

1

gpt-realtime-2 (xHigh)

48.45±4.59

1

tml-interaction-small

43.36±4.55

3

gpt-realtime-2

37.61±4.45

View Full Ranking

AudioMultiChallenge - Text Output

Evaluating spoken dialogue systems in multi-turn interaction

1

Inkling (Thinking)

NEW

56.64±2.70

1

Inkling-small

NEW

54.87±4.57

1

gemini-3-pro-preview (Thinking)

54.65±4.57

View Full Ranking

Professional Reasoning Benchmark - Finance

Evaluating Professional Reasoning in Finance

1

Muse Spark 1.3

NEW

59.54±1.66

2

Muse Spark 1.1

55.01±0.14

3

claude-fable 5

53.86±0.14

View Full Ranking

Professional Reasoning Benchmark - Legal

Evaluating Professional Reasoning in Legal Practice

1

Muse Spark 1.3

NEW

61.56±1.72

2

Muse Spark 1.1

57.05±0.22

3

claude-fable 5

52.56±0.54

View Full Ranking

Remote Labor Index (RLI)

Evaluating AI agents ability to perform real-world, economically valuable remote work

1

GPT 6 Astra

NEW

20.83

2

Fable 5.1

NEW

17.92

3

Fable 5

15.80

View Full Ranking

PropensityBench

Simulating real-world pressure to choose between safe or harmful behavior

1

o3-2025-04-16

10.50±0.60

2

claude-sonnet-4-20250514

12.20±0.20

3

o4-mini-2025-04-16

15.80±0.40

View Full Ranking

VisualToolBench (VTB)

Evaluating how LLMs can dynamically interact with and reason about visual information

1

Muse Spark 1.1

NEW

44.77±2.82

2

gpt-5.4-2026-03-05 (reasoning effort = high)

29.17±0.13

2

gemini-3.1-pro-preview

28.97±0.91

View Full Ranking

MultiNRC

Multilingual Native Reasoning Evaluation Benchmark for LLMs

1

Muse Spark 1.1

NEW

65.59±2.87

1

65.20±1.24

1

64.74±2.88

View Full Ranking

MultiChallenge

Assessing models across diverse, interdisciplinary challenges

1

Muse Spark

75.52±4.05

1

Muse Spark 1.1

NEW

75.30±0.60

3

gemini-3.1-pro-preview

71.37±1.74

View Full Ranking

Fortress

Frontier Risk Evaluation for National Security and Public Safety

1

8.24±1.93

1

9.63±2.11

3

12.40±1.48

View Full Ranking

MASK

Evaluate model honesty when pressured to lie

1

96.28±0.41

1

96.13±0.57

1

Claude Sonnet 4 (Thinking)

95.33±2.29

View Full Ranking

EnigmaEval

Evaluating model performance on complex, multi-step reasoning tasks

1

claude-fable-5-high

39.28±2.80

1

gpt-5.6-sol-high

37.12±2.80

1

gemini-3.1-pro-preview-high

36.78±2.71

View Full Ranking

VISTA

Vision-Language Understanding benchmark for multimodal models

1

Gemini 2.5 Pro Experimental (March 2025)

54.65±1.46

1

gemini-2.5-pro-preview-06-05

54.63±0.55

1

gpt-5.4-pro-2026-03-05

53.89±2.02

View Full Ranking

TutorBench

Evaluating model performance on common tutoring tasks for high school and AP-level subjects

1

Muse Spark

68.55±0.95

1

gpt-5.4-pro-2026-03-05

56.62±1.02

1

gemini-2.5-pro-preview-06-05

55.65±1.11

View Full Ranking

Frontier AI Model Evaluations & Benchmarks

We conduct high-complexity evaluations to expose model failures, prevent benchmark saturation, and push model capabilities--while continuously evaluating the latest frontier models.

Scaling with Human Expertise

Humans design complex evaluations and define precise criteria to assess models, while LLMs scale evaluations--ensuring efficiency and alignment with human judgment.

Robust Datasets for Reliable AI Benchmarks

Our leaderboards are built on carefully curated evaluation sets, combining private datasets to prevent overfitting and open-source datasets for broad benchmarking and comparability.

[EVALUATE YOUR MODEL]

If you'd like to add your model to this leaderboard or a future version, please contact [email protected]. To ensure leaderboard integrity, we require that models can only be featured the FIRST TIME when an organization encounters the prompts.