o11y-bench

The first observability benchmark for AI agents

A standardized evaluation suite measuring how well AI agents perform 63 real-world observability tasks across logs, metrics, traces, dashboards, and incident workflows.

Top Agents

View all →

#1

Base Model

Anthropic logo

claude-opus-4-8

79.4%

Pass^3

Thinking

High

Pass@3

88.9%

Tasks

50/63

Date

2026-08-21

Total Cost

$82.35

Avg Cost

$0.436

Avg Tokens

182k

#2

Base Model

Anthropic logo

claude-opus-4-7

79.4%

Pass^3

Thinking

Off

Pass@3

87.3%

Tasks

50/63

Date

2026-04-20

Total Cost

$60.39

Avg Cost

$0.320

Avg Tokens

157k

#3

Base Model

Anthropic logo

claude-opus-5

76.2%

Pass^3

Thinking

High

Pass@3

93.7%

Tasks

48/63

Date

2026-08-21

Total Cost

$152.28

Avg Cost

$0.806

Avg Tokens

289k

#4

Base Model

Anthropic logo

claude-opus-4-7

73.0%

Pass^3

Thinking

High

Pass@3

90.5%

Tasks

46/63

Date

2026-04-21

Total Cost

$74.34

Avg Cost

$0.393

Avg Tokens

170k

#5

Base Model

Anthropic logo

claude-opus-5

68.3%

Pass^3

Thinking

Low

Pass@3

92.1%

Tasks

43/63

Date

2026-08-21

Total Cost

$47.86

Avg Cost

$0.253

Avg Tokens

140k

Top 10 By Category

Category scores use Pass^3 consistency across the three benchmark trials per task. Green is 90%+, yellow is 70%+, and red is below 70%.

Swipe horizontally to compare category scores.

Model DashboardsGrafana APIInvestigationLogsMetricsTraces
claude-opus-4-8 57% 100% 82% 30% 100% 92%
claude-opus-4-7 57% 100% 73% 80% 88% 77%
claude-opus-5 71% 100% 36% 80% 81% 92%
claude-opus-4-7 43% 100% 45% 80% 88% 77%
claude-opus-5 57% 100% 64% 10% 81% 92%
claude-sonnet-4-6 29% 100% 45% 50% 94% 77%
gpt-5.6-sol 29% 100% 45% 60% 81% 77%
claude-opus-4-6 43% 100% 45% 60% 75% 77%
gpt-5.6-luna 43% 100% 55% 30% 88% 77%
gpt-5.6-terra 43% 100% 55% 30% 75% 85%

Featured Tasks

Browse all →