#1
Base Model
claude-opus-4-8
79.4%
Pass^3
Thinking
High
Pass@3
88.9%
Tasks
50/63
Date
2026-08-21
Total Cost
$82.35
Avg Cost
$0.436
Avg Tokens
182k
o11y-bench
A standardized evaluation suite measuring how well AI agents perform 63 real-world observability tasks across logs, metrics, traces, dashboards, and incident workflows.
#1
Base Model
79.4%
Pass^3
Thinking
High
Pass@3
88.9%
Tasks
50/63
Date
2026-08-21
Total Cost
$82.35
Avg Cost
$0.436
Avg Tokens
182k
#2
Base Model
79.4%
Pass^3
Thinking
Off
Pass@3
87.3%
Tasks
50/63
Date
2026-04-20
Total Cost
$60.39
Avg Cost
$0.320
Avg Tokens
157k
#3
Base Model
76.2%
Pass^3
Thinking
High
Pass@3
93.7%
Tasks
48/63
Date
2026-08-21
Total Cost
$152.28
Avg Cost
$0.806
Avg Tokens
289k
#4
Base Model
73.0%
Pass^3
Thinking
High
Pass@3
90.5%
Tasks
46/63
Date
2026-04-21
Total Cost
$74.34
Avg Cost
$0.393
Avg Tokens
170k
#5
Base Model
68.3%
Pass^3
Thinking
Low
Pass@3
92.1%
Tasks
43/63
Date
2026-08-21
Total Cost
$47.86
Avg Cost
$0.253
Avg Tokens
140k
| # | Agent | Model | Thinking | Tasks | Date | |||||
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Base Model | | High | 79.4% | 88.9% | 50/63 | $82.35 | $0.436 | 182k | 2026-08-21 |
| 2 | Base Model | | Off | 79.4% | 87.3% | 50/63 | $60.39 | $0.320 | 157k | 2026-04-20 |
| 3 | Base Model | | High | 76.2% | 93.7% | 48/63 | $152.28 | $0.806 | 289k | 2026-08-21 |
| 4 | Base Model | | High | 73.0% | 90.5% | 46/63 | $74.34 | $0.393 | 170k | 2026-04-21 |
| 5 | Base Model | | Low | 68.3% | 92.1% | 43/63 | $47.86 | $0.253 | 140k | 2026-08-21 |
Category scores use Pass^3 consistency across the three benchmark trials per task. Green is 90%+, yellow is 70%+, and red is below 70%.
Swipe horizontally to compare category scores.
| Model | Dashboards | Grafana API | Investigation | Logs | Metrics | Traces |
|---|---|---|---|---|---|---|
| claude-opus-4-8 | 57% | 100% | 82% | 30% | 100% | 92% |
| claude-opus-4-7 | 57% | 100% | 73% | 80% | 88% | 77% |
| claude-opus-5 | 71% | 100% | 36% | 80% | 81% | 92% |
| claude-opus-4-7 | 43% | 100% | 45% | 80% | 88% | 77% |
| claude-opus-5 | 57% | 100% | 64% | 10% | 81% | 92% |
| claude-sonnet-4-6 | 29% | 100% | 45% | 50% | 94% | 77% |
| gpt-5.6-sol | 29% | 100% | 45% | 60% | 81% | 77% |
| claude-opus-4-6 | 43% | 100% | 45% | 60% | 75% | 77% |
| gpt-5.6-luna | 43% | 100% | 55% | 30% | 88% | 77% |
| gpt-5.6-terra | 43% | 100% | 55% | 30% | 75% | 85% |