Benchmarks

Ranked on real usage from Trooper users, not synthetic suites: tasks their agents actually ran, what got done, what it cost, and how long it took.

Last run 2026-08-07

42,107 tasks · $142,942 measured spend · 2026-05-062026-08-06

6 of 10 agents have statistically significant coverage (≥500 tasks). Coverage expands as sessions grow.

Cheapest pair on the board · vs Hermes, same model

$0.11 / task·1.3× cheaper

Trooper × DeepSeek V4 Flash is the lowest typical cost on this snapshot — $0.11 vs $0.14 for Hermes on the same weights, at a higher Elo (1,096 vs 1,068).

Trooper × ChatGPT Luna · vs Codex, same window

$1.42 / task·1.3× cheaper

Luna on Trooper lands at $1.42 typical versus Codex × GPT-5.6 Sol at $1.88 — same production traffic, less spend per task.

Overall rankings

All task categories combined. Filter by harness, click a column to re-sort. Only same-harness rows are a like-for-like runtime comparison; swapping the model changes the price of a token.

Sorted by $ / task

#HarnessModel
1TrooperDeepSeek V4 Flash1,096↑1245.4%$0.11
2HermesDeepSeek V4 Flashprovisional1,068↓641.7%$0.14
3HermesDeepSeek V4 Pro1,104↓2146.8%$0.53
4TrooperChatGPT Luna1,142↑851.8%$1.42
5Claude CodeClaude Sonnet 4.61,090↑344.8%$1.52
6CodexGPT-5.6 Sol1,185↓1558.4%$1.88
7Claude CodeClaude Opus 4.81,130↑1850.5%$2.20
8HermesKimi K31,063↓641%$2.26
9Claude CodeClaude Opus 51,170↓2156.2%$3.63
10Claude CodeClaude Fable 51,20160.6%$4.86

◆ on the cost-efficiency frontier — no other pair is both cheaper and higher-Elo. Default sort is $ / task so Trooper’s cheaper pairs sit at the top.

Harness rankings

Volume-weighted Elo across every model a harness ran. Use this when you are choosing the runtime, not the weights.

#HarnessEloWin rate$ / taskTasks
1CodexBest model GPT-5.6 Sol · 1 models1,18558.4%$1.882,188
2Claude CodeBest model Claude Fable 5 · 4 models1,17957.5%$4.086,944
3TrooperBest model ChatGPT Luna · 2 models1,12749.7%$1.001,272
4HermesBest model DeepSeek V4 Pro · 3 models1,08844.6%$1.031,021

Model rankings

Each model at the harness it actually ran under. Same Elo as the overall board, grouped so you can scan weights first.

#ModelHarnessElo$ / taskTasks
1Claude Fable 5Claude Code1,201$4.864,240
2GPT-5.6 SolCodex1,185$1.882,188
3Claude Opus 5Claude Code1,170$3.631,458
4ChatGPT LunaTrooper1,142$1.42860
5Claude Opus 4.8Claude Code1,130$2.20754
6DeepSeek V4 ProHermes1,104$0.53615
7DeepSeek V4 FlashTrooper1,096$0.11412
8Claude Sonnet 4.6Claude Code1,090$1.52492
9DeepSeek V4 FlashHermes1,068$0.1490
10Kimi K3Hermes1,063$2.26316

Value per dollar

Elo against what a typical task costs. Up and to the left is the sweet spot.

Elo vs typical task cost

Every pair across all categories — pairs on the line are the efficient frontier.

Quality — Elo ↑

1,0251,0751,1251,1751,225
$0.10$1.00

Median cost per task — log scale

  • Trooper
  • Claude Code
  • Codex
  • Hermes

Coding

Software tasks — features, fixes, deploys3,029 tasks.

Elo rankings

Head-to-head rating on Coding work — longer bar is better.

Claude · Fable 5
1,208
Codex · GPT 5.6 Sol
1,186
Claude · Opus 5
1,181
Trooper · Luna
1,148
Claude · Opus 4.8
1,134
Claude · Sonnet 4.6
1,091
Trooper · DS Flash
1,082
Hermes · DS Pro
1,072
Hermes · Kimi K3
1,050
  • Trooper
  • Claude Code
  • Codex
  • Hermes

Elo vs cost

Up and to the left wins more for less.

Quality — Elo ↑

1,0251,0751,1251,1751,225
$0.10$1.00$10

Median cost per task — log scale

  • Trooper
  • Claude Code
  • Codex
  • Hermes

Cost per task

Light, typical, and heavy task cost for each pair.

  • light
  • typical
  • heavy
  1. 1Claude Code·Fable 5

    Elo 1,208 · 60.8% · 8m 38s · 1,152 tasks

    $2.87$7.91$23
  2. 2Codex·GPT 5.6 Sol

    Elo 1,186 · 57.7% · 6m 48s · 595 tasks

    $2.91$9.03
  3. 3Claude Code·Opus 5

    Elo 1,181 · 57% · 9m 58s · 475 tasks

    $2.35$5.87$19
  4. 4Trooper·Luna

    Elo 1,148 · 52.4% · 6m 10s · 96 tasks · provisional

    $2.05$6.40
  5. 5Claude Code·Opus 4.8

    Elo 1,134 · 50.3% · 7m 21s · 216 tasks · provisional

    $3.53$12
  6. 6Claude Code·Sonnet 4.6

    Elo 1,091 · 44.2% · 6m 45s · 210 tasks · provisional

    $2.23$7.54
  7. 7Trooper·DS Flash

    Elo 1,082 · 42.8% · 3m 12s · 38 tasks · provisional

    $0.18
  8. 8Hermes·DS Pro

    Elo 1,072 · 41.5% · 7m 19s · 119 tasks · provisional

    $0.87
  9. 9Hermes·Kimi K3

    Elo 1,050 · 38.5% · 11m 52s · 128 tasks · provisional

    $3.13$13

Research

Deep dives, competitive analysis, reports1,451 tasks.

Elo rankings

Head-to-head rating on Research work — longer bar is better.

Claude · Fable 5
1,196
Codex · GPT 5.6 Sol
1,186
Claude · Opus 5
1,166
Trooper · Luna
1,136
Hermes · DS Pro
1,119
Claude · Opus 4.8
1,119
Hermes · Kimi K3
1,081
Claude · Sonnet 4.6
1,079
Trooper · DS Flash
1,074
  • Trooper
  • Claude Code
  • Codex
  • Hermes

Elo vs cost

Up and to the left wins more for less.

Quality — Elo ↑

1,0251,0751,1251,1751,225
$0.10$1.00

Median cost per task — log scale

  • Trooper
  • Claude Code
  • Codex
  • Hermes

Cost per task

Light, typical, and heavy task cost for each pair.

  • light
  • typical
  • heavy
  1. 1Claude Code·Fable 5

    Elo 1,196 · 58.7% · 9m 46s · 502 tasks

    $1.95$5.22$16
  2. 2Codex·GPT 5.6 Sol

    Elo 1,186 · 57.3% · 9m 26s · 388 tasks

    $2.26$6.89
  3. 3Claude Code·Opus 5

    Elo 1,166 · 54.4% · 10m 38s · 185 tasks · provisional

    $3.69$14
  4. 4Trooper·Luna

    Elo 1,136 · 50.2% · 8m 4s · 48 tasks · provisional

    $1.68$5.40
  5. 5Hermes·DS Pro

    Elo 1,119 · 47.7% · 9m 17s · 131 tasks · provisional

    $0.63
  6. 6Claude Code·Opus 4.8

    Elo 1,119 · 47.7% · 8m 23s · 72 tasks · provisional

    $2.35$7.07
  7. 7Hermes·Kimi K3

    Elo 1,081 · 42.3% · 12m 37s · 66 tasks · provisional

    $1.96$6.72
  8. 8Claude Code·Sonnet 4.6

    Elo 1,079 · 42% · 6m 52s · 37 tasks · provisional

    $1.35$4.46
  9. 9Trooper·DS Flash

    Elo 1,074 · 41.6% · 4m 8s · 22 tasks · provisional

    $0.16

Customer support

Inbound questions answered end to end914 tasks.

Elo rankings

Head-to-head rating on Customer support work — longer bar is better.

Claude · Fable 5
1,185
Codex · GPT 5.6 Sol
1,177
Claude · Opus 5
1,154
Trooper · Luna
1,140
Claude · Opus 4.8
1,123
Hermes · DS Pro
1,113
Claude · Sonnet 4.6
1,102
Trooper · DS Flash
1,078
Hermes · DS Flash
1,058
  • Trooper
  • Claude Code
  • Codex
  • Hermes

Elo vs cost

Up and to the left wins more for less.

Quality — Elo ↑

1,0251,0751,1251,1751,225
$0.10$1.00

Median cost per task — log scale

  • Trooper
  • Claude Code
  • Codex
  • Hermes

Cost per task

Light, typical, and heavy task cost for each pair.

  • light
  • typical
  • heavy
  1. 1Claude Code·Fable 5

    Elo 1,185 · 57.8% · 4m 32s · 342 tasks

    $1.19$2.64$9.70
  2. 2Codex·GPT 5.6 Sol

    Elo 1,177 · 56.7% · 3m 36s · 179 tasks · provisional

    $0.98$3.71
  3. 3Claude Code·Opus 5

    Elo 1,154 · 53.4% · 4m 36s · 109 tasks · provisional

    $1.77$5.13
  4. 4Trooper·Luna

    Elo 1,140 · 51.1% · 3m 28s · 41 tasks · provisional

    $0.92$2.90
  5. 5Claude Code·Opus 4.8

    Elo 1,123 · 49% · 4m 10s · 57 tasks · provisional

    $1.25$4.70
  6. 6Hermes·DS Pro

    Elo 1,113 · 47.5% · 4m 19s · 89 tasks · provisional

    $0.32
  7. 7Claude Code·Sonnet 4.6

    Elo 1,102 · 45.9% · 3m 43s · 45 tasks · provisional

    $0.77$2.18
  8. 8Trooper·DS Flash

    Elo 1,078 · 42.1% · 2m 8s · 28 tasks · provisional

    $0.12
  9. 9Hermes·DS Flash

    Elo 1,058 · 39.7% · 2m 36s · 24 tasks · provisional

    $0.15

Marketing

Campaigns, positioning, launch plans838 tasks.

Elo rankings

Head-to-head rating on Marketing work — longer bar is better.

Claude · Fable 5
1,205
Codex · GPT 5.6 Sol
1,182
Claude · Opus 5
1,174
Trooper · Luna
1,138
Claude · Opus 4.8
1,124
Hermes · DS Pro
1,110
Claude · Sonnet 4.6
1,087
Trooper · DS Flash
1,076
Hermes · Kimi K3
1,071
  • Trooper
  • Claude Code
  • Codex
  • Hermes

Elo vs cost

Up and to the left wins more for less.

Quality — Elo ↑

1,0251,0751,1251,1751,225
$0.10$1.00

Median cost per task — log scale

  • Trooper
  • Claude Code
  • Codex
  • Hermes

Cost per task

Light, typical, and heavy task cost for each pair.

  • light
  • typical
  • heavy
  1. 1Claude Code·Fable 5

    Elo 1,205 · 59.8% · 7m 33s · 322 tasks

    $1.78$4.30$14
  2. 2Codex·GPT 5.6 Sol

    Elo 1,182 · 56.6% · 5m 48s · 157 tasks · provisional

    $1.55$4.57
  3. 3Claude Code·Opus 5

    Elo 1,174 · 55.4% · 8m 24s · 120 tasks · provisional

    $3.10$11
  4. 4Trooper·Luna

    Elo 1,138 · 50.6% · 5m 22s · 42 tasks · provisional

    $1.18$3.70
  5. 5Claude Code·Opus 4.8

    Elo 1,124 · 48.3% · 6m 21s · 59 tasks · provisional

    $1.90$6.63
  6. 6Hermes·DS Pro

    Elo 1,110 · 46.2% · 6m 47s · 55 tasks · provisional

    $0.50
  7. 7Claude Code·Sonnet 4.6

    Elo 1,087 · 43% · 5m 29s · 33 tasks · provisional

    $1.14$4.49
  8. 8Trooper·DS Flash

    Elo 1,076 · 41.4% · 3m 2s · 24 tasks · provisional

    $0.14
  9. 9Hermes·Kimi K3

    Elo 1,071 · 40.7% · 8m 56s · 26 tasks · provisional

    $1.51$5.67

Sales

Prospecting, outreach, and pipeline upkeep669 tasks.

Elo rankings

Head-to-head rating on Sales work — longer bar is better.

Claude · Fable 5
1,209
Codex · GPT 5.6 Sol
1,190
Claude · Opus 5
1,170
Trooper · Luna
1,144
Claude · Opus 4.8
1,131
Hermes · DS Pro
1,099
Claude · Sonnet 4.6
1,083
Trooper · DS Flash
1,070
Hermes · Kimi K3
1,059
  • Trooper
  • Claude Code
  • Codex
  • Hermes

Elo vs cost

Up and to the left wins more for less.

Quality — Elo ↑

1,0251,0751,1251,1751,225
$0.10$1.00

Median cost per task — log scale

  • Trooper
  • Claude Code
  • Codex
  • Hermes

Cost per task

Light, typical, and heavy task cost for each pair.

  • light
  • typical
  • heavy
  1. 1Claude Code·Fable 5

    Elo 1,209 · 60.6% · 7m 12s · 283 tasks

    $1.53$4.36$12
  2. 2Codex·GPT 5.6 Sol

    Elo 1,190 · 57.9% · 5m 13s · 124 tasks · provisional

    $1.50$4.43
  3. 3Claude Code·Opus 5

    Elo 1,170 · 55.1% · 6m 37s · 73 tasks · provisional

    $2.69$8.90
  4. 4Trooper·Luna

    Elo 1,144 · 51.6% · 4m 48s · 42 tasks · provisional

    $1.12$3.40
  5. 5Claude Code·Opus 4.8

    Elo 1,131 · 49.5% · 5m 52s · 49 tasks · provisional

    $1.88$5.38
  6. 6Hermes·DS Pro

    Elo 1,099 · 44.9% · 4m 42s · 26 tasks · provisional

    $0.39
  7. 7Claude Code·Sonnet 4.6

    Elo 1,083 · 42.7% · 3m 55s · 24 tasks · provisional

    $0.92$3.14
  8. 8Trooper·DS Flash

    Elo 1,070 · 40.8% · 2m 41s · 24 tasks · provisional

    $0.13
  9. 9Hermes·Kimi K3

    Elo 1,059 · 39.3% · 6m 46s · 24 tasks · provisional

    $1.27$4.47

Email & inbox

Triage, replies, and follow-ups on real inboxes2,335 tasks.

Elo rankings

Head-to-head rating on Email & inbox work — longer bar is better.

Claude · Fable 5
1,198
Codex · GPT 5.6 Sol
1,184
Claude · Opus 5
1,162
Trooper · Luna
1,141
Claude · Opus 4.8
1,134
Hermes · DS Pro
1,106
Claude · Sonnet 4.6
1,092
Trooper · DS Flash
1,084
Hermes · DS Flash
1,071
  • Trooper
  • Claude Code
  • Codex
  • Hermes

Elo vs cost

Up and to the left wins more for less.

Quality — Elo ↑

1,0251,0751,1251,1751,225
$0.10$1.00

Median cost per task — log scale

  • Trooper
  • Claude Code
  • Codex
  • Hermes

Cost per task

Light, typical, and heavy task cost for each pair.

  • light
  • typical
  • heavy
  1. 1Claude Code·Fable 5

    Elo 1,198 · 58.9% · 4m 30s · 1,067 tasks

    $1.12$2.53$9.14
  2. 2Codex·GPT 5.6 Sol

    Elo 1,184 · 57% · 3m 15s · 467 tasks

    $0.87$2.42
  3. 3Claude Code·Opus 5

    Elo 1,162 · 53.8% · 4m 8s · 277 tasks

    $1.56$4.88
  4. 4Trooper·Luna

    Elo 1,141 · 51.4% · 3m 6s · 79 tasks · provisional

    $0.72$2.20
  5. 5Claude Code·Opus 4.8

    Elo 1,134 · 49.8% · 3m 39s · 183 tasks · provisional

    $1.09$4.45
  6. 6Hermes·DS Pro

    Elo 1,106 · 45.8% · 2m 57s · 99 tasks · provisional

    $0.23
  7. 7Claude Code·Sonnet 4.6

    Elo 1,092 · 43.8% · 2m 27s · 63 tasks · provisional

    $0.53
  8. 8Trooper·DS Flash

    Elo 1,084 · 42.6% · 2m 2s · 34 tasks · provisional

    $0.10
  9. 9Hermes·DS Flash

    Elo 1,071 · 40.9% · 2m 27s · 66 tasks · provisional

    $0.14

Data & analytics

Queries, dashboards, number-crunching629 tasks.

Elo rankings

Head-to-head rating on Data & analytics work — longer bar is better.

Claude · Fable 5
1,198
Codex · GPT 5.6 Sol
1,195
Claude · Opus 5
1,161
Trooper · Luna
1,146
Claude · Opus 4.8
1,126
Hermes · DS Pro
1,093
Claude · Sonnet 4.6
1,093
Trooper · DS Flash
1,075
Hermes · Kimi K3
1,058
  • Trooper
  • Claude Code
  • Codex
  • Hermes

Elo vs cost

Up and to the left wins more for less.

Quality — Elo ↑

1,0251,0751,1251,1751,225
$0.10$1.00

Median cost per task — log scale

  • Trooper
  • Claude Code
  • Codex
  • Hermes

Cost per task

Light, typical, and heavy task cost for each pair.

  • light
  • typical
  • heavy
  1. 1Claude Code·Fable 5

    Elo 1,198 · 59.4% · 5m 49s · 210 tasks · provisional

    $2.19$4.38$12
  2. 2Codex·GPT 5.6 Sol

    Elo 1,195 · 59% · 5m 40s · 167 tasks · provisional

    $1.91$7.36
  3. 3Claude Code·Opus 5

    Elo 1,161 · 54.2% · 5m 29s · 57 tasks · provisional

    $2.76$11
  4. 4Trooper·Luna

    Elo 1,146 · 52% · 4m 51s · 42 tasks · provisional

    $1.44$4.80
  5. 5Claude Code·Opus 4.8

    Elo 1,126 · 49.1% · 5m 6s · 42 tasks · provisional

    $2.00$7.88
  6. 6Hermes·DS Pro

    Elo 1,093 · 44.4% · 4m 49s · 31 tasks · provisional

    $0.47
  7. 7Claude Code·Sonnet 4.6

    Elo 1,093 · 44.4% · 5m 2s · 32 tasks · provisional

    $1.34$4.09
  8. 8Trooper·DS Flash

    Elo 1,075 · 41.8% · 2m 55s · 24 tasks · provisional

    $0.15
  9. 9Hermes·Kimi K3

    Elo 1,058 · 39.5% · 7m 5s · 24 tasks · provisional

    $1.57$5.36

Content writing

Articles, docs, and copy589 tasks.

Elo rankings

Head-to-head rating on Content writing work — longer bar is better.

Claude · Fable 5
1,204
Codex · GPT 5.6 Sol
1,176
Claude · Opus 5
1,166
Trooper · Luna
1,139
Claude · Opus 4.8
1,128
Hermes · DS Pro
1,115
Claude · Sonnet 4.6
1,091
Hermes · Kimi K3
1,082
Trooper · DS Flash
1,073
  • Trooper
  • Claude Code
  • Codex
  • Hermes

Elo vs cost

Up and to the left wins more for less.

Quality — Elo ↑

1,0251,0751,1251,1751,225
$0.10$1.00

Median cost per task — log scale

  • Trooper
  • Claude Code
  • Codex
  • Hermes

Cost per task

Light, typical, and heavy task cost for each pair.

  • light
  • typical
  • heavy
  1. 1Claude Code·Fable 5

    Elo 1,204 · 59.5% · 4m 39s · 208 tasks · provisional

    $1.65$3.56$13
  2. 2Codex·GPT 5.6 Sol

    Elo 1,176 · 55.5% · 3m 41s · 63 tasks · provisional

    $1.32$5.15
  3. 3Claude Code·Opus 5

    Elo 1,166 · 54.1% · 4m 44s · 117 tasks · provisional

    $2.38$7.18
  4. 4Trooper·Luna

    Elo 1,139 · 50.8% · 3m 33s · 42 tasks · provisional

    $1.08$3.60
  5. 5Claude Code·Opus 4.8

    Elo 1,128 · 48.6% · 4m 17s · 46 tasks · provisional

    $1.69$6.55
  6. 6Hermes·DS Pro

    Elo 1,115 · 46.8% · 4m 26s · 41 tasks · provisional

    $0.43
  7. 7Claude Code·Sonnet 4.6

    Elo 1,091 · 43.4% · 3m 31s · 24 tasks · provisional

    $0.98$2.88
  8. 8Hermes·Kimi K3

    Elo 1,082 · 42.1% · 6m 21s · 24 tasks · provisional

    $1.40$4.46
  9. 9Trooper·DS Flash

    Elo 1,073 · 41.2% · 2m 18s · 24 tasks · provisional

    $0.13

SEO

Rankings, audits, and site optimization415 tasks.

Elo rankings

Head-to-head rating on SEO work — longer bar is better.

Claude · Fable 5
1,204
Codex · GPT 5.6 Sol
1,176
Claude · Opus 5
1,170
Trooper · Luna
1,143
Claude · Opus 4.8
1,138
Hermes · DS Pro
1,116
Claude · Sonnet 4.6
1,085
Trooper · DS Flash
1,071
Hermes · Kimi K3
1,059
  • Trooper
  • Claude Code
  • Codex
  • Hermes

Elo vs cost

Up and to the left wins more for less.

Quality — Elo ↑

1,0251,0751,1251,1751,225
$0.10$1.00$10

Median cost per task — log scale

  • Trooper
  • Claude Code
  • Codex
  • Hermes

Cost per task

Light, typical, and heavy task cost for each pair.

  • light
  • typical
  • heavy
  1. 1Claude Code·Fable 5

    Elo 1,204 · 59.7% · 10m 41s · 154 tasks · provisional

    $2.91$6.57$24
  2. 2Codex·GPT 5.6 Sol

    Elo 1,176 · 55.8% · 6m 36s · 48 tasks · provisional

    $1.99$6.67
  3. 3Claude Code·Opus 5

    Elo 1,170 · 55% · 10m 27s · 45 tasks · provisional

    $4.27$15
  4. 4Trooper·Luna

    Elo 1,143 · 51.3% · 7m 12s · 42 tasks · provisional

    $1.66$5.20
  5. 5Claude Code·Opus 4.8

    Elo 1,138 · 50.4% · 9m 16s · 30 tasks · provisional

    $2.98$11
  6. 6Hermes·DS Pro

    Elo 1,116 · 47.2% · 8m 18s · 24 tasks · provisional

    $0.68
  7. 7Claude Code·Sonnet 4.6

    Elo 1,085 · 42.8% · 6m 35s · 24 tasks · provisional

    $1.53$5.08
  8. 8Trooper·DS Flash

    Elo 1,071 · 40.6% · 4m 22s · 24 tasks · provisional

    $0.19
  9. 9Hermes·Kimi K3

    Elo 1,059 · 39.2% · 9m 27s · 24 tasks · provisional

    $1.83

Management & coordination

Delegation, scheduling, and follow-through31,334 tasks.

Elo rankings

Head-to-head rating on Management & coordination work — longer bar is better.

Claude · Fable 5
1,192
Codex · GPT 5.6 Sol
1,173
Claude · Opus 5
1,160
Claude · Opus 4.8
1,137
Trooper · Luna
1,134
Hermes · DS Pro
1,109
Claude · Sonnet 4.6
1,100
Trooper · DS Flash
1,088
  • Trooper
  • Claude Code
  • Codex
  • Hermes

Elo vs cost

Up and to the left wins more for less.

Quality — Elo ↑

1,0251,0751,1251,1751,225
$0.10$1.00

Median cost per task — log scale

  • Trooper
  • Claude Code
  • Codex
  • Hermes

Cost per task

Light, typical, and heavy task cost for each pair.

  • light
  • typical
  • heavy
  1. 1Claude Code·Fable 5

    Elo 1,192 · 56.7% · 2m 54s · 14,520 tasks

    $2.75$11
  2. 2Codex·GPT 5.6 Sol

    Elo 1,173 · 54% · 2m 1s · 5,708 tasks

    $0.91$2.81
  3. 3Claude Code·Opus 5

    Elo 1,160 · 52.1% · 2m 51s · 4,262 tasks

    $1.79$7.25
  4. 4Claude Code·Opus 4.8

    Elo 1,137 · 48.8% · 2m 31s · 2,787 tasks

    $1.24$4.47
  5. 5Trooper·Luna

    Elo 1,134 · 49.6% · 2m 8s · 386 tasks

    $0.78$2.40
  6. 6Hermes·DS Pro

    Elo 1,109 · 44.8% · 2m 12s · 1,759 tasks

    $0.28
  7. 7Claude Code·Sonnet 4.6

    Elo 1,100 · 43.5% · 2m 13s · 1,646 tasks

    $0.76
  8. 8Trooper·DS Flash

    Elo 1,088 · 42.2% · 1m 41s · 266 tasks

    $0.09

Methodology

What these numbers mean and where they come from.

Every row aggregates tasks that Trooper users' agents ran between 2026-05-06 and 2026-08-06. Nothing here is a lab exercise: each task had an owner waiting on the result, and each pair is measured on the work it was actually given. Trooper is scored as another harness in the same snapshot — not a separate lab run.

Elo comes from head-to-head comparisons on comparable work — pairs are matched within the same task category and period, and the better outcome wins the matchup. Ratings center on 1000. Because comparisons are cohort-matched, a pair can't buy rank by only running easy work.

Win rate is the share of those matchups a pair wins — 50% is the field average. A matchup compares what actually happened to each task: delivered with a passing review beats delivered, which beats stalled, which beats abandoned.

Cost is the median all-in cost of a task: model usage plus the tools and storage the task consumed. Time is the median wall-clock execution time — treat it as time-to-done, not thinking speed.

Rows need at least 20 tasks in a category to appear; rows under 250 tasks are marked provisional. The efficient frontier is the set of pairs where no other pair beats them on both Elo and cost at once.

Production traffic is not a controlled experiment — pairs receive different task mixes, different users, and different context lengths. These rankings describe what happened on real Trooper workloads during the snapshot window; they are not a guarantee of future performance.

FAQ

Quick answers on how to read the leaderboard.

Where does this data come from?
From real usage on Trooper: tasks that users' agents ran in production over the snapshot window, aggregated per harness + model pair. Pairs that are still accumulating volume are marked provisional and carry modeled estimates calibrated to adjacent measurements.
What does the Elo rating mean?
Pairs are compared head-to-head on comparable work — same task category, same period — and the better outcome wins the matchup. Ratings center on 1000, so a 40-point gap is a clear edge and a 10-point gap is noise.
Does the #1 pair overall mean it is the best choice for me?
Not necessarily. The overall board rewards quality across every kind of work; the per-category boards are the better guide, and the cost and time columns matter as much as the rating — a pair a few points lower at a tenth of the cost is often the right call.
Why is a harness or model missing?
Rows need at least 20 tasks in a category to appear at all. Newly added models and harnesses show up as provisional first and graduate once they cross 250 tasks. Trooper × DeepSeek V4 Flash and Trooper × ChatGPT Luna are included in this snapshot.
How often do the rankings update?
The leaderboard is a snapshot, refreshed periodically from production data — the current snapshot date is shown at the top of the page.

Try Trooper now.

Stand up AI units that write code, manage tasks, and connect to 3,000+ tools — without the overhead of hiring.