How long one DeepSWE task takes
Shorter is better. Pass@1 and repriced task cost remain beside the measured mean wall time.
| Model | Effort | Mean duration | Pass at one | Cost per task | Mean output tokens |
|---|
TrigrD Data turns public datasets into readable comparisons. We keep source values, derived values, and opinion separate so people and agents can inspect the same evidence. Write to us at contact@trigrd.tech.
DeepSWE v1.1
The public artifact records end-to-end duration for every scored attempt. This page covers 15 selected model families across labs, score bands, and price bands.
Shorter is better. Pass@1 and repriced task cost remain beside the measured mean wall time.
| Model | Effort | Mean duration | Pass at one | Cost per task | Mean output tokens |
|---|
Pass@1 starts at zero so the score differences stay in proportion. Use the configuration picker to compare individual model and effort combinations.
DeepSWE v1.1
Keep score on the y-axis. Compare it with cost, time, or a weighted x-axis for a cost vs time sensitive model choice.
The Pareto line connects configurations that no lower-burden point can match on score.
Artificial Analysis · 2026-09-04 snapshot
Compare 35 Artificial Analysis score leaders on two indices. Each model/effort configuration is its own datapoint, with cost, time, and a weighted choice of both on the x-axis.
Higher Intelligence Index is better. The same selector exposes current cost, time per task, and a normalized weighted burden for the selected leaders.
| Rank | Model | Effort | Intelligence Index | Cost per task | Time per task |
|---|
Higher Agentic Index is better. Move left to favor the lower task burden, or use the weighted view to set your own cost/time preference.
| Rank | Model | Effort | Agentic Index | Cost per task | Time per task |
|---|
CursorBench 3.2 · 2026-09-02 snapshot
Compare all 60 current CursorBench configurations. Keep score on the y-axis and test cost, tokens, steps, or your own cost/steps weighting on the x-axis.
Higher score is better. The Pareto line connects configurations that no lower-burden point can match on score.
| Rank | Model | Effort | Score | Average cost per task | Tokens | Steps |
|---|