How long one DeepSWE task takes
Shorter is better. Pass@1 and repriced task cost remain beside the measured mean wall time.
| Model | Effort | Mean duration | Pass at one | Cost per task | Mean output tokens |
|---|
TrigrD Data turns public datasets into readable comparisons. We keep source values, derived values, and opinion separate so people and agents can inspect the same evidence. Write to us at contact@trigrd.tech.
DeepSWE v1.1
The public artifact records end-to-end duration for every scored attempt. This page covers 14 selected model families across labs, score bands, and price bands.
Shorter is better. Pass@1 and repriced task cost remain beside the measured mean wall time.
| Model | Effort | Mean duration | Pass at one | Cost per task | Mean output tokens |
|---|
Pass@1 starts at zero so the score differences stay in proportion. Use the configuration picker to compare individual model and effort combinations.
DeepSWE v1.1
Keep score on the y-axis. Compare it with cost, time, or a weighted x-axis for a cost vs time sensitive model choice.
The Pareto line connects configurations that no lower-burden point can match on score.