Terminal-Bench 2.1 agent results
TerminalBench 2.1
A leaderboard for coding agents on TerminalBench-2.1: 89 tasks. All models run via OpenCode agent. Rankings sort by Final Score (pass@1 × 100). Multi-attempt metrics are shown as secondary diagnostics when available.
Best Final Score
31.46
OpenCode - MiniMax-M3
Completed Models
5
All via OpenCode agent
Tasks
89
TerminalBench-2.1
Exported
2026-07-08
TerminalBench 2.1 zh89
Benchmark Leaderboard
Final Score = pass@1 × 100. TerminalBench rows can be single-try or multi-try; Pass@3 and Pass^3 are secondary diagnostics and are not part of Final Score.
Agent
Notes:
All models run via OpenCode agent on TerminalBench 2.1 (zh89 variant, 89 tasks).
Final Score = pass@1 × 100 only; it does not use Pass@3 or Pass^3.
Pass@3 counts tasks solved at least once across 3 attempts. Only applicable for models with multiple attempts per task.
Pass^3 counts tasks solved in all 3 attempts. Only applicable for models with multiple attempts per task.
Pass@1 shows the primary solve rate (first attempt or best single attempt).
Step 3.7 Flash: 3 attempts per task, pass@3 = 17/89 (19.1%), pass@3 estimate = 21.6%.
MiniMax-M3, MiMo v2.5, MiMo v2.5 Pro, and GPT-5.4 Mini: selected pass@1 only — pass@3 and pass^3 are not applicable.
Tokens are summed across all attempts. Input includes cache tokens when available.
Visual Leaderboard
Switch metrics to compare models by different performance dimensions.
MiniMax-M3
OpenCode - MiniMax
MiMo v2.5
OpenCode - MiMo
MiMo v2.5 Pro
OpenCode - MiMo
Step 3.7 Flash
OpenCode - StepFun
GPT-5.4 Mini
OpenCode - OpenAI
Reach vs Consistency
Pass@3 shows whether an agent can solve a task at least once; Pass^3 shows whether it solves the same task all three times. Only models with multiple attempts are plotted.
Pass@3
OpenCode - Step 3.7 Flash
Pass^3