Terminal-Bench 2.1 agent results

TerminalBench 2.1

A leaderboard for coding agents on TerminalBench-2.1: 89 tasks. All models run via OpenCode agent. Rankings sort by Final Score (pass@1 × 100). Multi-attempt metrics are shown as secondary diagnostics when available.

Best Final Score 31.46 OpenCode - MiniMax-M3
Completed Models 5 All via OpenCode agent
Tasks 89 TerminalBench-2.1
Exported 2026-07-08 TerminalBench 2.1 zh89

Benchmark Leaderboard

Final Score = pass@1 × 100. TerminalBench rows can be single-try or multi-try; Pass@3 and Pass^3 are secondary diagnostics and are not part of Final Score.

Agent
# Agent / Model Agent Version Final Score Pass@1 Solved Tasks Attempts Tokens (In / Out / Total) Pass^3 Pass@3 Exported Model Dir
1
MiniMax-M3 OpenCodeMiniMax - minimax-cn-coding-plan/MiniMax-M3
opencode-cli 1.17.8 31.46
28/89 - 31.5%
28/89 1x In 302.0M Out 3.0M Total 305.1M N/A N/A 2026-07-06 opencode-minimax-m3
2
MiMo v2.5 OpenCodeMiMo - xiaomi-token-plan-cn/mimo-v2.5
opencode-cli 1.17.8 23.6
21/89 - 23.6%
21/89 1x In 178.9M Out 0.6M Total 179.5M N/A N/A 2026-07-08 opencode-mimo-v2.5
3
MiMo v2.5 Pro OpenCodeMiMo - xiaomi-token-plan-cn/mimo-v2.5-pro
opencode-cli 1.17.8 15.73
14/89 - 15.7%
14/89 1x In 107.4M Out 0.4M Total 107.8M N/A N/A 2026-07-07 opencode-mimo-v2.5-pro
4
Step 3.7 Flash OpenCodeStepFun - stepfun/step-3.7-flash
opencode-cli 1.17.8 13.48
12/89 - 13.5%
12/89 3x In 1.71B Out 29.5M Total 1.74B
0/89 - 0%
17/89 - 19.1%
2026-07-06 opencode-stepfun-3.7-flash
5
GPT-5.4 Mini OpenCodeOpenAI - openai/gpt-5.4-mini
opencode-cli 1.17.8 13.48
12/89 - 13.5%
12/89 1x In 22.6M Out 0.8M Total 23.4M N/A N/A 2026-07-07 opencode-gpt-5.4-mini
Notes: All models run via OpenCode agent on TerminalBench 2.1 (zh89 variant, 89 tasks). Final Score = pass@1 × 100 only; it does not use Pass@3 or Pass^3. Pass@3 counts tasks solved at least once across 3 attempts. Only applicable for models with multiple attempts per task. Pass^3 counts tasks solved in all 3 attempts. Only applicable for models with multiple attempts per task. Pass@1 shows the primary solve rate (first attempt or best single attempt). Step 3.7 Flash: 3 attempts per task, pass@3 = 17/89 (19.1%), pass@3 estimate = 21.6%. MiniMax-M3, MiMo v2.5, MiMo v2.5 Pro, and GPT-5.4 Mini: selected pass@1 only — pass@3 and pass^3 are not applicable. Tokens are summed across all attempts. Input includes cache tokens when available.

Visual Leaderboard

Switch metrics to compare models by different performance dimensions.

MiniMax-M3 OpenCode - MiniMax
MiMo v2.5 OpenCode - MiMo
MiMo v2.5 Pro OpenCode - MiMo
Step 3.7 Flash OpenCode - StepFun
GPT-5.4 Mini OpenCode - OpenAI

Reach vs Consistency

Pass@3 shows whether an agent can solve a task at least once; Pass^3 shows whether it solves the same task all three times. Only models with multiple attempts are plotted.

Pass@3
OpenCode - Step 3.7 Flash
Pass^3