Rank排名 12 Claude Code Xiaomi

MiMo v2.5 pro

A competitive mid-table result with 44/151 tasks solved at least once and 25/151 solved in all three attempts; strongest around large Python/Django application repairs plus localized Go security-scanner changes. 这是一个有竞争力的中游结果:151 题中至少一次解出 44 题,三次都解出 25 题;强项主要落在大型 Python/Django 应用修复以及边界相对清楚的 Go 漏洞扫描器改动。

claude-code 2.1.158 xiaomi/mimo-v2.5-pro Updated更新 2026-06-18

How to read this result可以这样读

  • MiMo v2.5 pro is best read as moderately stable: rank #12, 44 reached tasks, 25 stable solves.MiMo v2.5 pro 更适合读成中等稳定型:排名 #12,触达 44 题,稳定解出 25 题。
  • Best suite signal: Open Library · release 013 at 8/10 (80.0%).最强 suite 信号:Open Library · release 013,8/10(80.0%)。
  • Weakest visible area: qutebrowser · release 018 at 0/9 (0.0%).最弱可见区域:qutebrowser 浏览器 · release 018,0/9(0.0%)。
  • Claude Code gives this row a different orchestration profile from OpenCode and Qoder, which is useful when comparing the same model family across shells.Claude Code 让这一行有别于 OpenCode 和 Qoder 的编排形态,适合观察同类模型跨 shell 的差异。

MiMo v2.5 pro is a moderately stable row around the #12 slot. The useful reading is not just the 30.69 score, but the split between 44 reached tasks and 25 stable solves.

The closest family reference is MiMo v2.5 pro (high) at rank #13. Compared with that row, this one is 0.39 points ahead, with 2 fewer reached tasks and 3 more stable solves.

Most of the positive signal concentrates in Open Library · release 013 at 8/10 (80.0%). The opposing read is qutebrowser · release 018 at 0/9 (0.0%), which keeps the row from looking like a generalist. Claude Code gives this row a different orchestration profile from OpenCode and Qoder, which is useful when comparing the same model family across shells.

Selected high and low suites, grouped by pass-at-least-once rate.选取高分和低分 suite,按三次尝试至少解出一次的比例展示。
Open Library · release 013Open Library · release 013 8/10 · 80.0%

Best visible cluster for this row: 8/10 tasks reached.这一行最明显的强项簇:10 题中解出 8 题。

vuls · release 012vuls 漏洞扫描器 · release 012 3/4 · 75.0%
Open Library · release 015Open Library · release 015 5/10 · 50.0%
Flipt · release 007Flipt feature flag 服务 · release 007 4/10 · 40.0%
Open Library · release 016Open Library · release 016 2/5 · 40.0%
Ansible · release 004Ansible 自动化 · release 004 1/3 · 33.3%
qutebrowser · release 018qutebrowser 浏览器 · release 018 0/9 · 0.0%

Weak cluster: browser/runtime integration around QtWebEngine behavior resisted this model-agent pairing.弱项簇:围绕 QtWebEngine 行为的浏览器/runtime 集成对这个模型-agent 组合不友好。

Flipt · release 005Flipt feature flag 服务 · release 005 1/10 · 10.0%
Flipt · release 006Flipt feature flag 服务 · release 006 1/10 · 10.0%
Flipt · release 008Flipt feature flag 服务 · release 008 1/10 · 10.0%

For this row, the suite bars are a contrast tool. The distance between Open Library · release 013 at 8/10 (80.0%) and qutebrowser · release 018 at 0/9 (0.0%) is the model’s practical boundary.

The examples keep the middle-band story honest: Polling goroutines lack lifecycle management in storage backends is the upside, module_defaults of the underlying module are not applied when invoked via action plugins (gather_facts, package, service) is the failure surface, and the page should be read between those two poles.

The audit trims 29 solved attempts from MiMo v2.5 pro but still keeps 72% of the solved set, so the suite shape remains useful even where individual wins are debatable.

Original harness result vs verifier-backed audit sample原始 harness 结果 vs verifier-backed 复核样本
74 of 103 headline successes survived strict re-verification. 103 次初始成功里,74 次通过了更严格的复核。

The available audit keeps 74 of 103 initial solved attempts. Read this as a robustness check, especially when the audit sample is smaller than 453 attempts.当前可用复核保留了 103 次初始成功中的 74 次。这更适合作为稳健性检查,特别是在复核样本小于 453 次尝试时。

74 verifier-backed复核通过 29 strict rejected严格拒绝
30.69 30.69 +0.00 points+0.00 分

In practice, read it through the gap between Open Library · release 013 at 8/10 (80.0%) and qutebrowser · release 018 at 0/9 (0.0%). That gap is more actionable than the rank because it says which repo shape gets coherent patches. The 103/453 attempt score is the backdrop; the article above is about which parts of that score are repeatable enough to matter.

Supporting suite table
Suite Repo Solved Pass^3 Rate
release-zh-013-internetarchive-openlibrary internetarchive/openlibrary 8/10 4 80.0%
release-zh-012-future-architect-vuls future-architect/vuls 3/4 3 75.0%
release-zh-015-internetarchive-openlibrary internetarchive/openlibrary 5/10 2 50.0%
release-zh-007-flipt-io-flipt flipt-io/flipt 4/10 3 40.0%
release-zh-016-internetarchive-openlibrary internetarchive/openlibrary 2/5 1 40.0%
release-zh-004-ansible-ansible ansible/ansible 1/3 1 33.3%
release-zh-018-qutebrowser-qutebrowser qutebrowser/qutebrowser 0/9 0 0.0%
release-zh-005-flipt-io-flipt flipt-io/flipt 1/10 0 10.0%
release-zh-006-flipt-io-flipt flipt-io/flipt 1/10 1 10.0%
release-zh-008-flipt-io-flipt flipt-io/flipt 1/10 1 10.0%

MiMo v2.5 pro 是一个排名 #12 附近的中等稳定型结果。它的重点不只是 30.69 分,而是 44 道触达题和 25 道稳定题之间的差距。

最接近的同系参照是排名 #13 的 MiMo v2.5 pro (high)。和它相比,这一行最终分高 0.39 分,触达题少 2 个,稳定题多 3 个。

正面信号大多集中在Open Library · release 013,8/10(80.0%)。反向读法是qutebrowser 浏览器 · release 018,0/9(0.0%),它让这一行看起来不像通用型。Claude Code 让这一行有别于 OpenCode 和 Qoder 的编排形态,适合观察同类模型跨 shell 的差异。

Selected high and low suites, grouped by pass-at-least-once rate.选取高分和低分 suite,按三次尝试至少解出一次的比例展示。
Open Library · release 013Open Library · release 013 8/10 · 80.0%

Best visible cluster for this row: 8/10 tasks reached.这一行最明显的强项簇:10 题中解出 8 题。

vuls · release 012vuls 漏洞扫描器 · release 012 3/4 · 75.0%
Open Library · release 015Open Library · release 015 5/10 · 50.0%
Flipt · release 007Flipt feature flag 服务 · release 007 4/10 · 40.0%
Open Library · release 016Open Library · release 016 2/5 · 40.0%
Ansible · release 004Ansible 自动化 · release 004 1/3 · 33.3%
qutebrowser · release 018qutebrowser 浏览器 · release 018 0/9 · 0.0%

Weak cluster: browser/runtime integration around QtWebEngine behavior resisted this model-agent pairing.弱项簇:围绕 QtWebEngine 行为的浏览器/runtime 集成对这个模型-agent 组合不友好。

Flipt · release 005Flipt feature flag 服务 · release 005 1/10 · 10.0%
Flipt · release 006Flipt feature flag 服务 · release 006 1/10 · 10.0%
Flipt · release 008Flipt feature flag 服务 · release 008 1/10 · 10.0%

对这一行来说,suite 柱更像对比工具。Open Library · release 013,8/10(80.0%)和qutebrowser 浏览器 · release 018,0/9(0.0%)之间的距离,就是模型的实用边界。

这些案例让中段模型画像更具体:storage 后端中的 polling goroutine 缺少生命周期管理 是上限,通过 action plugins(gather_facts、package、service)调用时,底层 module 的 module_defaults 没有被应用 是失败面,这页应该在两者之间读。

复核从 MiMo v2.5 pro 中剔除了 29 次成功,但仍保留 72% 的成功集合,因此即便个别胜利有争议,suite 形状仍然有参考价值。

Original harness result vs verifier-backed audit sample原始 harness 结果 vs verifier-backed 复核样本
74 of 103 headline successes survived strict re-verification. 103 次初始成功里,74 次通过了更严格的复核。

The available audit keeps 74 of 103 initial solved attempts. Read this as a robustness check, especially when the audit sample is smaller than 453 attempts.当前可用复核保留了 103 次初始成功中的 74 次。这更适合作为稳健性检查,特别是在复核样本小于 453 次尝试时。

74 verifier-backed复核通过 29 strict rejected严格拒绝
30.69 30.69 +0.00 points+0.00 分

实际选择时,更应该通过Open Library · release 013,8/10(80.0%)和qutebrowser 浏览器 · release 018,0/9(0.0%)之间的落差来读它。这个落差比分数排名更可操作,因为它说明哪类代码库更容易得到连贯补丁。103/453 的单次尝试成功数只是背景;上面的文章重点是哪些部分足够可重复、值得当成能力看。

支撑这个判断的 suite 表
Suite Repo 解出 Pass^3 通过率
release-zh-013-internetarchive-openlibrary internetarchive/openlibrary 8/10 4 80.0%
release-zh-012-future-architect-vuls future-architect/vuls 3/4 3 75.0%
release-zh-015-internetarchive-openlibrary internetarchive/openlibrary 5/10 2 50.0%
release-zh-007-flipt-io-flipt flipt-io/flipt 4/10 3 40.0%
release-zh-016-internetarchive-openlibrary internetarchive/openlibrary 2/5 1 40.0%
release-zh-004-ansible-ansible ansible/ansible 1/3 1 33.3%
release-zh-018-qutebrowser-qutebrowser qutebrowser/qutebrowser 0/9 0 0.0%
release-zh-005-flipt-io-flipt flipt-io/flipt 1/10 0 10.0%
release-zh-006-flipt-io-flipt flipt-io/flipt 1/10 1 10.0%
release-zh-008-flipt-io-flipt flipt-io/flipt 1/10 1 10.0%