Rank排名 24 OpenCode Xiaomi

MiMo v2.5

A lower-table result with a few useful bright spots: 33/151 tasks solved at least once, 18/151 solved in all three attempts, with the clearest wins around automation and configuration-management work plus localized Go security-scanner changes. 这是一个排名靠后但仍有局部亮点的结果:151 题中至少一次解出 33 题,三次都解出 18 题;强项主要落在自动化和配置管理类改动以及边界相对清楚的 Go 漏洞扫描器改动。

opencode-cli 1.14.32 xiaomi-token-plan-cn/mimo-v2.5 Updated更新 2026-06-18

How to read this result可以这样读

  • MiMo v2.5 is best read as moderately stable: rank #24, 33 reached tasks, 18 stable solves.MiMo v2.5 更适合读成中等稳定型:排名 #24,触达 33 题,稳定解出 18 题。
  • Best suite signal: Ansible · release 002 at 7/10 (70.0%).最强 suite 信号:Ansible 自动化 · release 002,7/10(70.0%)。
  • Weakest visible area: Open Library · release 014 at 0/10 (0.0%).最弱可见区域:Open Library · release 014,0/10(0.0%)。
  • Because the agent shell is OpenCode, the result mostly exposes the underlying model's planning habits rather than a heavily opinionated workflow.因为 agent shell 是 OpenCode,这个结果更直接暴露底层模型的规划习惯,而不是强工作流包装后的表现。

MiMo v2.5 is a moderately stable row around the #24 slot. The useful reading is not just the 25.24 score, but the split between 33 reached tasks and 18 stable solves.

The closest family reference is MiMo v2.5 pro at rank #12. Compared with that row, this one is 5.45 points behind, with 11 fewer reached tasks and 7 fewer stable solves.

The suite split is asymmetric: Ansible · release 002 at 7/10 (70.0%) supplies the main body of wins, vuls · release 012 at 3/4 (75.0%) supplies the clean spike, and Open Library · release 014 at 0/10 (0.0%) is where that pattern stops. Because the agent shell is OpenCode, the result mostly exposes the underlying model’s planning habits rather than a heavily opinionated workflow.

Selected high and low suites, grouped by pass-at-least-once rate.选取高分和低分 suite,按三次尝试至少解出一次的比例展示。
vuls · release 012vuls 漏洞扫描器 · release 012 3/4 · 75.0%

Best visible cluster for this row: 3/4 tasks reached.这一行最明显的强项簇:4 题中解出 3 题。

Ansible · release 002Ansible 自动化 · release 002 7/10 · 70.0%
Ansible · release 003Ansible 自动化 · release 003 6/10 · 60.0%
Ansible · release 004Ansible 自动化 · release 004 1/3 · 33.3%
Ansible · release 001Ansible 自动化 · release 001 3/10 · 30.0%
vuls · release 010vuls 漏洞扫描器 · release 010 3/10 · 30.0%
Open Library · release 014Open Library · release 014 0/10 · 0.0%

Weak cluster: large Python/Django application repairs resisted this model-agent pairing.弱项簇:大型 Python/Django 应用修复对这个模型-agent 组合不友好。

Open Library · release 015Open Library · release 015 0/10 · 0.0%

Weak cluster: large Python/Django application repairs resisted this model-agent pairing.弱项簇:大型 Python/Django 应用修复对这个模型-agent 组合不友好。

qutebrowser · release 018qutebrowser 浏览器 · release 018 0/9 · 0.0%

Weak cluster: browser/runtime integration around QtWebEngine behavior resisted this model-agent pairing.弱项簇:围绕 QtWebEngine 行为的浏览器/runtime 集成对这个模型-agent 组合不友好。

Open Library · release 016Open Library · release 016 0/5 · 0.0%

Weak cluster: large Python/Django application repairs resisted this model-agent pairing.弱项簇:大型 Python/Django 应用修复对这个模型-agent 组合不友好。

The chart is not trying to crown a single strength; it shows how quickly the row falls from Ansible · release 002 at 7/10 (70.0%) to Open Library · release 014 at 0/10 (0.0%).

The examples keep the middle-band story honest: Introduce public methods to access PlayIterator._host_states is the upside, Add Reading-Log Counts to Solr Work Documents is the failure surface, and the page should be read between those two poles.

The audit changes how to read MiMo v2.5: only 60% of initial solved attempts survive, with 31 rejected attempts, while the exported score field stays flat. Treat the wins as leads that need stricter confirmation.

Original harness result vs verifier-backed audit sample原始 harness 结果 vs verifier-backed 复核样本
47 of 78 headline successes survived strict re-verification. 78 次初始成功里,47 次通过了更严格的复核。

The available audit keeps 47 of 78 initial solved attempts. Read this as a robustness check, especially when the audit sample is smaller than 453 attempts.当前可用复核保留了 78 次初始成功中的 47 次。这更适合作为稳健性检查,特别是在复核样本小于 453 次尝试时。

47 verifier-backed复核通过 31 strict rejected严格拒绝
25.24 25.24 +0.00 points+0.00 分

In practice, read it through the gap between Ansible · release 002 at 7/10 (70.0%) and Open Library · release 014 at 0/10 (0.0%). That gap is more actionable than the rank because it says which repo shape gets coherent patches. The 78/453 attempt score is the backdrop; the article above is about which parts of that score are repeatable enough to matter.

Supporting suite table
Suite Repo Solved Pass^3 Rate
release-zh-012-future-architect-vuls future-architect/vuls 3/4 2 75.0%
release-zh-002-ansible-ansible ansible/ansible 7/10 2 70.0%
release-zh-003-ansible-ansible ansible/ansible 6/10 4 60.0%
release-zh-004-ansible-ansible ansible/ansible 1/3 1 33.3%
release-zh-001-ansible-ansible ansible/ansible 3/10 2 30.0%
release-zh-010-future-architect-vuls future-architect/vuls 3/10 2 30.0%
release-zh-014-internetarchive-openlibrary internetarchive/openlibrary 0/10 0 0.0%
release-zh-015-internetarchive-openlibrary internetarchive/openlibrary 0/10 0 0.0%
release-zh-018-qutebrowser-qutebrowser qutebrowser/qutebrowser 0/9 0 0.0%
release-zh-016-internetarchive-openlibrary internetarchive/openlibrary 0/5 0 0.0%

MiMo v2.5 是一个排名 #24 附近的中等稳定型结果。它的重点不只是 25.24 分,而是 33 道触达题和 18 道稳定题之间的差距。

最接近的同系参照是排名 #12 的 MiMo v2.5 pro。和它相比,这一行最终分低 5.45 分,触达题少 11 个,稳定题少 7 个。

suite 分布是不对称的:Ansible 自动化 · release 002,7/10(70.0%)贡献主要胜利,vuls 漏洞扫描器 · release 012,3/4(75.0%)贡献最干净高点,而Open Library · release 014,0/10(0.0%)标出这种模式停止的地方。因为 agent shell 是 OpenCode,这个结果更直接暴露底层模型的规划习惯,而不是强工作流包装后的表现。

Selected high and low suites, grouped by pass-at-least-once rate.选取高分和低分 suite,按三次尝试至少解出一次的比例展示。
vuls · release 012vuls 漏洞扫描器 · release 012 3/4 · 75.0%

Best visible cluster for this row: 3/4 tasks reached.这一行最明显的强项簇:4 题中解出 3 题。

Ansible · release 002Ansible 自动化 · release 002 7/10 · 70.0%
Ansible · release 003Ansible 自动化 · release 003 6/10 · 60.0%
Ansible · release 004Ansible 自动化 · release 004 1/3 · 33.3%
Ansible · release 001Ansible 自动化 · release 001 3/10 · 30.0%
vuls · release 010vuls 漏洞扫描器 · release 010 3/10 · 30.0%
Open Library · release 014Open Library · release 014 0/10 · 0.0%

Weak cluster: large Python/Django application repairs resisted this model-agent pairing.弱项簇:大型 Python/Django 应用修复对这个模型-agent 组合不友好。

Open Library · release 015Open Library · release 015 0/10 · 0.0%

Weak cluster: large Python/Django application repairs resisted this model-agent pairing.弱项簇:大型 Python/Django 应用修复对这个模型-agent 组合不友好。

qutebrowser · release 018qutebrowser 浏览器 · release 018 0/9 · 0.0%

Weak cluster: browser/runtime integration around QtWebEngine behavior resisted this model-agent pairing.弱项簇:围绕 QtWebEngine 行为的浏览器/runtime 集成对这个模型-agent 组合不友好。

Open Library · release 016Open Library · release 016 0/5 · 0.0%

Weak cluster: large Python/Django application repairs resisted this model-agent pairing.弱项簇:大型 Python/Django 应用修复对这个模型-agent 组合不友好。

这张图不是为了给单一强项加冕,而是展示这一行从Ansible 自动化 · release 002,7/10(70.0%)滑到Open Library · release 014,0/10(0.0%)有多快。

这些案例让中段模型画像更具体:引入公共方法以访问 PlayIterator._host_states 是上限,向 Solr Work Documents 添加 Reading-Log Counts 是失败面,这页应该在两者之间读。

复核改变了 MiMo v2.5 的读法:初始成功只有 60% 保留下来,31 次被剔除,但当前导出的分数字段没有变化。原始胜利更适合作为线索,需要更严格确认。

Original harness result vs verifier-backed audit sample原始 harness 结果 vs verifier-backed 复核样本
47 of 78 headline successes survived strict re-verification. 78 次初始成功里,47 次通过了更严格的复核。

The available audit keeps 47 of 78 initial solved attempts. Read this as a robustness check, especially when the audit sample is smaller than 453 attempts.当前可用复核保留了 78 次初始成功中的 47 次。这更适合作为稳健性检查,特别是在复核样本小于 453 次尝试时。

47 verifier-backed复核通过 31 strict rejected严格拒绝
25.24 25.24 +0.00 points+0.00 分

实际选择时,更应该通过Ansible 自动化 · release 002,7/10(70.0%)和Open Library · release 014,0/10(0.0%)之间的落差来读它。这个落差比分数排名更可操作,因为它说明哪类代码库更容易得到连贯补丁。78/453 的单次尝试成功数只是背景;上面的文章重点是哪些部分足够可重复、值得当成能力看。

支撑这个判断的 suite 表
Suite Repo 解出 Pass^3 通过率
release-zh-012-future-architect-vuls future-architect/vuls 3/4 2 75.0%
release-zh-002-ansible-ansible ansible/ansible 7/10 2 70.0%
release-zh-003-ansible-ansible ansible/ansible 6/10 4 60.0%
release-zh-004-ansible-ansible ansible/ansible 1/3 1 33.3%
release-zh-001-ansible-ansible ansible/ansible 3/10 2 30.0%
release-zh-010-future-architect-vuls future-architect/vuls 3/10 2 30.0%
release-zh-014-internetarchive-openlibrary internetarchive/openlibrary 0/10 0 0.0%
release-zh-015-internetarchive-openlibrary internetarchive/openlibrary 0/10 0 0.0%
release-zh-018-qutebrowser-qutebrowser qutebrowser/qutebrowser 0/9 0 0.0%
release-zh-016-internetarchive-openlibrary internetarchive/openlibrary 0/5 0 0.0%