Rank排名 8 OpenCode Zhipu GLM

GLM 5.1

A competitive mid-table result with 45/151 tasks solved at least once and 27/151 solved in all three attempts; strongest around large Python/Django application repairs plus localized Go security-scanner changes. 这是一个有竞争力的中游结果:151 题中至少一次解出 45 题,三次都解出 27 题;强项主要落在大型 Python/Django 应用修复以及边界相对清楚的 Go 漏洞扫描器改动。

opencode-cli 1.14.32 zai-coding-plan/glm-5.1 Updated更新 2026-06-18

How to read this result可以这样读

  • GLM 5.1 is best read as moderately stable: rank #8, 45 reached tasks, 27 stable solves.GLM 5.1 更适合读成中等稳定型:排名 #8,触达 45 题,稳定解出 27 题。
  • Best suite signal: Open Library · release 013 at 8/10 (80.0%).最强 suite 信号:Open Library · release 013,8/10(80.0%)。
  • Weakest visible area: qutebrowser · release 018 at 0/9 (0.0%).最弱可见区域:qutebrowser 浏览器 · release 018,0/9(0.0%)。
  • Because the agent shell is OpenCode, the result mostly exposes the underlying model's planning habits rather than a heavily opinionated workflow.因为 agent shell 是 OpenCode,这个结果更直接暴露底层模型的规划习惯,而不是强工作流包装后的表现。

GLM 5.1 is a moderately stable row around the #8 slot. The useful reading is not just the 31.74 score, but the split between 45 reached tasks and 27 stable solves.

The closest family reference is GLM 5.2 at rank #1. Compared with that row, this one is 5.85 points behind, with 12 fewer reached tasks and 9 fewer stable solves.

The profile has one obvious anchor: Open Library · release 013 at 8/10 (80.0%). That anchor matters because qutebrowser · release 018 at 0/9 (0.0%) shows the score does not generalize evenly across the benchmark. Because the agent shell is OpenCode, the result mostly exposes the underlying model’s planning habits rather than a heavily opinionated workflow.

Selected high and low suites, grouped by pass-at-least-once rate.选取高分和低分 suite,按三次尝试至少解出一次的比例展示。
Open Library · release 013Open Library · release 013 8/10 · 80.0%

Best visible cluster for this row: 8/10 tasks reached.这一行最明显的强项簇:10 题中解出 8 题。

vuls · release 012vuls 漏洞扫描器 · release 012 3/4 · 75.0%
Open Library · release 014Open Library · release 014 6/10 · 60.0%
Ansible · release 003Ansible 自动化 · release 003 4/10 · 40.0%
Flipt · release 007Flipt feature flag 服务 · release 007 4/10 · 40.0%
vuls · release 010vuls 漏洞扫描器 · release 010 4/10 · 40.0%
qutebrowser · release 018qutebrowser 浏览器 · release 018 0/9 · 0.0%

Weak cluster: browser/runtime integration around QtWebEngine behavior resisted this model-agent pairing.弱项簇:围绕 QtWebEngine 行为的浏览器/runtime 集成对这个模型-agent 组合不友好。

Navidrome · release 017Navidrome 音乐服务 · release 017 0/5 · 0.0%

Weak cluster: Go service work with persistence and API behavior resisted this model-agent pairing.弱项簇:涉及持久化和 API 行为的 Go 服务改动对这个模型-agent 组合不友好。

Ansible · release 002Ansible 自动化 · release 002 1/10 · 10.0%
Flipt · release 005Flipt feature flag 服务 · release 005 1/10 · 10.0%

The chart is not trying to crown a single strength; it shows how quickly the row falls from Open Library · release 013 at 8/10 (80.0%) to qutebrowser · release 018 at 0/9 (0.0%).

The examples keep the middle-band story honest: Feature Request: Add caching support for evaluation rollouts is the upside, Python module shebang not honored; interpreter forced to /usr/bin/python is the failure surface, and the page should be read between those two poles.

The verifier audit keeps 110/110 solved attempts for GLM 5.1, so the interesting question is not score inflation; it is where the model repeatedly finds the same kind of patch.

Original harness result vs verifier-backed audit sample原始 harness 结果 vs verifier-backed 复核样本
110 of 110 headline successes survived strict re-verification. 110 次初始成功里,110 次通过了更严格的复核。

The available audit keeps 110 of 110 initial solved attempts. Read this as a robustness check, especially when the audit sample is smaller than 453 attempts.当前可用复核保留了 110 次初始成功中的 110 次。这更适合作为稳健性检查,特别是在复核样本小于 453 次尝试时。

110 verifier-backed复核通过 0 strict rejected严格拒绝
31.74 31.74 +0.00 points+0.00 分

In practice, read it through the gap between Open Library · release 013 at 8/10 (80.0%) and qutebrowser · release 018 at 0/9 (0.0%). That gap is more actionable than the rank because it says which repo shape gets coherent patches. The 110/453 attempt score is the backdrop; the article above is about which parts of that score are repeatable enough to matter.

Supporting suite table
Suite Repo Solved Pass^3 Rate
release-zh-013-internetarchive-openlibrary internetarchive/openlibrary 8/10 6 80.0%
release-zh-012-future-architect-vuls future-architect/vuls 3/4 3 75.0%
release-zh-014-internetarchive-openlibrary internetarchive/openlibrary 6/10 3 60.0%
release-zh-003-ansible-ansible ansible/ansible 4/10 2 40.0%
release-zh-007-flipt-io-flipt flipt-io/flipt 4/10 3 40.0%
release-zh-010-future-architect-vuls future-architect/vuls 4/10 2 40.0%
release-zh-018-qutebrowser-qutebrowser qutebrowser/qutebrowser 0/9 0 0.0%
release-zh-017-navidrome-navidrome navidrome/navidrome 0/5 0 0.0%
release-zh-002-ansible-ansible ansible/ansible 1/10 1 10.0%
release-zh-005-flipt-io-flipt flipt-io/flipt 1/10 0 10.0%

GLM 5.1 是一个排名 #8 附近的中等稳定型结果。它的重点不只是 31.74 分,而是 45 道触达题和 27 道稳定题之间的差距。

最接近的同系参照是排名 #1 的 GLM 5.2。和它相比,这一行最终分低 5.85 分,触达题少 12 个,稳定题少 9 个。

这组画像有一个明显锚点:Open Library · release 013,8/10(80.0%)。这个锚点重要,是因为qutebrowser 浏览器 · release 018,0/9(0.0%)说明分数没有均匀迁移到整套 benchmark。因为 agent shell 是 OpenCode,这个结果更直接暴露底层模型的规划习惯,而不是强工作流包装后的表现。

Selected high and low suites, grouped by pass-at-least-once rate.选取高分和低分 suite,按三次尝试至少解出一次的比例展示。
Open Library · release 013Open Library · release 013 8/10 · 80.0%

Best visible cluster for this row: 8/10 tasks reached.这一行最明显的强项簇:10 题中解出 8 题。

vuls · release 012vuls 漏洞扫描器 · release 012 3/4 · 75.0%
Open Library · release 014Open Library · release 014 6/10 · 60.0%
Ansible · release 003Ansible 自动化 · release 003 4/10 · 40.0%
Flipt · release 007Flipt feature flag 服务 · release 007 4/10 · 40.0%
vuls · release 010vuls 漏洞扫描器 · release 010 4/10 · 40.0%
qutebrowser · release 018qutebrowser 浏览器 · release 018 0/9 · 0.0%

Weak cluster: browser/runtime integration around QtWebEngine behavior resisted this model-agent pairing.弱项簇:围绕 QtWebEngine 行为的浏览器/runtime 集成对这个模型-agent 组合不友好。

Navidrome · release 017Navidrome 音乐服务 · release 017 0/5 · 0.0%

Weak cluster: Go service work with persistence and API behavior resisted this model-agent pairing.弱项簇:涉及持久化和 API 行为的 Go 服务改动对这个模型-agent 组合不友好。

Ansible · release 002Ansible 自动化 · release 002 1/10 · 10.0%
Flipt · release 005Flipt feature flag 服务 · release 005 1/10 · 10.0%

这张图不是为了给单一强项加冕,而是展示这一行从Open Library · release 013,8/10(80.0%)滑到qutebrowser 浏览器 · release 018,0/9(0.0%)有多快。

这些案例让中段模型画像更具体:Feature Request:为 evaluation rollouts 添加缓存支持 是上限,Python module shebang 未被遵守;interpreter 被强制为 /usr/bin/python 是失败面,这页应该在两者之间读。

GLM 5.1 的复核保留了 110 次成功中的 110 次,所以重点不是分数膨胀,而是模型在哪些地方能反复找到同类补丁。

Original harness result vs verifier-backed audit sample原始 harness 结果 vs verifier-backed 复核样本
110 of 110 headline successes survived strict re-verification. 110 次初始成功里,110 次通过了更严格的复核。

The available audit keeps 110 of 110 initial solved attempts. Read this as a robustness check, especially when the audit sample is smaller than 453 attempts.当前可用复核保留了 110 次初始成功中的 110 次。这更适合作为稳健性检查,特别是在复核样本小于 453 次尝试时。

110 verifier-backed复核通过 0 strict rejected严格拒绝
31.74 31.74 +0.00 points+0.00 分

实际选择时,更应该通过Open Library · release 013,8/10(80.0%)和qutebrowser 浏览器 · release 018,0/9(0.0%)之间的落差来读它。这个落差比分数排名更可操作,因为它说明哪类代码库更容易得到连贯补丁。110/453 的单次尝试成功数只是背景;上面的文章重点是哪些部分足够可重复、值得当成能力看。

支撑这个判断的 suite 表
Suite Repo 解出 Pass^3 通过率
release-zh-013-internetarchive-openlibrary internetarchive/openlibrary 8/10 6 80.0%
release-zh-012-future-architect-vuls future-architect/vuls 3/4 3 75.0%
release-zh-014-internetarchive-openlibrary internetarchive/openlibrary 6/10 3 60.0%
release-zh-003-ansible-ansible ansible/ansible 4/10 2 40.0%
release-zh-007-flipt-io-flipt flipt-io/flipt 4/10 3 40.0%
release-zh-010-future-architect-vuls future-architect/vuls 4/10 2 40.0%
release-zh-018-qutebrowser-qutebrowser qutebrowser/qutebrowser 0/9 0 0.0%
release-zh-017-navidrome-navidrome navidrome/navidrome 0/5 0 0.0%
release-zh-002-ansible-ansible ansible/ansible 1/10 1 10.0%
release-zh-005-flipt-io-flipt flipt-io/flipt 1/10 0 10.0%