GLM 5.1
A competitive mid-table result with 45/151 tasks solved at least once and 27/151 solved in all three attempts; strongest around large Python/Django application repairs plus localized Go security-scanner changes. 这是一个有竞争力的中游结果:151 题中至少一次解出 45 题,三次都解出 27 题;强项主要落在大型 Python/Django 应用修复以及边界相对清楚的 Go 漏洞扫描器改动。
How to read this result可以这样读
- GLM 5.1 is best read as moderately stable: rank #8, 45 reached tasks, 27 stable solves.GLM 5.1 更适合读成中等稳定型:排名 #8,触达 45 题,稳定解出 27 题。
- Best suite signal: Open Library · release 013 at 8/10 (80.0%).最强 suite 信号:Open Library · release 013,8/10(80.0%)。
- Weakest visible area: qutebrowser · release 018 at 0/9 (0.0%).最弱可见区域:qutebrowser 浏览器 · release 018,0/9(0.0%)。
- Because the agent shell is OpenCode, the result mostly exposes the underlying model's planning habits rather than a heavily opinionated workflow.因为 agent shell 是 OpenCode,这个结果更直接暴露底层模型的规划习惯,而不是强工作流包装后的表现。
GLM 5.1 is a moderately stable row around the #8 slot. The useful reading is not just the 31.74 score, but the split between 45 reached tasks and 27 stable solves.
The closest family reference is GLM 5.2 at rank #1. Compared with that row, this one is 5.85 points behind, with 12 fewer reached tasks and 9 fewer stable solves.
The profile has one obvious anchor: Open Library · release 013 at 8/10 (80.0%). That anchor matters because qutebrowser · release 018 at 0/9 (0.0%) shows the score does not generalize evenly across the benchmark. Because the agent shell is OpenCode, the result mostly exposes the underlying model’s planning habits rather than a heavily opinionated workflow.
The chart is not trying to crown a single strength; it shows how quickly the row falls from Open Library · release 013 at 8/10 (80.0%) to qutebrowser · release 018 at 0/9 (0.0%).
The examples keep the middle-band story honest: Feature Request: Add caching support for evaluation rollouts is the upside, Python module shebang not honored; interpreter forced to /usr/bin/python is the failure surface, and the page should be read between those two poles.
The verifier audit keeps 110/110 solved attempts for GLM 5.1, so the interesting question is not score inflation; it is where the model repeatedly finds the same kind of patch.
The available audit keeps 110 of 110 initial solved attempts. Read this as a robustness check, especially when the audit sample is smaller than 453 attempts.当前可用复核保留了 110 次初始成功中的 110 次。这更适合作为稳健性检查,特别是在复核样本小于 453 次尝试时。
In practice, read it through the gap between Open Library · release 013 at 8/10 (80.0%) and qutebrowser · release 018 at 0/9 (0.0%). That gap is more actionable than the rank because it says which repo shape gets coherent patches. The 110/453 attempt score is the backdrop; the article above is about which parts of that score are repeatable enough to matter.
Supporting suite table
| Suite | Repo | Solved | Pass^3 | Rate |
|---|---|---|---|---|
release-zh-013-internetarchive-openlibrary |
internetarchive/openlibrary | 8/10 | 6 | 80.0% |
release-zh-012-future-architect-vuls |
future-architect/vuls | 3/4 | 3 | 75.0% |
release-zh-014-internetarchive-openlibrary |
internetarchive/openlibrary | 6/10 | 3 | 60.0% |
release-zh-003-ansible-ansible |
ansible/ansible | 4/10 | 2 | 40.0% |
release-zh-007-flipt-io-flipt |
flipt-io/flipt | 4/10 | 3 | 40.0% |
release-zh-010-future-architect-vuls |
future-architect/vuls | 4/10 | 2 | 40.0% |
release-zh-018-qutebrowser-qutebrowser |
qutebrowser/qutebrowser | 0/9 | 0 | 0.0% |
release-zh-017-navidrome-navidrome |
navidrome/navidrome | 0/5 | 0 | 0.0% |
release-zh-002-ansible-ansible |
ansible/ansible | 1/10 | 1 | 10.0% |
release-zh-005-flipt-io-flipt |
flipt-io/flipt | 1/10 | 0 | 10.0% |
GLM 5.1 是一个排名 #8 附近的中等稳定型结果。它的重点不只是 31.74 分,而是 45 道触达题和 27 道稳定题之间的差距。
最接近的同系参照是排名 #1 的 GLM 5.2。和它相比,这一行最终分低 5.85 分,触达题少 12 个,稳定题少 9 个。
这组画像有一个明显锚点:Open Library · release 013,8/10(80.0%)。这个锚点重要,是因为qutebrowser 浏览器 · release 018,0/9(0.0%)说明分数没有均匀迁移到整套 benchmark。因为 agent shell 是 OpenCode,这个结果更直接暴露底层模型的规划习惯,而不是强工作流包装后的表现。
这张图不是为了给单一强项加冕,而是展示这一行从Open Library · release 013,8/10(80.0%)滑到qutebrowser 浏览器 · release 018,0/9(0.0%)有多快。
这些案例让中段模型画像更具体:Feature Request:为 evaluation rollouts 添加缓存支持 是上限,Python module shebang 未被遵守;interpreter 被强制为 /usr/bin/python 是失败面,这页应该在两者之间读。
GLM 5.1 的复核保留了 110 次成功中的 110 次,所以重点不是分数膨胀,而是模型在哪些地方能反复找到同类补丁。
The available audit keeps 110 of 110 initial solved attempts. Read this as a robustness check, especially when the audit sample is smaller than 453 attempts.当前可用复核保留了 110 次初始成功中的 110 次。这更适合作为稳健性检查,特别是在复核样本小于 453 次尝试时。
实际选择时,更应该通过Open Library · release 013,8/10(80.0%)和qutebrowser 浏览器 · release 018,0/9(0.0%)之间的落差来读它。这个落差比分数排名更可操作,因为它说明哪类代码库更容易得到连贯补丁。110/453 的单次尝试成功数只是背景;上面的文章重点是哪些部分足够可重复、值得当成能力看。
支撑这个判断的 suite 表
| Suite | Repo | 解出 | Pass^3 | 通过率 |
|---|---|---|---|---|
release-zh-013-internetarchive-openlibrary |
internetarchive/openlibrary | 8/10 | 6 | 80.0% |
release-zh-012-future-architect-vuls |
future-architect/vuls | 3/4 | 3 | 75.0% |
release-zh-014-internetarchive-openlibrary |
internetarchive/openlibrary | 6/10 | 3 | 60.0% |
release-zh-003-ansible-ansible |
ansible/ansible | 4/10 | 2 | 40.0% |
release-zh-007-flipt-io-flipt |
flipt-io/flipt | 4/10 | 3 | 40.0% |
release-zh-010-future-architect-vuls |
future-architect/vuls | 4/10 | 2 | 40.0% |
release-zh-018-qutebrowser-qutebrowser |
qutebrowser/qutebrowser | 0/9 | 0 | 0.0% |
release-zh-017-navidrome-navidrome |
navidrome/navidrome | 0/5 | 0 | 0.0% |
release-zh-002-ansible-ansible |
ansible/ansible | 1/10 | 1 | 10.0% |
release-zh-005-flipt-io-flipt |
flipt-io/flipt | 1/10 | 0 | 10.0% |