Rank排名 16 OpenCode Zhipu GLM

GLM 5 turbo

A lower-table result with a few useful bright spots: 44/151 tasks solved at least once, 19/151 solved in all three attempts, with the clearest wins around large Python/Django application repairs plus localized Go security-scanner changes. 这是一个排名靠后但仍有局部亮点的结果:151 题中至少一次解出 44 题,三次都解出 19 题;强项主要落在大型 Python/Django 应用修复以及边界相对清楚的 Go 漏洞扫描器改动。

opencode-cli 1.14.32 zai-coding-plan/glm-5-turbo Updated更新 2026-06-18

How to read this result可以这样读

  • GLM 5 turbo is best read as volatile explorer: rank #16, 44 reached tasks, 19 stable solves.GLM 5 turbo 更适合读成探索型但波动较大:排名 #16,触达 44 题,稳定解出 19 题。
  • Best suite signal: Open Library · release 013 at 6/10 (60.0%).最强 suite 信号:Open Library · release 013,6/10(60.0%)。
  • Weakest visible area: qutebrowser · release 018 at 0/9 (0.0%).最弱可见区域:qutebrowser 浏览器 · release 018,0/9(0.0%)。
  • Because the agent shell is OpenCode, the result mostly exposes the underlying model's planning habits rather than a heavily opinionated workflow.因为 agent shell 是 OpenCode,这个结果更直接暴露底层模型的规划习惯,而不是强工作流包装后的表现。

GLM 5 turbo is a volatile explorer row around the #16 slot. The useful reading is not just the 28.95 score, but the split between 44 reached tasks and 19 stable solves.

The closest family reference is GLM 5.2 at rank #1. Compared with that row, this one is 8.63 points behind, with 13 fewer reached tasks and 17 fewer stable solves.

The suite split is asymmetric: Open Library · release 013 at 6/10 (60.0%) supplies the main body of wins, vuls · release 012 at 3/4 (75.0%) supplies the clean spike, and qutebrowser · release 018 at 0/9 (0.0%) is where that pattern stops. Because the agent shell is OpenCode, the result mostly exposes the underlying model’s planning habits rather than a heavily opinionated workflow.

Selected high and low suites, grouped by pass-at-least-once rate.选取高分和低分 suite,按三次尝试至少解出一次的比例展示。
vuls · release 012vuls 漏洞扫描器 · release 012 3/4 · 75.0%

Best visible cluster for this row: 3/4 tasks reached.这一行最明显的强项簇:4 题中解出 3 题。

Open Library · release 013Open Library · release 013 6/10 · 60.0%
Open Library · release 014Open Library · release 014 6/10 · 60.0%
Open Library · release 015Open Library · release 015 5/10 · 50.0%
Flipt · release 007Flipt feature flag 服务 · release 007 4/10 · 40.0%
vuls · release 010vuls 漏洞扫描器 · release 010 4/10 · 40.0%
qutebrowser · release 018qutebrowser 浏览器 · release 018 0/9 · 0.0%

Weak cluster: browser/runtime integration around QtWebEngine behavior resisted this model-agent pairing.弱项簇:围绕 QtWebEngine 行为的浏览器/runtime 集成对这个模型-agent 组合不友好。

Navidrome · release 017Navidrome 音乐服务 · release 017 0/5 · 0.0%

Weak cluster: Go service work with persistence and API behavior resisted this model-agent pairing.弱项簇:涉及持久化和 API 行为的 Go 服务改动对这个模型-agent 组合不友好。

Flipt · release 005Flipt feature flag 服务 · release 005 1/10 · 10.0%
Flipt · release 006Flipt feature flag 服务 · release 006 1/10 · 10.0%

The chart is not trying to crown a single strength; it shows how quickly the row falls from Open Library · release 013 at 6/10 (60.0%) to qutebrowser · release 018 at 0/9 (0.0%).

The examples keep the middle-band story honest: psrp connection plugin accepts undocumented extras, causing ambiguous and inconsistent configuration. is the upside, vuls report fails to parse legacy scan results due to incompatible listenPorts field format is the failure surface, and the page should be read between those two poles.

The verifier audit keeps 99/99 solved attempts for GLM 5 turbo, so the interesting question is not score inflation; it is where the model repeatedly finds the same kind of patch.

Original harness result vs verifier-backed audit sample原始 harness 结果 vs verifier-backed 复核样本
99 of 99 headline successes survived strict re-verification. 99 次初始成功里,99 次通过了更严格的复核。

The available audit keeps 99 of 99 initial solved attempts. Read this as a robustness check, especially when the audit sample is smaller than 453 attempts.当前可用复核保留了 99 次初始成功中的 99 次。这更适合作为稳健性检查,特别是在复核样本小于 453 次尝试时。

99 verifier-backed复核通过 0 strict rejected严格拒绝
28.95 28.95 +0.00 points+0.00 分

In practice, read it through the gap between Open Library · release 013 at 6/10 (60.0%) and qutebrowser · release 018 at 0/9 (0.0%). That gap is more actionable than the rank because it says which repo shape gets coherent patches. The 99/453 attempt score is the backdrop; the article above is about which parts of that score are repeatable enough to matter.

Supporting suite table
Suite Repo Solved Pass^3 Rate
release-zh-012-future-architect-vuls future-architect/vuls 3/4 2 75.0%
release-zh-013-internetarchive-openlibrary internetarchive/openlibrary 6/10 4 60.0%
release-zh-014-internetarchive-openlibrary internetarchive/openlibrary 6/10 3 60.0%
release-zh-015-internetarchive-openlibrary internetarchive/openlibrary 5/10 1 50.0%
release-zh-007-flipt-io-flipt flipt-io/flipt 4/10 2 40.0%
release-zh-010-future-architect-vuls future-architect/vuls 4/10 0 40.0%
release-zh-018-qutebrowser-qutebrowser qutebrowser/qutebrowser 0/9 0 0.0%
release-zh-017-navidrome-navidrome navidrome/navidrome 0/5 0 0.0%
release-zh-005-flipt-io-flipt flipt-io/flipt 1/10 0 10.0%
release-zh-006-flipt-io-flipt flipt-io/flipt 1/10 0 10.0%

GLM 5 turbo 是一个排名 #16 附近的探索型但波动较大结果。它的重点不只是 28.95 分,而是 44 道触达题和 19 道稳定题之间的差距。

最接近的同系参照是排名 #1 的 GLM 5.2。和它相比,这一行最终分低 8.63 分,触达题少 13 个,稳定题少 17 个。

suite 分布是不对称的:Open Library · release 013,6/10(60.0%)贡献主要胜利,vuls 漏洞扫描器 · release 012,3/4(75.0%)贡献最干净高点,而qutebrowser 浏览器 · release 018,0/9(0.0%)标出这种模式停止的地方。因为 agent shell 是 OpenCode,这个结果更直接暴露底层模型的规划习惯,而不是强工作流包装后的表现。

Selected high and low suites, grouped by pass-at-least-once rate.选取高分和低分 suite,按三次尝试至少解出一次的比例展示。
vuls · release 012vuls 漏洞扫描器 · release 012 3/4 · 75.0%

Best visible cluster for this row: 3/4 tasks reached.这一行最明显的强项簇:4 题中解出 3 题。

Open Library · release 013Open Library · release 013 6/10 · 60.0%
Open Library · release 014Open Library · release 014 6/10 · 60.0%
Open Library · release 015Open Library · release 015 5/10 · 50.0%
Flipt · release 007Flipt feature flag 服务 · release 007 4/10 · 40.0%
vuls · release 010vuls 漏洞扫描器 · release 010 4/10 · 40.0%
qutebrowser · release 018qutebrowser 浏览器 · release 018 0/9 · 0.0%

Weak cluster: browser/runtime integration around QtWebEngine behavior resisted this model-agent pairing.弱项簇:围绕 QtWebEngine 行为的浏览器/runtime 集成对这个模型-agent 组合不友好。

Navidrome · release 017Navidrome 音乐服务 · release 017 0/5 · 0.0%

Weak cluster: Go service work with persistence and API behavior resisted this model-agent pairing.弱项簇:涉及持久化和 API 行为的 Go 服务改动对这个模型-agent 组合不友好。

Flipt · release 005Flipt feature flag 服务 · release 005 1/10 · 10.0%
Flipt · release 006Flipt feature flag 服务 · release 006 1/10 · 10.0%

这张图不是为了给单一强项加冕,而是展示这一行从Open Library · release 013,6/10(60.0%)滑到qutebrowser 浏览器 · release 018,0/9(0.0%)有多快。

这些案例让中段模型画像更具体:psrp connection plugin 接受未文档化 extras,导致配置含糊且不一致。 是上限,vuls report fails to parse legacy scan results due to incompatible listenPorts field format 是失败面,这页应该在两者之间读。

GLM 5 turbo 的复核保留了 99 次成功中的 99 次,所以重点不是分数膨胀,而是模型在哪些地方能反复找到同类补丁。

Original harness result vs verifier-backed audit sample原始 harness 结果 vs verifier-backed 复核样本
99 of 99 headline successes survived strict re-verification. 99 次初始成功里,99 次通过了更严格的复核。

The available audit keeps 99 of 99 initial solved attempts. Read this as a robustness check, especially when the audit sample is smaller than 453 attempts.当前可用复核保留了 99 次初始成功中的 99 次。这更适合作为稳健性检查,特别是在复核样本小于 453 次尝试时。

99 verifier-backed复核通过 0 strict rejected严格拒绝
28.95 28.95 +0.00 points+0.00 分

实际选择时,更应该通过Open Library · release 013,6/10(60.0%)和qutebrowser 浏览器 · release 018,0/9(0.0%)之间的落差来读它。这个落差比分数排名更可操作,因为它说明哪类代码库更容易得到连贯补丁。99/453 的单次尝试成功数只是背景;上面的文章重点是哪些部分足够可重复、值得当成能力看。

支撑这个判断的 suite 表
Suite Repo 解出 Pass^3 通过率
release-zh-012-future-architect-vuls future-architect/vuls 3/4 2 75.0%
release-zh-013-internetarchive-openlibrary internetarchive/openlibrary 6/10 4 60.0%
release-zh-014-internetarchive-openlibrary internetarchive/openlibrary 6/10 3 60.0%
release-zh-015-internetarchive-openlibrary internetarchive/openlibrary 5/10 1 50.0%
release-zh-007-flipt-io-flipt flipt-io/flipt 4/10 2 40.0%
release-zh-010-future-architect-vuls future-architect/vuls 4/10 0 40.0%
release-zh-018-qutebrowser-qutebrowser qutebrowser/qutebrowser 0/9 0 0.0%
release-zh-017-navidrome-navidrome navidrome/navidrome 0/5 0 0.0%
release-zh-005-flipt-io-flipt flipt-io/flipt 1/10 0 10.0%
release-zh-006-flipt-io-flipt flipt-io/flipt 1/10 0 10.0%