DeepSeek v4 pro (max)
A competitive mid-table result with 48/151 tasks solved at least once and 28/151 solved in all three attempts; strongest around large Python/Django application repairs plus localized Go security-scanner changes. 这是一个有竞争力的中游结果:151 题中至少一次解出 48 题,三次都解出 28 题;强项主要落在大型 Python/Django 应用修复以及边界相对清楚的 Go 漏洞扫描器改动。
How to read this result可以这样读
- DeepSeek v4 pro (max) is best read as moderately stable: rank #7, 48 reached tasks, 28 stable solves.DeepSeek v4 pro (max) 更适合读成中等稳定型:排名 #7,触达 48 题,稳定解出 28 题。
- Best suite signal: Open Library · release 013 at 8/10 (80.0%).最强 suite 信号:Open Library · release 013,8/10(80.0%)。
- Weakest visible area: qutebrowser · release 018 at 0/9 (0.0%).最弱可见区域:qutebrowser 浏览器 · release 018,0/9(0.0%)。
- Because the agent shell is OpenCode, the result mostly exposes the underlying model's planning habits rather than a heavily opinionated workflow.因为 agent shell 是 OpenCode,这个结果更直接暴露底层模型的规划习惯,而不是强工作流包装后的表现。
DeepSeek v4 pro (max) is a moderately stable row around the #7 slot. The useful reading is not just the 32.95 score, but the split between 48 reached tasks and 28 stable solves.
The closest family reference is DeepSeek v4 flash (max) at rank #15. Compared with that row, this one is 3.96 points ahead, with 4 more reached tasks and 8 more stable solves.
Most of the positive signal concentrates in Open Library · release 013 at 8/10 (80.0%). The opposing read is qutebrowser · release 018 at 0/9 (0.0%), which keeps the row from looking like a generalist. Because the agent shell is OpenCode, the result mostly exposes the underlying model’s planning habits rather than a heavily opinionated workflow.
For this row, the suite bars are a contrast tool. The distance between Open Library · release 013 at 8/10 (80.0%) and qutebrowser · release 018 at 0/9 (0.0%) is the model’s practical boundary.
The examples keep the middle-band story honest: Add Reading-Log Counts to Solr Work Documents is the upside, Reversible Password Encryption in Navidrome is the failure surface, and the page should be read between those two poles.
The verifier audit keeps 117/117 solved attempts for DeepSeek v4 pro (max), so the interesting question is not score inflation; it is where the model repeatedly finds the same kind of patch.
The available audit keeps 117 of 117 initial solved attempts. Read this as a robustness check, especially when the audit sample is smaller than 453 attempts.当前可用复核保留了 117 次初始成功中的 117 次。这更适合作为稳健性检查,特别是在复核样本小于 453 次尝试时。
In practice, read it through the gap between Open Library · release 013 at 8/10 (80.0%) and qutebrowser · release 018 at 0/9 (0.0%). That gap is more actionable than the rank because it says which repo shape gets coherent patches. The 117/453 attempt score is the backdrop; the article above is about which parts of that score are repeatable enough to matter.
Supporting suite table
| Suite | Repo | Solved | Pass^3 | Rate |
|---|---|---|---|---|
release-zh-013-internetarchive-openlibrary |
internetarchive/openlibrary | 8/10 | 5 | 80.0% |
release-zh-012-future-architect-vuls |
future-architect/vuls | 3/4 | 3 | 75.0% |
release-zh-014-internetarchive-openlibrary |
internetarchive/openlibrary | 6/10 | 3 | 60.0% |
release-zh-015-internetarchive-openlibrary |
internetarchive/openlibrary | 6/10 | 2 | 60.0% |
release-zh-003-ansible-ansible |
ansible/ansible | 4/10 | 3 | 40.0% |
release-zh-007-flipt-io-flipt |
flipt-io/flipt | 4/10 | 2 | 40.0% |
release-zh-018-qutebrowser-qutebrowser |
qutebrowser/qutebrowser | 0/9 | 0 | 0.0% |
release-zh-017-navidrome-navidrome |
navidrome/navidrome | 0/5 | 0 | 0.0% |
release-zh-005-flipt-io-flipt |
flipt-io/flipt | 1/10 | 0 | 10.0% |
release-zh-006-flipt-io-flipt |
flipt-io/flipt | 1/10 | 1 | 10.0% |
DeepSeek v4 pro (max) 是一个排名 #7 附近的中等稳定型结果。它的重点不只是 32.95 分,而是 48 道触达题和 28 道稳定题之间的差距。
最接近的同系参照是排名 #15 的 DeepSeek v4 flash (max)。和它相比,这一行最终分高 3.96 分,触达题多 4 个,稳定题多 8 个。
正面信号大多集中在Open Library · release 013,8/10(80.0%)。反向读法是qutebrowser 浏览器 · release 018,0/9(0.0%),它让这一行看起来不像通用型。因为 agent shell 是 OpenCode,这个结果更直接暴露底层模型的规划习惯,而不是强工作流包装后的表现。
对这一行来说,suite 柱更像对比工具。Open Library · release 013,8/10(80.0%)和qutebrowser 浏览器 · release 018,0/9(0.0%)之间的距离,就是模型的实用边界。
这些案例让中段模型画像更具体:向 Solr Work Documents 添加 Reading-Log Counts 是上限,Navidrome 中的可逆密码加密 是失败面,这页应该在两者之间读。
DeepSeek v4 pro (max) 的复核保留了 117 次成功中的 117 次,所以重点不是分数膨胀,而是模型在哪些地方能反复找到同类补丁。
The available audit keeps 117 of 117 initial solved attempts. Read this as a robustness check, especially when the audit sample is smaller than 453 attempts.当前可用复核保留了 117 次初始成功中的 117 次。这更适合作为稳健性检查,特别是在复核样本小于 453 次尝试时。
实际选择时,更应该通过Open Library · release 013,8/10(80.0%)和qutebrowser 浏览器 · release 018,0/9(0.0%)之间的落差来读它。这个落差比分数排名更可操作,因为它说明哪类代码库更容易得到连贯补丁。117/453 的单次尝试成功数只是背景;上面的文章重点是哪些部分足够可重复、值得当成能力看。
支撑这个判断的 suite 表
| Suite | Repo | 解出 | Pass^3 | 通过率 |
|---|---|---|---|---|
release-zh-013-internetarchive-openlibrary |
internetarchive/openlibrary | 8/10 | 5 | 80.0% |
release-zh-012-future-architect-vuls |
future-architect/vuls | 3/4 | 3 | 75.0% |
release-zh-014-internetarchive-openlibrary |
internetarchive/openlibrary | 6/10 | 3 | 60.0% |
release-zh-015-internetarchive-openlibrary |
internetarchive/openlibrary | 6/10 | 2 | 60.0% |
release-zh-003-ansible-ansible |
ansible/ansible | 4/10 | 3 | 40.0% |
release-zh-007-flipt-io-flipt |
flipt-io/flipt | 4/10 | 2 | 40.0% |
release-zh-018-qutebrowser-qutebrowser |
qutebrowser/qutebrowser | 0/9 | 0 | 0.0% |
release-zh-017-navidrome-navidrome |
navidrome/navidrome | 0/5 | 0 | 0.0% |
release-zh-005-flipt-io-flipt |
flipt-io/flipt | 1/10 | 0 | 10.0% |
release-zh-006-flipt-io-flipt |
flipt-io/flipt | 1/10 | 1 | 10.0% |