Qwen 3.7 Plus (1m)
A competitive mid-table result with 43/151 tasks solved at least once and 28/151 solved in all three attempts; strongest around large Python/Django application repairs plus localized Go security-scanner changes. 这是一个有竞争力的中游结果:151 题中至少一次解出 43 题,三次都解出 28 题;强项主要落在大型 Python/Django 应用修复以及边界相对清楚的 Go 漏洞扫描器改动。
How to read this result可以这样读
- Qwen 3.7 Plus (1m) is best read as moderately stable: rank #11, 43 reached tasks, 28 stable solves.Qwen 3.7 Plus (1m) 更适合读成中等稳定型:排名 #11,触达 43 题,稳定解出 28 题。
- Best suite signal: Open Library · release 013 at 6/10 (60.0%).最强 suite 信号:Open Library · release 013,6/10(60.0%)。
- Weakest visible area: qutebrowser · release 018 at 0/9 (0.0%).最弱可见区域:qutebrowser 浏览器 · release 018,0/9(0.0%)。
- Qoder adds more workflow structure around the model, so its stable wins should be read as model-plus-shell behavior.Qoder 给模型外面加了更强的工作流结构,因此稳定胜利更适合读成 model-plus-shell 的组合效果。
Qwen 3.7 Plus (1m) is a moderately stable row around the #11 slot. The useful reading is not just the 31.21 score, but the split between 43 reached tasks and 28 stable solves.
The closest family reference is Qwen 3.7 Max (1m) at rank #9. Compared with that row, this one is 0.41 points behind, with 4 fewer reached tasks and 3 more stable solves.
The volume win is Open Library · release 013 at 6/10 (60.0%), while the cleanest pass-rate spike is vuls · release 012 at 3/4 (75.0%). The warning label is qutebrowser · release 018 at 0/9 (0.0%), so the contrast is not generic strength versus weakness; it is large Python/Django application repairs holding together better than browser/runtime integration around QtWebEngine behavior on this run. Qoder adds more workflow structure around the model, so its stable wins should be read as model-plus-shell behavior.
The Qoder shell tends to turn some model guesses into more disciplined patch attempts. That is why the suite profile should be compared with direct Qwen/OpenCode rows, not read as pure model capability.
Look at Forked output from ‘Display.display’ is unreliable and exposes shutdown deadlock risk and Feature Request: Add flag key to batch evaluation response as shell-behavior examples. The difference is not only model knowledge; it is whether the workflow keeps the patch disciplined enough to pass.
The verifier audit keeps 104/104 solved attempts for Qwen 3.7 Plus (1m), so the interesting question is not score inflation; it is where the model repeatedly finds the same kind of patch.
The available audit keeps 104 of 104 initial solved attempts. Read this as a robustness check, especially when the audit sample is smaller than 453 attempts.当前可用复核保留了 104 次初始成功中的 104 次。这更适合作为稳健性检查,特别是在复核样本小于 453 次尝试时。
For Qoder-style use, the interesting part is how the shell converts model guesses into patches. Compare Open Library · release 013 at 6/10 (60.0%) with qutebrowser · release 018 at 0/9 (0.0%) before attributing the result to the base model alone. The 104/452 attempt score should be read as model plus Qoder workflow, especially when comparing it with direct Qwen rows.
Supporting suite table
| Suite | Repo | Solved | Pass^3 | Rate |
|---|---|---|---|---|
release-zh-012-future-architect-vuls |
future-architect/vuls | 3/4 | 3 | 75.0% |
release-zh-013-internetarchive-openlibrary |
internetarchive/openlibrary | 6/10 | 3 | 60.0% |
release-zh-015-internetarchive-openlibrary |
internetarchive/openlibrary | 6/10 | 4 | 60.0% |
release-zh-003-ansible-ansible |
ansible/ansible | 4/10 | 4 | 40.0% |
release-zh-010-future-architect-vuls |
future-architect/vuls | 4/10 | 4 | 40.0% |
release-zh-016-internetarchive-openlibrary |
internetarchive/openlibrary | 2/5 | 1 | 40.0% |
release-zh-018-qutebrowser-qutebrowser |
qutebrowser/qutebrowser | 0/9 | 0 | 0.0% |
release-zh-017-navidrome-navidrome |
navidrome/navidrome | 0/5 | 0 | 0.0% |
release-zh-006-flipt-io-flipt |
flipt-io/flipt | 1/10 | 1 | 10.0% |
release-zh-008-flipt-io-flipt |
flipt-io/flipt | 1/10 | 1 | 10.0% |
Qwen 3.7 Plus (1m) 是一个排名 #11 附近的中等稳定型结果。它的重点不只是 31.21 分,而是 43 道触达题和 28 道稳定题之间的差距。
最接近的同系参照是排名 #9 的 Qwen 3.7 Max (1m)。和它相比,这一行最终分低 0.41 分,触达题少 4 个,稳定题多 3 个。
从数量看,主要胜利来自Open Library · release 013,6/10(60.0%);从通过率看,最干净的高点是vuls 漏洞扫描器 · release 012,3/4(75.0%)。需要警惕的是qutebrowser 浏览器 · release 018,0/9(0.0%),所以这里不是泛泛地说强弱项,而是大型 Python/Django 应用修复在这次运行中比围绕 QtWebEngine 行为的浏览器/runtime 集成更能闭环。Qoder 给模型外面加了更强的工作流结构,因此稳定胜利更适合读成 model-plus-shell 的组合效果。
Qoder shell 往往会把部分模型猜测压成更规整的补丁尝试。所以这张 suite 图更适合和 Qwen CLI / OpenCode 行对照,而不是当作纯模型能力。
可以把 "# 从 fork 进程调用 Display.display 的输出不可靠,并暴露 shutdown 死锁风险 和 Feature Request: Add flag key to batch evaluation response 当成 shell 行为样本:差异不只是模型懂不懂,也在于工作流能否把补丁约束到可通过状态。
Qwen 3.7 Plus (1m) 的复核保留了 104 次成功中的 104 次,所以重点不是分数膨胀,而是模型在哪些地方能反复找到同类补丁。
The available audit keeps 104 of 104 initial solved attempts. Read this as a robustness check, especially when the audit sample is smaller than 453 attempts.当前可用复核保留了 104 次初始成功中的 104 次。这更适合作为稳健性检查,特别是在复核样本小于 453 次尝试时。
对 Qoder-style 使用来说,重点是 shell 如何把模型猜测压成补丁。在把结果完全归因到底座模型之前,应先对照Open Library · release 013,6/10(60.0%)和qutebrowser 浏览器 · release 018,0/9(0.0%)。104/452 的单次尝试成功数应读成模型加 Qoder 工作流的结果,尤其要和直接 Qwen 行对照。
支撑这个判断的 suite 表
| Suite | Repo | 解出 | Pass^3 | 通过率 |
|---|---|---|---|---|
release-zh-012-future-architect-vuls |
future-architect/vuls | 3/4 | 3 | 75.0% |
release-zh-013-internetarchive-openlibrary |
internetarchive/openlibrary | 6/10 | 3 | 60.0% |
release-zh-015-internetarchive-openlibrary |
internetarchive/openlibrary | 6/10 | 4 | 60.0% |
release-zh-003-ansible-ansible |
ansible/ansible | 4/10 | 4 | 40.0% |
release-zh-010-future-architect-vuls |
future-architect/vuls | 4/10 | 4 | 40.0% |
release-zh-016-internetarchive-openlibrary |
internetarchive/openlibrary | 2/5 | 1 | 40.0% |
release-zh-018-qutebrowser-qutebrowser |
qutebrowser/qutebrowser | 0/9 | 0 | 0.0% |
release-zh-017-navidrome-navidrome |
navidrome/navidrome | 0/5 | 0 | 0.0% |
release-zh-006-flipt-io-flipt |
flipt-io/flipt | 1/10 | 1 | 10.0% |
release-zh-008-flipt-io-flipt |
flipt-io/flipt | 1/10 | 1 | 10.0% |