Qwen 3.6 plus (180k)
A lower-table result with a few useful bright spots: 38/151 tasks solved at least once, 16/151 solved in all three attempts, with the clearest wins around large Python/Django application repairs plus localized Go security-scanner changes. 这是一个排名靠后但仍有局部亮点的结果:151 题中至少一次解出 38 题,三次都解出 16 题;强项主要落在大型 Python/Django 应用修复以及边界相对清楚的 Go 漏洞扫描器改动。
How to read this result可以这样读
- Qwen 3.6 plus (180k) is best read as volatile explorer: rank #22, 38 reached tasks, 16 stable solves.Qwen 3.6 plus (180k) 更适合读成探索型但波动较大:排名 #22,触达 38 题,稳定解出 16 题。
- Best suite signal: Open Library · release 013 at 7/10 (70.0%).最强 suite 信号:Open Library · release 013,7/10(70.0%)。
- Weakest visible area: Flipt · release 005 at 0/10 (0.0%).最弱可见区域:Flipt feature flag 服务 · release 005,0/10(0.0%)。
- Qoder adds more workflow structure around the model, so its stable wins should be read as model-plus-shell behavior.Qoder 给模型外面加了更强的工作流结构,因此稳定胜利更适合读成 model-plus-shell 的组合效果。
Qwen 3.6 plus (180k) is a volatile explorer row around the #22 slot. The useful reading is not just the 25.81 score, but the split between 38 reached tasks and 16 stable solves.
The closest family reference is Qwen 3.7 Max (1m) at rank #9. Compared with that row, this one is 5.81 points behind, with 9 fewer reached tasks and 9 fewer stable solves.
The volume win is Open Library · release 013 at 7/10 (70.0%), while the cleanest pass-rate spike is vuls · release 012 at 3/4 (75.0%). The warning label is Flipt · release 005 at 0/10 (0.0%), so the contrast is not generic strength versus weakness; it is large Python/Django application repairs holding together better than Go product plumbing across configuration, storage, and service APIs on this run. Qoder adds more workflow structure around the model, so its stable wins should be read as model-plus-shell behavior.
The Qoder shell tends to turn some model guesses into more disciplined patch attempts. That is why the suite profile should be compared with direct Qwen/OpenCode rows, not read as pure model capability.
Look at pip module fails when executable and virtualenv are unset and no pip binary is found and Enhance Kernel Version Handling for Debian Scans in Docker, or when the kernel version cannot be obtained as shell-behavior examples. The difference is not only model knowledge; it is whether the workflow keeps the patch disciplined enough to pass.
The audit trims 27 solved attempts from Qwen 3.6 plus (180k) but still keeps 66% of the solved set, so the suite shape remains useful even where individual wins are debatable.
The available audit keeps 53 of 80 initial solved attempts. Read this as a robustness check, especially when the audit sample is smaller than 453 attempts.当前可用复核保留了 80 次初始成功中的 53 次。这更适合作为稳健性检查,特别是在复核样本小于 453 次尝试时。
For Qoder-style use, the interesting part is how the shell converts model guesses into patches. Compare Open Library · release 013 at 7/10 (70.0%) with Flipt · release 005 at 0/10 (0.0%) before attributing the result to the base model alone. The 80/453 attempt score should be read as model plus Qoder workflow, especially when comparing it with direct Qwen rows.
Supporting suite table
| Suite | Repo | Solved | Pass^3 | Rate |
|---|---|---|---|---|
release-zh-012-future-architect-vuls |
future-architect/vuls | 3/4 | 3 | 75.0% |
release-zh-013-internetarchive-openlibrary |
internetarchive/openlibrary | 7/10 | 1 | 70.0% |
release-zh-003-ansible-ansible |
ansible/ansible | 5/10 | 2 | 50.0% |
release-zh-010-future-architect-vuls |
future-architect/vuls | 4/10 | 2 | 40.0% |
release-zh-015-internetarchive-openlibrary |
internetarchive/openlibrary | 4/10 | 1 | 40.0% |
release-zh-004-ansible-ansible |
ansible/ansible | 1/3 | 1 | 33.3% |
release-zh-005-flipt-io-flipt |
flipt-io/flipt | 0/10 | 0 | 0.0% |
release-zh-018-qutebrowser-qutebrowser |
qutebrowser/qutebrowser | 0/9 | 0 | 0.0% |
release-zh-002-ansible-ansible |
ansible/ansible | 1/10 | 1 | 10.0% |
release-zh-006-flipt-io-flipt |
flipt-io/flipt | 1/10 | 1 | 10.0% |
Qwen 3.6 plus (180k) 是一个排名 #22 附近的探索型但波动较大结果。它的重点不只是 25.81 分,而是 38 道触达题和 16 道稳定题之间的差距。
最接近的同系参照是排名 #9 的 Qwen 3.7 Max (1m)。和它相比,这一行最终分低 5.81 分,触达题少 9 个,稳定题少 9 个。
从数量看,主要胜利来自Open Library · release 013,7/10(70.0%);从通过率看,最干净的高点是vuls 漏洞扫描器 · release 012,3/4(75.0%)。需要警惕的是Flipt feature flag 服务 · release 005,0/10(0.0%),所以这里不是泛泛地说强弱项,而是大型 Python/Django 应用修复在这次运行中比横跨配置、存储和服务 API 的 Go 产品工程更能闭环。Qoder 给模型外面加了更强的工作流结构,因此稳定胜利更适合读成 model-plus-shell 的组合效果。
Qoder shell 往往会把部分模型猜测压成更规整的补丁尝试。所以这张 suite 图更适合和 Qwen CLI / OpenCode 行对照,而不是当作纯模型能力。
可以把 当 executable 和 virtualenv 未设置且找不到 pip binary 时,pip module 会失败 和 增强 Docker 中 Debian 扫描的 Kernel 版本处理,或在无法获取 kernel 版本时 当成 shell 行为样本:差异不只是模型懂不懂,也在于工作流能否把补丁约束到可通过状态。
复核从 Qwen 3.6 plus (180k) 中剔除了 27 次成功,但仍保留 66% 的成功集合,因此即便个别胜利有争议,suite 形状仍然有参考价值。
The available audit keeps 53 of 80 initial solved attempts. Read this as a robustness check, especially when the audit sample is smaller than 453 attempts.当前可用复核保留了 80 次初始成功中的 53 次。这更适合作为稳健性检查,特别是在复核样本小于 453 次尝试时。
对 Qoder-style 使用来说,重点是 shell 如何把模型猜测压成补丁。在把结果完全归因到底座模型之前,应先对照Open Library · release 013,7/10(70.0%)和Flipt feature flag 服务 · release 005,0/10(0.0%)。80/453 的单次尝试成功数应读成模型加 Qoder 工作流的结果,尤其要和直接 Qwen 行对照。
支撑这个判断的 suite 表
| Suite | Repo | 解出 | Pass^3 | 通过率 |
|---|---|---|---|---|
release-zh-012-future-architect-vuls |
future-architect/vuls | 3/4 | 3 | 75.0% |
release-zh-013-internetarchive-openlibrary |
internetarchive/openlibrary | 7/10 | 1 | 70.0% |
release-zh-003-ansible-ansible |
ansible/ansible | 5/10 | 2 | 50.0% |
release-zh-010-future-architect-vuls |
future-architect/vuls | 4/10 | 2 | 40.0% |
release-zh-015-internetarchive-openlibrary |
internetarchive/openlibrary | 4/10 | 1 | 40.0% |
release-zh-004-ansible-ansible |
ansible/ansible | 1/3 | 1 | 33.3% |
release-zh-005-flipt-io-flipt |
flipt-io/flipt | 0/10 | 0 | 0.0% |
release-zh-018-qutebrowser-qutebrowser |
qutebrowser/qutebrowser | 0/9 | 0 | 0.0% |
release-zh-002-ansible-ansible |
ansible/ansible | 1/10 | 1 | 10.0% |
release-zh-006-flipt-io-flipt |
flipt-io/flipt | 1/10 | 1 | 10.0% |