Qwen 3.5 plus
A competitive mid-table result with 51/151 tasks solved at least once and 22/151 solved in all three attempts; strongest around large Python/Django application repairs plus automation and configuration-management work. 这是一个有竞争力的中游结果:151 题中至少一次解出 51 题,三次都解出 22 题;强项主要落在大型 Python/Django 应用修复以及自动化和配置管理类改动。
How to read this result可以这样读
- Qwen 3.5 plus is best read as volatile explorer: rank #10, 51 reached tasks, 22 stable solves.Qwen 3.5 plus 更适合读成探索型但波动较大:排名 #10,触达 51 题,稳定解出 22 题。
- Best suite signal: Open Library · release 013 at 8/10 (80.0%).最强 suite 信号:Open Library · release 013,8/10(80.0%)。
- Weakest visible area: qutebrowser · release 018 at 0/9 (0.0%).最弱可见区域:qutebrowser 浏览器 · release 018,0/9(0.0%)。
- The Qwen CLI setup is closer to a direct model readout: when it misses, the miss is less hidden behind orchestration.Qwen CLI 更接近直接读模型本身:它失败时,失败也较少被编排层遮住。
Qwen 3.5 plus is a volatile explorer row around the #10 slot. The useful reading is not just the 31.39 score, but the split between 51 reached tasks and 22 stable solves.
The closest family reference is Qwen 3.7 Max (1m) at rank #9. Compared with that row, this one is 0.23 points behind, with 4 more reached tasks and 3 fewer stable solves.
The volume win is Open Library · release 013 at 8/10 (80.0%), while the cleanest pass-rate spike is vuls · release 012 at 4/4 (100.0%). The warning label is qutebrowser · release 018 at 0/9 (0.0%), so the contrast is not generic strength versus weakness; it is large Python/Django application repairs holding together better than browser/runtime integration around QtWebEngine behavior on this run. The Qwen CLI setup is closer to a direct model readout: when it misses, the miss is less hidden behind orchestration.
Because this is the Qwen CLI row, the suite shape is unusually direct: strong and weak bars are closer to model behavior than to shell behavior.
Avoid double calculation of loops and delegate_to in TaskExecutor is the clean read of what the model can do directly; Configuration Logic for Qt Arguments and Environment Setup Is Overloaded and Hard to Maintain is the corresponding negative read, with less agent machinery to hide the miss.
The audit changes how to read Qwen 3.5 plus: only 57% of initial solved attempts survive, with 45 rejected attempts, while the exported score field stays flat. Treat the wins as leads that need stricter confirmation.
The available audit keeps 60 of 105 initial solved attempts. Read this as a robustness check, especially when the audit sample is smaller than 453 attempts.当前可用复核保留了 105 次初始成功中的 60 次。这更适合作为稳健性检查,特别是在复核样本小于 453 次尝试时。
As a direct CLI row, it is most valuable for reading the model itself: Open Library · release 013 at 8/10 (80.0%) is the positive sample, and qutebrowser · release 018 at 0/9 (0.0%) is the boundary. The 105/453 attempt score is useful because there is less shell behavior between the model and the verifier result.
Supporting suite table
| Suite | Repo | Solved | Pass^3 | Rate |
|---|---|---|---|---|
release-zh-012-future-architect-vuls |
future-architect/vuls | 4/4 | 2 | 100.0% |
release-zh-013-internetarchive-openlibrary |
internetarchive/openlibrary | 8/10 | 3 | 80.0% |
release-zh-015-internetarchive-openlibrary |
internetarchive/openlibrary | 5/10 | 2 | 50.0% |
release-zh-001-ansible-ansible |
ansible/ansible | 4/10 | 2 | 40.0% |
release-zh-002-ansible-ansible |
ansible/ansible | 4/10 | 2 | 40.0% |
release-zh-003-ansible-ansible |
ansible/ansible | 4/10 | 2 | 40.0% |
release-zh-018-qutebrowser-qutebrowser |
qutebrowser/qutebrowser | 0/9 | 0 | 0.0% |
release-zh-005-flipt-io-flipt |
flipt-io/flipt | 1/10 | 0 | 10.0% |
release-zh-008-flipt-io-flipt |
flipt-io/flipt | 1/10 | 1 | 10.0% |
release-zh-006-flipt-io-flipt |
flipt-io/flipt | 2/10 | 1 | 20.0% |
Qwen 3.5 plus 是一个排名 #10 附近的探索型但波动较大结果。它的重点不只是 31.39 分,而是 51 道触达题和 22 道稳定题之间的差距。
最接近的同系参照是排名 #9 的 Qwen 3.7 Max (1m)。和它相比,这一行最终分低 0.23 分,触达题多 4 个,稳定题少 3 个。
从数量看,主要胜利来自Open Library · release 013,8/10(80.0%);从通过率看,最干净的高点是vuls 漏洞扫描器 · release 012,4/4(100.0%)。需要警惕的是qutebrowser 浏览器 · release 018,0/9(0.0%),所以这里不是泛泛地说强弱项,而是大型 Python/Django 应用修复在这次运行中比围绕 QtWebEngine 行为的浏览器/runtime 集成更能闭环。Qwen CLI 更接近直接读模型本身:它失败时,失败也较少被编排层遮住。
因为这是 Qwen CLI 行,suite 形状相对直接:高低柱更接近模型本身,而不是 shell 工作流的效果。
避免在 TaskExecutor 中重复计算 loops 和 delegate_to 是模型直接能力的正面样本;Qt 参数和环境初始化配置逻辑过载且难以维护 是相应的反面样本,中间没有太多 agent 编排可以掩盖失败。
复核改变了 Qwen 3.5 plus 的读法:初始成功只有 57% 保留下来,45 次被剔除,但当前导出的分数字段没有变化。原始胜利更适合作为线索,需要更严格确认。
The available audit keeps 60 of 105 initial solved attempts. Read this as a robustness check, especially when the audit sample is smaller than 453 attempts.当前可用复核保留了 105 次初始成功中的 60 次。这更适合作为稳健性检查,特别是在复核样本小于 453 次尝试时。
作为直接 CLI 行,它最有价值的是读模型本身:Open Library · release 013,8/10(80.0%)是正面样本,qutebrowser 浏览器 · release 018,0/9(0.0%)是边界。105/453 的单次尝试成功数有价值,是因为模型和 verifier 结果之间隔着的 shell 行为更少。
支撑这个判断的 suite 表
| Suite | Repo | 解出 | Pass^3 | 通过率 |
|---|---|---|---|---|
release-zh-012-future-architect-vuls |
future-architect/vuls | 4/4 | 2 | 100.0% |
release-zh-013-internetarchive-openlibrary |
internetarchive/openlibrary | 8/10 | 3 | 80.0% |
release-zh-015-internetarchive-openlibrary |
internetarchive/openlibrary | 5/10 | 2 | 50.0% |
release-zh-001-ansible-ansible |
ansible/ansible | 4/10 | 2 | 40.0% |
release-zh-002-ansible-ansible |
ansible/ansible | 4/10 | 2 | 40.0% |
release-zh-003-ansible-ansible |
ansible/ansible | 4/10 | 2 | 40.0% |
release-zh-018-qutebrowser-qutebrowser |
qutebrowser/qutebrowser | 0/9 | 0 | 0.0% |
release-zh-005-flipt-io-flipt |
flipt-io/flipt | 1/10 | 0 | 10.0% |
release-zh-008-flipt-io-flipt |
flipt-io/flipt | 1/10 | 1 | 10.0% |
release-zh-006-flipt-io-flipt |
flipt-io/flipt | 2/10 | 1 | 20.0% |