Qwen 3.6 plus
A lower-table result with a few useful bright spots: 44/151 tasks solved at least once, 18/151 solved in all three attempts, with the clearest wins around large Python/Django application repairs plus automation and configuration-management work. 这是一个排名靠后但仍有局部亮点的结果:151 题中至少一次解出 44 题,三次都解出 18 题;强项主要落在大型 Python/Django 应用修复以及自动化和配置管理类改动。
How to read this result可以这样读
- Qwen 3.6 plus is best read as volatile explorer: rank #17, 44 reached tasks, 18 stable solves.Qwen 3.6 plus 更适合读成探索型但波动较大:排名 #17,触达 44 题,稳定解出 18 题。
- Best suite signal: Open Library · release 015 at 7/10 (70.0%).最强 suite 信号:Open Library · release 015,7/10(70.0%)。
- Weakest visible area: Flipt · release 005 at 0/10 (0.0%).最弱可见区域:Flipt feature flag 服务 · release 005,0/10(0.0%)。
- The Qwen CLI setup is closer to a direct model readout: when it misses, the miss is less hidden behind orchestration.Qwen CLI 更接近直接读模型本身:它失败时,失败也较少被编排层遮住。
Qwen 3.6 plus is broad but volatile. It can touch 44/151 tasks, yet only 18 become 3/3 solves, so much of its value comes from retrying the same benchmark surface.
The closest family reference is Qwen 3.7 Max (1m) at rank #9. Compared with that row, this one is 3.19 points behind, with 3 fewer reached tasks and 7 fewer stable solves.
The suite split is asymmetric: Open Library · release 015 at 7/10 (70.0%) supplies the main body of wins, vuls · release 012 at 3/4 (75.0%) supplies the clean spike, and Flipt · release 005 at 0/10 (0.0%) is where that pattern stops. The Qwen CLI setup is closer to a direct model readout: when it misses, the miss is less hidden behind orchestration.
Read the bars as a volatility chart: Open Library · release 015 at 7/10 (70.0%) shows the upside, while the Pass^3 gap explains why the same row can feel much weaker on a single rerun.
The useful contrast is between Password lookup plugin ignores key=value parameters such as seed, resulting in non-deterministic output (ansible/ansible · solved 3/3) and Inconsistency in author identifier generation when comparing editions. (internetarchive/openlibrary · solved 2/3). The model reaches both kinds of problems, but only one becomes dependable.
The verifier audit keeps 95/95 solved attempts for Qwen 3.6 plus, so the interesting question is not score inflation; it is where the model repeatedly finds the same kind of patch.
The available audit keeps 95 of 95 initial solved attempts. Read this as a robustness check, especially when the audit sample is smaller than 453 attempts.当前可用复核保留了 95 次初始成功中的 95 次。这更适合作为稳健性检查,特别是在复核样本小于 453 次尝试时。
Use it when breadth matters more than deterministic replay. It can find openings around Open Library · release 015 at 7/10 (70.0%), but the 26-task reach gap says a second or third run may tell a different story. The 95/453 attempt score is best read as exploration bandwidth: 44 tasks are reachable, but many need retry luck.
Supporting suite table
| Suite | Repo | Solved | Pass^3 | Rate |
|---|---|---|---|---|
release-zh-012-future-architect-vuls |
future-architect/vuls | 3/4 | 3 | 75.0% |
release-zh-015-internetarchive-openlibrary |
internetarchive/openlibrary | 7/10 | 1 | 70.0% |
release-zh-014-internetarchive-openlibrary |
internetarchive/openlibrary | 6/10 | 2 | 60.0% |
release-zh-003-ansible-ansible |
ansible/ansible | 5/10 | 4 | 50.0% |
release-zh-013-internetarchive-openlibrary |
internetarchive/openlibrary | 4/10 | 0 | 40.0% |
release-zh-016-internetarchive-openlibrary |
internetarchive/openlibrary | 2/5 | 1 | 40.0% |
release-zh-005-flipt-io-flipt |
flipt-io/flipt | 0/10 | 0 | 0.0% |
release-zh-018-qutebrowser-qutebrowser |
qutebrowser/qutebrowser | 0/9 | 0 | 0.0% |
release-zh-017-navidrome-navidrome |
navidrome/navidrome | 0/5 | 0 | 0.0% |
release-zh-002-ansible-ansible |
ansible/ansible | 1/10 | 1 | 10.0% |
Qwen 3.6 plus 的特点是覆盖不窄但波动较大。它能至少一次摸到 44/151 题,但只有 18 题能做到 3/3,因此很大一部分价值来自重试。
最接近的同系参照是排名 #9 的 Qwen 3.7 Max (1m)。和它相比,这一行最终分低 3.19 分,触达题少 3 个,稳定题少 7 个。
suite 分布是不对称的:Open Library · release 015,7/10(70.0%)贡献主要胜利,vuls 漏洞扫描器 · release 012,3/4(75.0%)贡献最干净高点,而Flipt feature flag 服务 · release 005,0/10(0.0%)标出这种模式停止的地方。Qwen CLI 更接近直接读模型本身:它失败时,失败也较少被编排层遮住。
这张图更像波动率图:Open Library · release 015,7/10(70.0%)展示上限,而 Pass^3 落差解释了为什么单次重跑会显得弱很多。
最有用的对比是 Password lookup plugin 忽略 seed 等 key=value 参数,导致输出非确定性(ansible/ansible · 3 次中成功 3 次)和 比较 editions 时 author identifier 生成不一致。(internetarchive/openlibrary · 3 次中成功 2 次):模型都能触达,但只有前者变成可靠结果。
Qwen 3.6 plus 的复核保留了 95 次成功中的 95 次,所以重点不是分数膨胀,而是模型在哪些地方能反复找到同类补丁。
The available audit keeps 95 of 95 initial solved attempts. Read this as a robustness check, especially when the audit sample is smaller than 453 attempts.当前可用复核保留了 95 次初始成功中的 95 次。这更适合作为稳健性检查,特别是在复核样本小于 453 次尝试时。
当你更看重覆盖面而不是确定复现时,它更合适。它能在Open Library · release 015,7/10(70.0%)附近找到入口,但 26 题的覆盖-稳定差说明第二、第三次运行可能给出不同结果。95/453 的单次尝试成功数更像探索带宽:44 道题能触达,但很多仍需要重试运气。
支撑这个判断的 suite 表
| Suite | Repo | 解出 | Pass^3 | 通过率 |
|---|---|---|---|---|
release-zh-012-future-architect-vuls |
future-architect/vuls | 3/4 | 3 | 75.0% |
release-zh-015-internetarchive-openlibrary |
internetarchive/openlibrary | 7/10 | 1 | 70.0% |
release-zh-014-internetarchive-openlibrary |
internetarchive/openlibrary | 6/10 | 2 | 60.0% |
release-zh-003-ansible-ansible |
ansible/ansible | 5/10 | 4 | 50.0% |
release-zh-013-internetarchive-openlibrary |
internetarchive/openlibrary | 4/10 | 0 | 40.0% |
release-zh-016-internetarchive-openlibrary |
internetarchive/openlibrary | 2/5 | 1 | 40.0% |
release-zh-005-flipt-io-flipt |
flipt-io/flipt | 0/10 | 0 | 0.0% |
release-zh-018-qutebrowser-qutebrowser |
qutebrowser/qutebrowser | 0/9 | 0 | 0.0% |
release-zh-017-navidrome-navidrome |
navidrome/navidrome | 0/5 | 0 | 0.0% |
release-zh-002-ansible-ansible |
ansible/ansible | 1/10 | 1 | 10.0% |