Rank排名 21 Qoder Qwen

Qwen 3.5 plus (180k)

A lower-table result with a few useful bright spots: 41/151 tasks solved at least once, 14/151 solved in all three attempts, with the clearest wins around large Python/Django application repairs plus localized Go security-scanner changes. 这是一个排名靠后但仍有局部亮点的结果:151 题中至少一次解出 41 题,三次都解出 14 题;强项主要落在大型 Python/Django 应用修复以及边界相对清楚的 Go 漏洞扫描器改动。

qodercli 1.0.14 qwen3.5-plus-cp#context-window=180000 Updated更新 2026-06-18

How to read this result可以这样读

  • Qwen 3.5 plus (180k) is best read as wide but retry-sensitive: rank #21, 41 reached tasks, 14 stable solves.Qwen 3.5 plus (180k) 更适合读成覆盖不窄但依赖重试:排名 #21,触达 41 题,稳定解出 14 题。
  • Best suite signal: Open Library · release 013 at 6/10 (60.0%).最强 suite 信号:Open Library · release 013,6/10(60.0%)。
  • Weakest visible area: Flipt · release 006 at 0/10 (0.0%).最弱可见区域:Flipt feature flag 服务 · release 006,0/10(0.0%)。
  • Qoder adds more workflow structure around the model, so its stable wins should be read as model-plus-shell behavior.Qoder 给模型外面加了更强的工作流结构,因此稳定胜利更适合读成 model-plus-shell 的组合效果。

Qwen 3.5 plus (180k) is broad but volatile. It can touch 41/151 tasks, yet only 14 become 3/3 solves, so much of its value comes from retrying the same benchmark surface.

The closest family reference is Qwen 3.7 Max (1m) at rank #9. Compared with that row, this one is 5.71 points behind, with 6 fewer reached tasks and 11 fewer stable solves.

The result is easiest to understand as a three-point shape: volume at Open Library · release 013 at 6/10 (60.0%), efficiency at vuls · release 012 at 3/4 (75.0%), and resistance at Flipt · release 006 at 0/10 (0.0%). Qoder adds more workflow structure around the model, so its stable wins should be read as model-plus-shell behavior.

Selected high and low suites, grouped by pass-at-least-once rate.选取高分和低分 suite,按三次尝试至少解出一次的比例展示。
vuls · release 012vuls 漏洞扫描器 · release 012 3/4 · 75.0%

Best visible cluster for this row: 3/4 tasks reached.这一行最明显的强项簇:4 题中解出 3 题。

Open Library · release 013Open Library · release 013 6/10 · 60.0%
Ansible · release 003Ansible 自动化 · release 003 4/10 · 40.0%
vuls · release 010vuls 漏洞扫描器 · release 010 4/10 · 40.0%
vuls · release 011vuls 漏洞扫描器 · release 011 4/10 · 40.0%
Open Library · release 014Open Library · release 014 4/10 · 40.0%
Flipt · release 006Flipt feature flag 服务 · release 006 0/10 · 0.0%

Weak cluster: Go product plumbing across configuration, storage, and service APIs resisted this model-agent pairing.弱项簇:横跨配置、存储和服务 API 的 Go 产品工程对这个模型-agent 组合不友好。

qutebrowser · release 018qutebrowser 浏览器 · release 018 0/9 · 0.0%

Weak cluster: browser/runtime integration around QtWebEngine behavior resisted this model-agent pairing.弱项簇:围绕 QtWebEngine 行为的浏览器/runtime 集成对这个模型-agent 组合不友好。

Navidrome · release 017Navidrome 音乐服务 · release 017 0/5 · 0.0%

Weak cluster: Go service work with persistence and API behavior resisted this model-agent pairing.弱项簇:涉及持久化和 API 行为的 Go 服务改动对这个模型-agent 组合不友好。

Ansible · release 001Ansible 自动化 · release 001 1/10 · 10.0%

The bars say this row has search reach, not settled mastery. The model gets into the right repos often enough, but the repeatability line is still thin.

The useful contrast is between Feature Request: Add flag key to batch evaluation response (flipt-io/flipt · solved 3/3) and Inconsistency in author identifier generation when comparing editions. (internetarchive/openlibrary · solved 2/3). The model reaches both kinds of problems, but only one becomes dependable.

The audit trims 28 solved attempts from Qwen 3.5 plus (180k) but still keeps 66% of the solved set, so the suite shape remains useful even where individual wins are debatable.

Original harness result vs verifier-backed audit sample原始 harness 结果 vs verifier-backed 复核样本
55 of 83 headline successes survived strict re-verification. 83 次初始成功里,55 次通过了更严格的复核。

The available audit keeps 55 of 83 initial solved attempts. Read this as a robustness check, especially when the audit sample is smaller than 453 attempts.当前可用复核保留了 83 次初始成功中的 55 次。这更适合作为稳健性检查,特别是在复核样本小于 453 次尝试时。

55 verifier-backed复核通过 28 strict rejected严格拒绝
25.91 25.91 +0.00 points+0.00 分

Use it when breadth matters more than deterministic replay. It can find openings around Open Library · release 013 at 6/10 (60.0%), but the 27-task reach gap says a second or third run may tell a different story. The 83/453 attempt score is best read as exploration bandwidth: 41 tasks are reachable, but many need retry luck.

Supporting suite table
Suite Repo Solved Pass^3 Rate
release-zh-012-future-architect-vuls future-architect/vuls 3/4 2 75.0%
release-zh-013-internetarchive-openlibrary internetarchive/openlibrary 6/10 1 60.0%
release-zh-003-ansible-ansible ansible/ansible 4/10 2 40.0%
release-zh-010-future-architect-vuls future-architect/vuls 4/10 2 40.0%
release-zh-011-future-architect-vuls future-architect/vuls 4/10 0 40.0%
release-zh-014-internetarchive-openlibrary internetarchive/openlibrary 4/10 0 40.0%
release-zh-006-flipt-io-flipt flipt-io/flipt 0/10 0 0.0%
release-zh-018-qutebrowser-qutebrowser qutebrowser/qutebrowser 0/9 0 0.0%
release-zh-017-navidrome-navidrome navidrome/navidrome 0/5 0 0.0%
release-zh-001-ansible-ansible ansible/ansible 1/10 1 10.0%

Qwen 3.5 plus (180k) 的特点是覆盖不窄但波动较大。它能至少一次摸到 41/151 题,但只有 14 题能做到 3/3,因此很大一部分价值来自重试。

最接近的同系参照是排名 #9 的 Qwen 3.7 Max (1m)。和它相比,这一行最终分低 5.71 分,触达题少 6 个,稳定题少 11 个。

这个结果最容易读成三点形状:数量在Open Library · release 013,6/10(60.0%),效率在vuls 漏洞扫描器 · release 012,3/4(75.0%),阻力在Flipt feature flag 服务 · release 006,0/10(0.0%)。Qoder 给模型外面加了更强的工作流结构,因此稳定胜利更适合读成 model-plus-shell 的组合效果。

Selected high and low suites, grouped by pass-at-least-once rate.选取高分和低分 suite,按三次尝试至少解出一次的比例展示。
vuls · release 012vuls 漏洞扫描器 · release 012 3/4 · 75.0%

Best visible cluster for this row: 3/4 tasks reached.这一行最明显的强项簇:4 题中解出 3 题。

Open Library · release 013Open Library · release 013 6/10 · 60.0%
Ansible · release 003Ansible 自动化 · release 003 4/10 · 40.0%
vuls · release 010vuls 漏洞扫描器 · release 010 4/10 · 40.0%
vuls · release 011vuls 漏洞扫描器 · release 011 4/10 · 40.0%
Open Library · release 014Open Library · release 014 4/10 · 40.0%
Flipt · release 006Flipt feature flag 服务 · release 006 0/10 · 0.0%

Weak cluster: Go product plumbing across configuration, storage, and service APIs resisted this model-agent pairing.弱项簇:横跨配置、存储和服务 API 的 Go 产品工程对这个模型-agent 组合不友好。

qutebrowser · release 018qutebrowser 浏览器 · release 018 0/9 · 0.0%

Weak cluster: browser/runtime integration around QtWebEngine behavior resisted this model-agent pairing.弱项簇:围绕 QtWebEngine 行为的浏览器/runtime 集成对这个模型-agent 组合不友好。

Navidrome · release 017Navidrome 音乐服务 · release 017 0/5 · 0.0%

Weak cluster: Go service work with persistence and API behavior resisted this model-agent pairing.弱项簇:涉及持久化和 API 行为的 Go 服务改动对这个模型-agent 组合不友好。

Ansible · release 001Ansible 自动化 · release 001 1/10 · 10.0%

这些柱子说明这一行有搜索触达,不等于已经掌握。模型经常能进入正确代码库,但可重复通过的线仍然偏细。

最有用的对比是 Feature Request: Add flag key to batch evaluation response(flipt-io/flipt · 3 次中成功 3 次)和 比较 editions 时 author identifier 生成不一致。(internetarchive/openlibrary · 3 次中成功 2 次):模型都能触达,但只有前者变成可靠结果。

复核从 Qwen 3.5 plus (180k) 中剔除了 28 次成功,但仍保留 66% 的成功集合,因此即便个别胜利有争议,suite 形状仍然有参考价值。

Original harness result vs verifier-backed audit sample原始 harness 结果 vs verifier-backed 复核样本
55 of 83 headline successes survived strict re-verification. 83 次初始成功里,55 次通过了更严格的复核。

The available audit keeps 55 of 83 initial solved attempts. Read this as a robustness check, especially when the audit sample is smaller than 453 attempts.当前可用复核保留了 83 次初始成功中的 55 次。这更适合作为稳健性检查,特别是在复核样本小于 453 次尝试时。

55 verifier-backed复核通过 28 strict rejected严格拒绝
25.91 25.91 +0.00 points+0.00 分

当你更看重覆盖面而不是确定复现时,它更合适。它能在Open Library · release 013,6/10(60.0%)附近找到入口,但 27 题的覆盖-稳定差说明第二、第三次运行可能给出不同结果。83/453 的单次尝试成功数更像探索带宽:41 道题能触达,但很多仍需要重试运气。

支撑这个判断的 suite 表
Suite Repo 解出 Pass^3 通过率
release-zh-012-future-architect-vuls future-architect/vuls 3/4 2 75.0%
release-zh-013-internetarchive-openlibrary internetarchive/openlibrary 6/10 1 60.0%
release-zh-003-ansible-ansible ansible/ansible 4/10 2 40.0%
release-zh-010-future-architect-vuls future-architect/vuls 4/10 2 40.0%
release-zh-011-future-architect-vuls future-architect/vuls 4/10 0 40.0%
release-zh-014-internetarchive-openlibrary internetarchive/openlibrary 4/10 0 40.0%
release-zh-006-flipt-io-flipt flipt-io/flipt 0/10 0 0.0%
release-zh-018-qutebrowser-qutebrowser qutebrowser/qutebrowser 0/9 0 0.0%
release-zh-017-navidrome-navidrome navidrome/navidrome 0/5 0 0.0%
release-zh-001-ansible-ansible ansible/ansible 1/10 1 10.0%