Rank排名 28 OpenCode StreamLake

KAT Coder Pro v2

A lower-table result with a few useful bright spots: 29/151 tasks solved at least once, 17/151 solved in all three attempts, with the clearest wins around automation and configuration-management work plus localized Go security-scanner changes. 这是一个排名靠后但仍有局部亮点的结果:151 题中至少一次解出 29 题,三次都解出 17 题;强项主要落在自动化和配置管理类改动以及边界相对清楚的 Go 漏洞扫描器改动。

opencode-cli 1.14.32 streamlake/kat-coder-pro-v2 Updated更新 2026-06-18

How to read this result可以这样读

  • KAT Coder Pro v2 is best read as moderately stable: rank #28, 29 reached tasks, 17 stable solves.KAT Coder Pro v2 更适合读成中等稳定型:排名 #28,触达 29 题,稳定解出 17 题。
  • Best suite signal: Ansible · release 001 at 4/10 (40.0%).最强 suite 信号:Ansible 自动化 · release 001,4/10(40.0%)。
  • Weakest visible area: Open Library · release 013 at 0/10 (0.0%).最弱可见区域:Open Library · release 013,0/10(0.0%)。
  • Because the agent shell is OpenCode, the result mostly exposes the underlying model's planning habits rather than a heavily opinionated workflow.因为 agent shell 是 OpenCode,这个结果更直接暴露底层模型的规划习惯,而不是强工作流包装后的表现。

KAT Coder Pro v2 is a moderately stable row around the #28 slot. The useful reading is not just the 23.47 score, but the split between 29 reached tasks and 17 stable solves.

The closest family reference is KAT Coder Pro v2 at rank #19. Compared with that row, this one is 4.50 points behind, with 10 fewer reached tasks and 4 fewer stable solves.

The volume win is Ansible · release 001 at 4/10 (40.0%), while the cleanest pass-rate spike is vuls · release 012 at 3/4 (75.0%). The warning label is Open Library · release 013 at 0/10 (0.0%), so the contrast is not generic strength versus weakness; it is automation and configuration-management work holding together better than large Python/Django application repairs on this run. Because the agent shell is OpenCode, the result mostly exposes the underlying model’s planning habits rather than a heavily opinionated workflow.

Selected high and low suites, grouped by pass-at-least-once rate.选取高分和低分 suite,按三次尝试至少解出一次的比例展示。
vuls · release 012vuls 漏洞扫描器 · release 012 3/4 · 75.0%

Best visible cluster for this row: 3/4 tasks reached.这一行最明显的强项簇:4 题中解出 3 题。

Ansible · release 001Ansible 自动化 · release 001 4/10 · 40.0%
Ansible · release 002Ansible 自动化 · release 002 4/10 · 40.0%
Ansible · release 003Ansible 自动化 · release 003 4/10 · 40.0%
vuls · release 011vuls 漏洞扫描器 · release 011 4/10 · 40.0%
Ansible · release 004Ansible 自动化 · release 004 1/3 · 33.3%
Open Library · release 013Open Library · release 013 0/10 · 0.0%

Weak cluster: large Python/Django application repairs resisted this model-agent pairing.弱项簇:大型 Python/Django 应用修复对这个模型-agent 组合不友好。

Open Library · release 014Open Library · release 014 0/10 · 0.0%

Weak cluster: large Python/Django application repairs resisted this model-agent pairing.弱项簇:大型 Python/Django 应用修复对这个模型-agent 组合不友好。

Open Library · release 015Open Library · release 015 0/10 · 0.0%

Weak cluster: large Python/Django application repairs resisted this model-agent pairing.弱项簇:大型 Python/Django 应用修复对这个模型-agent 组合不友好。

qutebrowser · release 018qutebrowser 浏览器 · release 018 0/9 · 0.0%

Weak cluster: browser/runtime integration around QtWebEngine behavior resisted this model-agent pairing.弱项簇:围绕 QtWebEngine 行为的浏览器/runtime 集成对这个模型-agent 组合不友好。

At this end of the table, the weak bars are more informative than the wins. They show which task families break first when the model-agent loop runs out of reliable planning.

At this rank, Host blocking does not apply to subdomains when only the parent domain is listed matters as much as the wins. It shows the task shape where the model-agent loop fails before it can produce a meaningful verifier-backed patch.

The verifier audit keeps 68/68 solved attempts for KAT Coder Pro v2, so the interesting question is not score inflation; it is where the model repeatedly finds the same kind of patch.

Original harness result vs verifier-backed audit sample原始 harness 结果 vs verifier-backed 复核样本
68 of 68 headline successes survived strict re-verification. 68 次初始成功里,68 次通过了更严格的复核。

The available audit keeps 68 of 68 initial solved attempts. Read this as a robustness check, especially when the audit sample is smaller than 453 attempts.当前可用复核保留了 68 次初始成功中的 68 次。这更适合作为稳健性检查,特别是在复核样本小于 453 次尝试时。

68 verifier-backed复核通过 0 strict rejected严格拒绝
23.47 23.47 +0.00 points+0.00 分

This row is more useful as a failure map than as a default choice. Look at Open Library · release 013 at 0/10 (0.0%) first: it shows the task shape where the loop loses traction. With 68/453 solved attempts, the page is most useful for seeing where the agent loop breaks before it becomes a dependable option.

Supporting suite table
Suite Repo Solved Pass^3 Rate
release-zh-012-future-architect-vuls future-architect/vuls 3/4 3 75.0%
release-zh-001-ansible-ansible ansible/ansible 4/10 2 40.0%
release-zh-002-ansible-ansible ansible/ansible 4/10 2 40.0%
release-zh-003-ansible-ansible ansible/ansible 4/10 3 40.0%
release-zh-011-future-architect-vuls future-architect/vuls 4/10 1 40.0%
release-zh-004-ansible-ansible ansible/ansible 1/3 0 33.3%
release-zh-013-internetarchive-openlibrary internetarchive/openlibrary 0/10 0 0.0%
release-zh-014-internetarchive-openlibrary internetarchive/openlibrary 0/10 0 0.0%
release-zh-015-internetarchive-openlibrary internetarchive/openlibrary 0/10 0 0.0%
release-zh-018-qutebrowser-qutebrowser qutebrowser/qutebrowser 0/9 0 0.0%

KAT Coder Pro v2 是一个排名 #28 附近的中等稳定型结果。它的重点不只是 23.47 分,而是 29 道触达题和 17 道稳定题之间的差距。

最接近的同系参照是排名 #19 的 KAT Coder Pro v2。和它相比,这一行最终分低 4.50 分,触达题少 10 个,稳定题少 4 个。

从数量看,主要胜利来自Ansible 自动化 · release 001,4/10(40.0%);从通过率看,最干净的高点是vuls 漏洞扫描器 · release 012,3/4(75.0%)。需要警惕的是Open Library · release 013,0/10(0.0%),所以这里不是泛泛地说强弱项,而是自动化和配置管理类改动在这次运行中比大型 Python/Django 应用修复更能闭环。因为 agent shell 是 OpenCode,这个结果更直接暴露底层模型的规划习惯,而不是强工作流包装后的表现。

Selected high and low suites, grouped by pass-at-least-once rate.选取高分和低分 suite,按三次尝试至少解出一次的比例展示。
vuls · release 012vuls 漏洞扫描器 · release 012 3/4 · 75.0%

Best visible cluster for this row: 3/4 tasks reached.这一行最明显的强项簇:4 题中解出 3 题。

Ansible · release 001Ansible 自动化 · release 001 4/10 · 40.0%
Ansible · release 002Ansible 自动化 · release 002 4/10 · 40.0%
Ansible · release 003Ansible 自动化 · release 003 4/10 · 40.0%
vuls · release 011vuls 漏洞扫描器 · release 011 4/10 · 40.0%
Ansible · release 004Ansible 自动化 · release 004 1/3 · 33.3%
Open Library · release 013Open Library · release 013 0/10 · 0.0%

Weak cluster: large Python/Django application repairs resisted this model-agent pairing.弱项簇:大型 Python/Django 应用修复对这个模型-agent 组合不友好。

Open Library · release 014Open Library · release 014 0/10 · 0.0%

Weak cluster: large Python/Django application repairs resisted this model-agent pairing.弱项簇:大型 Python/Django 应用修复对这个模型-agent 组合不友好。

Open Library · release 015Open Library · release 015 0/10 · 0.0%

Weak cluster: large Python/Django application repairs resisted this model-agent pairing.弱项簇:大型 Python/Django 应用修复对这个模型-agent 组合不友好。

qutebrowser · release 018qutebrowser 浏览器 · release 018 0/9 · 0.0%

Weak cluster: browser/runtime integration around QtWebEngine behavior resisted this model-agent pairing.弱项簇:围绕 QtWebEngine 行为的浏览器/runtime 集成对这个模型-agent 组合不友好。

在榜单后段,低柱子往往比胜利更有信息量。它们说明模型-agent 循环在哪些任务家族上最先失去可靠规划。

在这个排名段,当只在主域名中列出父域时,基于 hosts 的屏蔽方式不会对其子域名生效。请修复该问题,使父域名规则能正确影响相关子域,同时仍然遵守现有的内容屏蔽开关和白名单逻辑。 和成功案例一样重要。它说明模型-agent 循环在哪种任务形态上还没形成有效 verifier-backed patch。

KAT Coder Pro v2 的复核保留了 68 次成功中的 68 次,所以重点不是分数膨胀,而是模型在哪些地方能反复找到同类补丁。

Original harness result vs verifier-backed audit sample原始 harness 结果 vs verifier-backed 复核样本
68 of 68 headline successes survived strict re-verification. 68 次初始成功里,68 次通过了更严格的复核。

The available audit keeps 68 of 68 initial solved attempts. Read this as a robustness check, especially when the audit sample is smaller than 453 attempts.当前可用复核保留了 68 次初始成功中的 68 次。这更适合作为稳健性检查,特别是在复核样本小于 453 次尝试时。

68 verifier-backed复核通过 0 strict rejected严格拒绝
23.47 23.47 +0.00 points+0.00 分

这一行更适合作为失败地图,而不是默认选择。先看Open Library · release 013,0/10(0.0%):它展示了模型-agent 循环最容易失去抓手的任务形态。在 453 次尝试中只成功 68 次时,这页最有价值的是看 agent loop 在哪里先断掉。

支撑这个判断的 suite 表
Suite Repo 解出 Pass^3 通过率
release-zh-012-future-architect-vuls future-architect/vuls 3/4 3 75.0%
release-zh-001-ansible-ansible ansible/ansible 4/10 2 40.0%
release-zh-002-ansible-ansible ansible/ansible 4/10 2 40.0%
release-zh-003-ansible-ansible ansible/ansible 4/10 3 40.0%
release-zh-011-future-architect-vuls future-architect/vuls 4/10 1 40.0%
release-zh-004-ansible-ansible ansible/ansible 1/3 0 33.3%
release-zh-013-internetarchive-openlibrary internetarchive/openlibrary 0/10 0 0.0%
release-zh-014-internetarchive-openlibrary internetarchive/openlibrary 0/10 0 0.0%
release-zh-015-internetarchive-openlibrary internetarchive/openlibrary 0/10 0 0.0%
release-zh-018-qutebrowser-qutebrowser qutebrowser/qutebrowser 0/9 0 0.0%