KAT Coder Pro v2
A lower-table result with a few useful bright spots: 39/151 tasks solved at least once, 21/151 solved in all three attempts, with the clearest wins around large Python/Django application repairs plus localized Go security-scanner changes. 这是一个排名靠后但仍有局部亮点的结果:151 题中至少一次解出 39 题,三次都解出 21 题;强项主要落在大型 Python/Django 应用修复以及边界相对清楚的 Go 漏洞扫描器改动。
How to read this result可以这样读
- KAT Coder Pro v2 is best read as moderately stable: rank #19, 39 reached tasks, 21 stable solves.KAT Coder Pro v2 更适合读成中等稳定型:排名 #19,触达 39 题,稳定解出 21 题。
- Best suite signal: Open Library · release 013 at 5/10 (50.0%).最强 suite 信号:Open Library · release 013,5/10(50.0%)。
- Weakest visible area: Flipt · release 005 at 0/10 (0.0%).最弱可见区域:Flipt feature flag 服务 · release 005,0/10(0.0%)。
- Because the agent shell is OpenCode, the result mostly exposes the underlying model's planning habits rather than a heavily opinionated workflow.因为 agent shell 是 OpenCode,这个结果更直接暴露底层模型的规划习惯,而不是强工作流包装后的表现。
KAT Coder Pro v2 is a moderately stable row around the #19 slot. The useful reading is not just the 27.97 score, but the split between 39 reached tasks and 21 stable solves.
The closest family reference is KAT Coder Pro v2 at rank #28. Compared with that row, this one is 4.50 points ahead, with 10 more reached tasks and 4 more stable solves.
The volume win is Open Library · release 013 at 5/10 (50.0%), while the cleanest pass-rate spike is vuls · release 012 at 3/4 (75.0%). The warning label is Flipt · release 005 at 0/10 (0.0%), so the contrast is not generic strength versus weakness; it is large Python/Django application repairs holding together better than Go product plumbing across configuration, storage, and service APIs on this run. Because the agent shell is OpenCode, the result mostly exposes the underlying model’s planning habits rather than a heavily opinionated workflow.
This is a middle-band profile: the useful signal is the slope between Open Library · release 013 at 5/10 (50.0%) and Flipt · release 005 at 0/10 (0.0%).
The examples keep the middle-band story honest: Scan results miss Package URL (PURL) information in library output is the upside, Alpine Linux vulnerability detection incorrectly handles source vs binary packages is the failure surface, and the page should be read between those two poles.
The audit changes how to read KAT Coder Pro v2: only 60% of initial solved attempts survive, with 36 rejected attempts, while the exported score field stays flat. Treat the wins as leads that need stricter confirmation.
The available audit keeps 54 of 90 initial solved attempts. Read this as a robustness check, especially when the audit sample is smaller than 453 attempts.当前可用复核保留了 90 次初始成功中的 54 次。这更适合作为稳健性检查,特别是在复核样本小于 453 次尝试时。
In practice, read it through the gap between Open Library · release 013 at 5/10 (50.0%) and Flipt · release 005 at 0/10 (0.0%). That gap is more actionable than the rank because it says which repo shape gets coherent patches. The 90/453 attempt score is the backdrop; the article above is about which parts of that score are repeatable enough to matter.
Supporting suite table
| Suite | Repo | Solved | Pass^3 | Rate |
|---|---|---|---|---|
release-zh-012-future-architect-vuls |
future-architect/vuls | 3/4 | 3 | 75.0% |
release-zh-013-internetarchive-openlibrary |
internetarchive/openlibrary | 5/10 | 2 | 50.0% |
release-zh-015-internetarchive-openlibrary |
internetarchive/openlibrary | 5/10 | 2 | 50.0% |
release-zh-010-future-architect-vuls |
future-architect/vuls | 4/10 | 2 | 40.0% |
release-zh-011-future-architect-vuls |
future-architect/vuls | 4/10 | 0 | 40.0% |
release-zh-016-internetarchive-openlibrary |
internetarchive/openlibrary | 2/5 | 1 | 40.0% |
release-zh-005-flipt-io-flipt |
flipt-io/flipt | 0/10 | 0 | 0.0% |
release-zh-018-qutebrowser-qutebrowser |
qutebrowser/qutebrowser | 0/9 | 0 | 0.0% |
release-zh-017-navidrome-navidrome |
navidrome/navidrome | 0/5 | 0 | 0.0% |
release-zh-002-ansible-ansible |
ansible/ansible | 1/10 | 0 | 10.0% |
KAT Coder Pro v2 是一个排名 #19 附近的中等稳定型结果。它的重点不只是 27.97 分,而是 39 道触达题和 21 道稳定题之间的差距。
最接近的同系参照是排名 #28 的 KAT Coder Pro v2。和它相比,这一行最终分高 4.50 分,触达题多 10 个,稳定题多 4 个。
从数量看,主要胜利来自Open Library · release 013,5/10(50.0%);从通过率看,最干净的高点是vuls 漏洞扫描器 · release 012,3/4(75.0%)。需要警惕的是Flipt feature flag 服务 · release 005,0/10(0.0%),所以这里不是泛泛地说强弱项,而是大型 Python/Django 应用修复在这次运行中比横跨配置、存储和服务 API 的 Go 产品工程更能闭环。因为 agent shell 是 OpenCode,这个结果更直接暴露底层模型的规划习惯,而不是强工作流包装后的表现。
这是一个中段模型画像:真正有用的信号,是Open Library · release 013,5/10(50.0%)到Flipt feature flag 服务 · release 005,0/10(0.0%)之间的落差。
这些案例让中段模型画像更具体:Library 输出中的扫描结果缺少 Package URL (PURL) 信息 是上限,Alpine Linux vulnerability detection 错误处理 source packages 与 binary packages 是失败面,这页应该在两者之间读。
复核改变了 KAT Coder Pro v2 的读法:初始成功只有 60% 保留下来,36 次被剔除,但当前导出的分数字段没有变化。原始胜利更适合作为线索,需要更严格确认。
The available audit keeps 54 of 90 initial solved attempts. Read this as a robustness check, especially when the audit sample is smaller than 453 attempts.当前可用复核保留了 90 次初始成功中的 54 次。这更适合作为稳健性检查,特别是在复核样本小于 453 次尝试时。
实际选择时,更应该通过Open Library · release 013,5/10(50.0%)和Flipt feature flag 服务 · release 005,0/10(0.0%)之间的落差来读它。这个落差比分数排名更可操作,因为它说明哪类代码库更容易得到连贯补丁。90/453 的单次尝试成功数只是背景;上面的文章重点是哪些部分足够可重复、值得当成能力看。
支撑这个判断的 suite 表
| Suite | Repo | 解出 | Pass^3 | 通过率 |
|---|---|---|---|---|
release-zh-012-future-architect-vuls |
future-architect/vuls | 3/4 | 3 | 75.0% |
release-zh-013-internetarchive-openlibrary |
internetarchive/openlibrary | 5/10 | 2 | 50.0% |
release-zh-015-internetarchive-openlibrary |
internetarchive/openlibrary | 5/10 | 2 | 50.0% |
release-zh-010-future-architect-vuls |
future-architect/vuls | 4/10 | 2 | 40.0% |
release-zh-011-future-architect-vuls |
future-architect/vuls | 4/10 | 0 | 40.0% |
release-zh-016-internetarchive-openlibrary |
internetarchive/openlibrary | 2/5 | 1 | 40.0% |
release-zh-005-flipt-io-flipt |
flipt-io/flipt | 0/10 | 0 | 0.0% |
release-zh-018-qutebrowser-qutebrowser |
qutebrowser/qutebrowser | 0/9 | 0 | 0.0% |
release-zh-017-navidrome-navidrome |
navidrome/navidrome | 0/5 | 0 | 0.0% |
release-zh-002-ansible-ansible |
ansible/ansible | 1/10 | 0 | 10.0% |