Step 3.7 flash
A lower-table result with a few useful bright spots: 43/151 tasks solved at least once, 18/151 solved in all three attempts, with the clearest wins around large Python/Django application repairs plus localized Go security-scanner changes. 这是一个排名靠后但仍有局部亮点的结果:151 题中至少一次解出 43 题,三次都解出 18 题;强项主要落在大型 Python/Django 应用修复以及边界相对清楚的 Go 漏洞扫描器改动。
How to read this result可以这样读
- Step 3.7 flash is best read as volatile explorer: rank #18, 43 reached tasks, 18 stable solves.Step 3.7 flash 更适合读成探索型但波动较大:排名 #18,触达 43 题,稳定解出 18 题。
- Best suite signal: Open Library · release 013 at 7/10 (70.0%).最强 suite 信号:Open Library · release 013,7/10(70.0%)。
- Weakest visible area: qutebrowser · release 018 at 0/9 (0.0%).最弱可见区域:qutebrowser 浏览器 · release 018,0/9(0.0%)。
- Because the agent shell is OpenCode, the result mostly exposes the underlying model's planning habits rather than a heavily opinionated workflow.因为 agent shell 是 OpenCode,这个结果更直接暴露底层模型的规划习惯,而不是强工作流包装后的表现。
Step 3.7 flash is broad but volatile. It can touch 43/151 tasks, yet only 18 become 3/3 solves, so much of its value comes from retrying the same benchmark surface.
The closest family reference is Step 3.5 flash 2603 at rank #26. Compared with that row, this one is 3.79 points ahead, with 6 more reached tasks and 5 more stable solves.
The volume win is Open Library · release 013 at 7/10 (70.0%), while the cleanest pass-rate spike is vuls · release 012 at 3/4 (75.0%). The warning label is qutebrowser · release 018 at 0/9 (0.0%), so the contrast is not generic strength versus weakness; it is large Python/Django application repairs holding together better than browser/runtime integration around QtWebEngine behavior on this run. Because the agent shell is OpenCode, the result mostly exposes the underlying model’s planning habits rather than a heavily opinionated workflow.
The important visual cue is the gap between high-reach suites and low Pass^3 counts. This model can often locate the neighborhood of the fix, but many patches do not survive three independent runs.
The useful contrast is between Enhance Kernel Version Handling for Debian Scans in Docker, or when the kernel version cannot be obtained (future-architect/vuls · solved 3/3) and Evaluation responses lack contextual reason for the result (flipt-io/flipt · solved 2/3). The model reaches both kinds of problems, but only one becomes dependable.
The audit changes how to read Step 3.7 flash: only 63% of initial solved attempts survive, with 34 rejected attempts, while the exported score field stays flat. Treat the wins as leads that need stricter confirmation.
The available audit keeps 57 of 91 initial solved attempts. Read this as a robustness check, especially when the audit sample is smaller than 453 attempts.当前可用复核保留了 91 次初始成功中的 57 次。这更适合作为稳健性检查,特别是在复核样本小于 453 次尝试时。
Use it when breadth matters more than deterministic replay. It can find openings around Open Library · release 013 at 7/10 (70.0%), but the 25-task reach gap says a second or third run may tell a different story. The 91/453 attempt score is best read as exploration bandwidth: 43 tasks are reachable, but many need retry luck.
Supporting suite table
| Suite | Repo | Solved | Pass^3 | Rate |
|---|---|---|---|---|
release-zh-012-future-architect-vuls |
future-architect/vuls | 3/4 | 2 | 75.0% |
release-zh-013-internetarchive-openlibrary |
internetarchive/openlibrary | 7/10 | 1 | 70.0% |
release-zh-015-internetarchive-openlibrary |
internetarchive/openlibrary | 5/10 | 2 | 50.0% |
release-zh-003-ansible-ansible |
ansible/ansible | 4/10 | 2 | 40.0% |
release-zh-010-future-architect-vuls |
future-architect/vuls | 4/10 | 3 | 40.0% |
release-zh-014-internetarchive-openlibrary |
internetarchive/openlibrary | 4/10 | 0 | 40.0% |
release-zh-018-qutebrowser-qutebrowser |
qutebrowser/qutebrowser | 0/9 | 0 | 0.0% |
release-zh-017-navidrome-navidrome |
navidrome/navidrome | 0/5 | 0 | 0.0% |
release-zh-005-flipt-io-flipt |
flipt-io/flipt | 1/10 | 0 | 10.0% |
release-zh-006-flipt-io-flipt |
flipt-io/flipt | 1/10 | 1 | 10.0% |
Step 3.7 flash 的特点是覆盖不窄但波动较大。它能至少一次摸到 43/151 题,但只有 18 题能做到 3/3,因此很大一部分价值来自重试。
最接近的同系参照是排名 #26 的 Step 3.5 flash 2603。和它相比,这一行最终分高 3.79 分,触达题多 6 个,稳定题多 5 个。
从数量看,主要胜利来自Open Library · release 013,7/10(70.0%);从通过率看,最干净的高点是vuls 漏洞扫描器 · release 012,3/4(75.0%)。需要警惕的是qutebrowser 浏览器 · release 018,0/9(0.0%),所以这里不是泛泛地说强弱项,而是大型 Python/Django 应用修复在这次运行中比围绕 QtWebEngine 行为的浏览器/runtime 集成更能闭环。因为 agent shell 是 OpenCode,这个结果更直接暴露底层模型的规划习惯,而不是强工作流包装后的表现。
这张图最重要的信号,是高触达 suite 和较低 Pass^3 之间的落差。模型经常能找到修复附近的位置,但很多补丁不能在三次独立运行中稳定复现。
最有用的对比是 增强 Docker 中 Debian 扫描的 Kernel 版本处理,或在无法获取 kernel 版本时(future-architect/vuls · 3 次中成功 3 次)和 Evaluation responses lack contextual reason for the result(flipt-io/flipt · 3 次中成功 2 次):模型都能触达,但只有前者变成可靠结果。
复核改变了 Step 3.7 flash 的读法:初始成功只有 63% 保留下来,34 次被剔除,但当前导出的分数字段没有变化。原始胜利更适合作为线索,需要更严格确认。
The available audit keeps 57 of 91 initial solved attempts. Read this as a robustness check, especially when the audit sample is smaller than 453 attempts.当前可用复核保留了 91 次初始成功中的 57 次。这更适合作为稳健性检查,特别是在复核样本小于 453 次尝试时。
当你更看重覆盖面而不是确定复现时,它更合适。它能在Open Library · release 013,7/10(70.0%)附近找到入口,但 25 题的覆盖-稳定差说明第二、第三次运行可能给出不同结果。91/453 的单次尝试成功数更像探索带宽:43 道题能触达,但很多仍需要重试运气。
支撑这个判断的 suite 表
| Suite | Repo | 解出 | Pass^3 | 通过率 |
|---|---|---|---|---|
release-zh-012-future-architect-vuls |
future-architect/vuls | 3/4 | 2 | 75.0% |
release-zh-013-internetarchive-openlibrary |
internetarchive/openlibrary | 7/10 | 1 | 70.0% |
release-zh-015-internetarchive-openlibrary |
internetarchive/openlibrary | 5/10 | 2 | 50.0% |
release-zh-003-ansible-ansible |
ansible/ansible | 4/10 | 2 | 40.0% |
release-zh-010-future-architect-vuls |
future-architect/vuls | 4/10 | 3 | 40.0% |
release-zh-014-internetarchive-openlibrary |
internetarchive/openlibrary | 4/10 | 0 | 40.0% |
release-zh-018-qutebrowser-qutebrowser |
qutebrowser/qutebrowser | 0/9 | 0 | 0.0% |
release-zh-017-navidrome-navidrome |
navidrome/navidrome | 0/5 | 0 | 0.0% |
release-zh-005-flipt-io-flipt |
flipt-io/flipt | 1/10 | 0 | 10.0% |
release-zh-006-flipt-io-flipt |
flipt-io/flipt | 1/10 | 1 | 10.0% |