Rank排名 31 OpenCode StepFun

Step 3.5 flash

A lower-table result with a few useful bright spots: 21/151 tasks solved at least once, 8/151 solved in all three attempts, with the clearest wins around automation and configuration-management work plus Go product plumbing across configuration, storage, and service APIs. 这是一个排名靠后但仍有局部亮点的结果:151 题中至少一次解出 21 题,三次都解出 8 题;强项主要落在自动化和配置管理类改动以及横跨配置、存储和服务 API 的 Go 产品工程。

opencode-cli 1.14.32 stepfun/step-3.5-flash Updated更新 2026-06-18

How to read this result可以这样读

  • Step 3.5 flash is best read as volatile explorer: rank #31, 21 reached tasks, 8 stable solves.Step 3.5 flash 更适合读成探索型但波动较大:排名 #31,触达 21 题,稳定解出 8 题。
  • Best suite signal: Ansible · release 003 at 4/10 (40.0%).最强 suite 信号:Ansible 自动化 · release 003,4/10(40.0%)。
  • Weakest visible area: vuls · release 011 at 0/10 (0.0%).最弱可见区域:vuls 漏洞扫描器 · release 011,0/10(0.0%)。
  • Because the agent shell is OpenCode, the result mostly exposes the underlying model's planning habits rather than a heavily opinionated workflow.因为 agent shell 是 OpenCode,这个结果更直接暴露底层模型的规划习惯,而不是强工作流包装后的表现。

Step 3.5 flash is a volatile explorer row around the #31 slot. The useful reading is not just the 17.04 score, but the split between 21 reached tasks and 8 stable solves.

The closest family reference is Step 3.7 flash at rank #18. Compared with that row, this one is 10.97 points behind, with 22 fewer reached tasks and 10 fewer stable solves.

The volume win is Ansible · release 003 at 4/10 (40.0%), while the cleanest pass-rate spike is vuls · release 012 at 2/4 (50.0%). The warning label is vuls · release 011 at 0/10 (0.0%), so the contrast is not generic strength versus weakness; it is automation and configuration-management work holding together better than localized Go security-scanner changes on this run. Because the agent shell is OpenCode, the result mostly exposes the underlying model’s planning habits rather than a heavily opinionated workflow.

Selected high and low suites, grouped by pass-at-least-once rate.选取高分和低分 suite,按三次尝试至少解出一次的比例展示。
vuls · release 012vuls 漏洞扫描器 · release 012 2/4 · 50.0%

Best visible cluster for this row: 2/4 tasks reached.这一行最明显的强项簇:4 题中解出 2 题。

Ansible · release 003Ansible 自动化 · release 003 4/10 · 40.0%
Ansible · release 004Ansible 自动化 · release 004 1/3 · 33.3%
Flipt · release 007Flipt feature flag 服务 · release 007 3/10 · 30.0%
vuls · release 010vuls 漏洞扫描器 · release 010 3/10 · 30.0%
Ansible · release 001Ansible 自动化 · release 001 2/10 · 20.0%
vuls · release 011vuls 漏洞扫描器 · release 011 0/10 · 0.0%

Weak cluster: localized Go security-scanner changes resisted this model-agent pairing.弱项簇:边界相对清楚的 Go 漏洞扫描器改动对这个模型-agent 组合不友好。

Open Library · release 014Open Library · release 014 0/10 · 0.0%

Weak cluster: large Python/Django application repairs resisted this model-agent pairing.弱项簇:大型 Python/Django 应用修复对这个模型-agent 组合不友好。

Open Library · release 015Open Library · release 015 0/10 · 0.0%

Weak cluster: large Python/Django application repairs resisted this model-agent pairing.弱项簇:大型 Python/Django 应用修复对这个模型-agent 组合不友好。

qutebrowser · release 018qutebrowser 浏览器 · release 018 0/9 · 0.0%

Weak cluster: browser/runtime integration around QtWebEngine behavior resisted this model-agent pairing.弱项簇:围绕 QtWebEngine 行为的浏览器/runtime 集成对这个模型-agent 组合不友好。

At this end of the table, the weak bars are more informative than the wins. They show which task families break first when the model-agent loop runs out of reliable planning.

At this rank, Redis cache backend cannot connect to TLS-enabled Redis servers without additional configuration options matters as much as the wins. It shows the task shape where the model-agent loop fails before it can produce a meaningful verifier-backed patch.

The audit trims 3 solved attempts from Step 3.5 flash but still keeps 93% of the solved set, so the suite shape remains useful even where individual wins are debatable.

Original harness result vs verifier-backed audit sample原始 harness 结果 vs verifier-backed 复核样本
39 of 42 headline successes survived strict re-verification. 42 次初始成功里,39 次通过了更严格的复核。

The available audit keeps 39 of 42 initial solved attempts. Read this as a robustness check, especially when the audit sample is smaller than 453 attempts.当前可用复核保留了 42 次初始成功中的 39 次。这更适合作为稳健性检查,特别是在复核样本小于 453 次尝试时。

39 verifier-backed复核通过 3 strict rejected严格拒绝
17.04 17.04 +0.00 points+0.00 分

This row is more useful as a failure map than as a default choice. Look at vuls · release 011 at 0/10 (0.0%) first: it shows the task shape where the loop loses traction. With 42/453 solved attempts, the page is most useful for seeing where the agent loop breaks before it becomes a dependable option.

Supporting suite table
Suite Repo Solved Pass^3 Rate
release-zh-012-future-architect-vuls future-architect/vuls 2/4 2 50.0%
release-zh-003-ansible-ansible ansible/ansible 4/10 1 40.0%
release-zh-004-ansible-ansible ansible/ansible 1/3 0 33.3%
release-zh-007-flipt-io-flipt flipt-io/flipt 3/10 1 30.0%
release-zh-010-future-architect-vuls future-architect/vuls 3/10 1 30.0%
release-zh-001-ansible-ansible ansible/ansible 2/10 1 20.0%
release-zh-011-future-architect-vuls future-architect/vuls 0/10 0 0.0%
release-zh-014-internetarchive-openlibrary internetarchive/openlibrary 0/10 0 0.0%
release-zh-015-internetarchive-openlibrary internetarchive/openlibrary 0/10 0 0.0%
release-zh-018-qutebrowser-qutebrowser qutebrowser/qutebrowser 0/9 0 0.0%

Step 3.5 flash 是一个排名 #31 附近的探索型但波动较大结果。它的重点不只是 17.04 分,而是 21 道触达题和 8 道稳定题之间的差距。

最接近的同系参照是排名 #18 的 Step 3.7 flash。和它相比,这一行最终分低 10.97 分,触达题少 22 个,稳定题少 10 个。

从数量看,主要胜利来自Ansible 自动化 · release 003,4/10(40.0%);从通过率看,最干净的高点是vuls 漏洞扫描器 · release 012,2/4(50.0%)。需要警惕的是vuls 漏洞扫描器 · release 011,0/10(0.0%),所以这里不是泛泛地说强弱项,而是自动化和配置管理类改动在这次运行中比边界相对清楚的 Go 漏洞扫描器改动更能闭环。因为 agent shell 是 OpenCode,这个结果更直接暴露底层模型的规划习惯,而不是强工作流包装后的表现。

Selected high and low suites, grouped by pass-at-least-once rate.选取高分和低分 suite,按三次尝试至少解出一次的比例展示。
vuls · release 012vuls 漏洞扫描器 · release 012 2/4 · 50.0%

Best visible cluster for this row: 2/4 tasks reached.这一行最明显的强项簇:4 题中解出 2 题。

Ansible · release 003Ansible 自动化 · release 003 4/10 · 40.0%
Ansible · release 004Ansible 自动化 · release 004 1/3 · 33.3%
Flipt · release 007Flipt feature flag 服务 · release 007 3/10 · 30.0%
vuls · release 010vuls 漏洞扫描器 · release 010 3/10 · 30.0%
Ansible · release 001Ansible 自动化 · release 001 2/10 · 20.0%
vuls · release 011vuls 漏洞扫描器 · release 011 0/10 · 0.0%

Weak cluster: localized Go security-scanner changes resisted this model-agent pairing.弱项簇:边界相对清楚的 Go 漏洞扫描器改动对这个模型-agent 组合不友好。

Open Library · release 014Open Library · release 014 0/10 · 0.0%

Weak cluster: large Python/Django application repairs resisted this model-agent pairing.弱项簇:大型 Python/Django 应用修复对这个模型-agent 组合不友好。

Open Library · release 015Open Library · release 015 0/10 · 0.0%

Weak cluster: large Python/Django application repairs resisted this model-agent pairing.弱项簇:大型 Python/Django 应用修复对这个模型-agent 组合不友好。

qutebrowser · release 018qutebrowser 浏览器 · release 018 0/9 · 0.0%

Weak cluster: browser/runtime integration around QtWebEngine behavior resisted this model-agent pairing.弱项簇:围绕 QtWebEngine 行为的浏览器/runtime 集成对这个模型-agent 组合不友好。

在榜单后段,低柱子往往比胜利更有信息量。它们说明模型-agent 循环在哪些任务家族上最先失去可靠规划。

在这个排名段,Redis 缓存后端缺少额外配置选项,无法连接启用 TLS 的 Redis 服务器 和成功案例一样重要。它说明模型-agent 循环在哪种任务形态上还没形成有效 verifier-backed patch。

复核从 Step 3.5 flash 中剔除了 3 次成功,但仍保留 93% 的成功集合,因此即便个别胜利有争议,suite 形状仍然有参考价值。

Original harness result vs verifier-backed audit sample原始 harness 结果 vs verifier-backed 复核样本
39 of 42 headline successes survived strict re-verification. 42 次初始成功里,39 次通过了更严格的复核。

The available audit keeps 39 of 42 initial solved attempts. Read this as a robustness check, especially when the audit sample is smaller than 453 attempts.当前可用复核保留了 42 次初始成功中的 39 次。这更适合作为稳健性检查,特别是在复核样本小于 453 次尝试时。

39 verifier-backed复核通过 3 strict rejected严格拒绝
17.04 17.04 +0.00 points+0.00 分

这一行更适合作为失败地图,而不是默认选择。先看vuls 漏洞扫描器 · release 011,0/10(0.0%):它展示了模型-agent 循环最容易失去抓手的任务形态。在 453 次尝试中只成功 42 次时,这页最有价值的是看 agent loop 在哪里先断掉。

支撑这个判断的 suite 表
Suite Repo 解出 Pass^3 通过率
release-zh-012-future-architect-vuls future-architect/vuls 2/4 2 50.0%
release-zh-003-ansible-ansible ansible/ansible 4/10 1 40.0%
release-zh-004-ansible-ansible ansible/ansible 1/3 0 33.3%
release-zh-007-flipt-io-flipt flipt-io/flipt 3/10 1 30.0%
release-zh-010-future-architect-vuls future-architect/vuls 3/10 1 30.0%
release-zh-001-ansible-ansible ansible/ansible 2/10 1 20.0%
release-zh-011-future-architect-vuls future-architect/vuls 0/10 0 0.0%
release-zh-014-internetarchive-openlibrary internetarchive/openlibrary 0/10 0 0.0%
release-zh-015-internetarchive-openlibrary internetarchive/openlibrary 0/10 0 0.0%
release-zh-018-qutebrowser-qutebrowser qutebrowser/qutebrowser 0/9 0 0.0%