doubao seed 2.0 code
A lower-table result with a few useful bright spots: 27/151 tasks solved at least once, 9/151 solved in all three attempts, with the clearest wins around large Python/Django application repairs plus automation and configuration-management work. 这是一个排名靠后但仍有局部亮点的结果:151 题中至少一次解出 27 题,三次都解出 9 题;强项主要落在大型 Python/Django 应用修复以及自动化和配置管理类改动。
How to read this result可以这样读
- doubao seed 2.0 code is best read as volatile explorer: rank #30, 27 reached tasks, 9 stable solves.doubao seed 2.0 code 更适合读成探索型但波动较大:排名 #30,触达 27 题,稳定解出 9 题。
- Best suite signal: Open Library · release 013 at 7/10 (70.0%).最强 suite 信号:Open Library · release 013,7/10(70.0%)。
- Weakest visible area: Flipt · release 005 at 0/10 (0.0%).最弱可见区域:Flipt feature flag 服务 · release 005,0/10(0.0%)。
- Because the agent shell is OpenCode, the result mostly exposes the underlying model's planning habits rather than a heavily opinionated workflow.因为 agent shell 是 OpenCode,这个结果更直接暴露底层模型的规划习惯,而不是强工作流包装后的表现。
doubao seed 2.0 code is a volatile explorer row around the #30 slot. The useful reading is not just the 19.42 score, but the split between 27 reached tasks and 9 stable solves.
The closest family reference is SenseNova 6.7 flash lite at rank #29. Compared with that row, this one is 3.72 points behind, with 6 fewer reached tasks and 4 fewer stable solves.
Most of the positive signal concentrates in Open Library · release 013 at 7/10 (70.0%). The opposing read is Flipt · release 005 at 0/10 (0.0%), which keeps the row from looking like a generalist. Because the agent shell is OpenCode, the result mostly exposes the underlying model’s planning habits rather than a heavily opinionated workflow.
Do not read the chart as a small version of the top rows. It is a map of early failure surfaces with a few recoverable pockets.
At this rank, Function read_subjects() in get_subjects.py exceeds acceptable complexity thresholds and includes unused logic matters as much as the wins. It shows the task shape where the model-agent loop fails before it can produce a meaningful verifier-backed patch.
The audit trims 1 solved attempt from doubao seed 2.0 code but still keeps 98% of the solved set, so the suite shape remains useful even where individual wins are debatable.
The available audit keeps 51 of 52 initial solved attempts. Read this as a robustness check, especially when the audit sample is smaller than 453 attempts.当前可用复核保留了 52 次初始成功中的 51 次。这更适合作为稳健性检查,特别是在复核样本小于 453 次尝试时。
This row is more useful as a failure map than as a default choice. Look at Flipt · release 005 at 0/10 (0.0%) first: it shows the task shape where the loop loses traction. With 52/453 solved attempts, the page is most useful for seeing where the agent loop breaks before it becomes a dependable option.
Supporting suite table
| Suite | Repo | Solved | Pass^3 | Rate |
|---|---|---|---|---|
release-zh-013-internetarchive-openlibrary |
internetarchive/openlibrary | 7/10 | 3 | 70.0% |
release-zh-003-ansible-ansible |
ansible/ansible | 4/10 | 2 | 40.0% |
release-zh-004-ansible-ansible |
ansible/ansible | 1/3 | 0 | 33.3% |
release-zh-001-ansible-ansible |
ansible/ansible | 3/10 | 1 | 30.0% |
release-zh-015-internetarchive-openlibrary |
internetarchive/openlibrary | 3/10 | 0 | 30.0% |
release-zh-012-future-architect-vuls |
future-architect/vuls | 1/4 | 0 | 25.0% |
release-zh-005-flipt-io-flipt |
flipt-io/flipt | 0/10 | 0 | 0.0% |
release-zh-007-flipt-io-flipt |
flipt-io/flipt | 0/10 | 0 | 0.0% |
release-zh-008-flipt-io-flipt |
flipt-io/flipt | 0/10 | 0 | 0.0% |
release-zh-018-qutebrowser-qutebrowser |
qutebrowser/qutebrowser | 0/9 | 0 | 0.0% |
doubao seed 2.0 code 是一个排名 #30 附近的探索型但波动较大结果。它的重点不只是 19.42 分,而是 27 道触达题和 9 道稳定题之间的差距。
最接近的同系参照是排名 #29 的 SenseNova 6.7 flash lite。和它相比,这一行最终分低 3.72 分,触达题少 6 个,稳定题少 4 个。
正面信号大多集中在Open Library · release 013,7/10(70.0%)。反向读法是Flipt feature flag 服务 · release 005,0/10(0.0%),它让这一行看起来不像通用型。因为 agent shell 是 OpenCode,这个结果更直接暴露底层模型的规划习惯,而不是强工作流包装后的表现。
不要把这张图读成头部模型的小号版本。它更像早期失败面的地图,中间夹着少数可恢复区域。
在这个排名段,get_subjects.py 中的函数 read_subjects() 超出可接受的复杂度阈值,并包含未使用逻辑 和成功案例一样重要。它说明模型-agent 循环在哪种任务形态上还没形成有效 verifier-backed patch。
复核从 doubao seed 2.0 code 中剔除了 1 次成功,但仍保留 98% 的成功集合,因此即便个别胜利有争议,suite 形状仍然有参考价值。
The available audit keeps 51 of 52 initial solved attempts. Read this as a robustness check, especially when the audit sample is smaller than 453 attempts.当前可用复核保留了 52 次初始成功中的 51 次。这更适合作为稳健性检查,特别是在复核样本小于 453 次尝试时。
这一行更适合作为失败地图,而不是默认选择。先看Flipt feature flag 服务 · release 005,0/10(0.0%):它展示了模型-agent 循环最容易失去抓手的任务形态。在 453 次尝试中只成功 52 次时,这页最有价值的是看 agent loop 在哪里先断掉。
支撑这个判断的 suite 表
| Suite | Repo | 解出 | Pass^3 | 通过率 |
|---|---|---|---|---|
release-zh-013-internetarchive-openlibrary |
internetarchive/openlibrary | 7/10 | 3 | 70.0% |
release-zh-003-ansible-ansible |
ansible/ansible | 4/10 | 2 | 40.0% |
release-zh-004-ansible-ansible |
ansible/ansible | 1/3 | 0 | 33.3% |
release-zh-001-ansible-ansible |
ansible/ansible | 3/10 | 1 | 30.0% |
release-zh-015-internetarchive-openlibrary |
internetarchive/openlibrary | 3/10 | 0 | 30.0% |
release-zh-012-future-architect-vuls |
future-architect/vuls | 1/4 | 0 | 25.0% |
release-zh-005-flipt-io-flipt |
flipt-io/flipt | 0/10 | 0 | 0.0% |
release-zh-007-flipt-io-flipt |
flipt-io/flipt | 0/10 | 0 | 0.0% |
release-zh-008-flipt-io-flipt |
flipt-io/flipt | 0/10 | 0 | 0.0% |
release-zh-018-qutebrowser-qutebrowser |
qutebrowser/qutebrowser | 0/9 | 0 | 0.0% |