LongCat 2.0 Preview
A lower-table result with a few useful bright spots: 33/151 tasks solved at least once, 17/151 solved in all three attempts, with the clearest wins around automation and configuration-management work plus Go product plumbing across configuration, storage, and service APIs. 这是一个排名靠后但仍有局部亮点的结果:151 题中至少一次解出 33 题,三次都解出 17 题;强项主要落在自动化和配置管理类改动以及横跨配置、存储和服务 API 的 Go 产品工程。
How to read this result可以这样读
- LongCat 2.0 Preview is best read as volatile explorer: rank #25, 33 reached tasks, 17 stable solves.LongCat 2.0 Preview 更适合读成探索型但波动较大:排名 #25,触达 33 题,稳定解出 17 题。
- Best suite signal: Ansible · release 003 at 5/10 (50.0%).最强 suite 信号:Ansible 自动化 · release 003,5/10(50.0%)。
- Weakest visible area: Open Library · release 015 at 0/10 (0.0%).最弱可见区域:Open Library · release 015,0/10(0.0%)。
- Because the agent shell is OpenCode, the result mostly exposes the underlying model's planning habits rather than a heavily opinionated workflow.因为 agent shell 是 OpenCode,这个结果更直接暴露底层模型的规划习惯,而不是强工作流包装后的表现。
LongCat 2.0 Preview is a volatile explorer row around the #25 slot. The useful reading is not just the 24.78 score, but the split between 33 reached tasks and 17 stable solves.
The closest family reference is MiMo v2.5 at rank #24. Compared with that row, this one is 0.46 points behind, with the same reached tasks and 1 fewer stable solves.
The result is easiest to understand as a three-point shape: volume at Ansible · release 003 at 5/10 (50.0%), efficiency at vuls · release 012 at 3/4 (75.0%), and resistance at Open Library · release 015 at 0/10 (0.0%). Because the agent shell is OpenCode, the result mostly exposes the underlying model’s planning habits rather than a heavily opinionated workflow.
For this row, the suite bars are a contrast tool. The distance between Ansible · release 003 at 5/10 (50.0%) and Open Library · release 015 at 0/10 (0.0%) is the model’s practical boundary.
The examples keep the middle-band story honest: Refactor build_marc() into expand_record() and relocate to catalog/utils for clarity and reuse is the upside, Add Support for Galaxy Server Configuration in ansible-config Command is the failure surface, and the page should be read between those two poles.
The audit trims 6 solved attempts from LongCat 2.0 Preview but still keeps 92% of the solved set, so the suite shape remains useful even where individual wins are debatable.
The available audit keeps 69 of 75 initial solved attempts. Read this as a robustness check, especially when the audit sample is smaller than 453 attempts.当前可用复核保留了 75 次初始成功中的 69 次。这更适合作为稳健性检查,特别是在复核样本小于 453 次尝试时。
In practice, read it through the gap between Ansible · release 003 at 5/10 (50.0%) and Open Library · release 015 at 0/10 (0.0%). That gap is more actionable than the rank because it says which repo shape gets coherent patches. The 75/453 attempt score is the backdrop; the article above is about which parts of that score are repeatable enough to matter.
Supporting suite table
| Suite | Repo | Solved | Pass^3 | Rate |
|---|---|---|---|---|
release-zh-012-future-architect-vuls |
future-architect/vuls | 3/4 | 3 | 75.0% |
release-zh-003-ansible-ansible |
ansible/ansible | 5/10 | 3 | 50.0% |
release-zh-002-ansible-ansible |
ansible/ansible | 4/10 | 1 | 40.0% |
release-zh-004-ansible-ansible |
ansible/ansible | 1/3 | 0 | 33.3% |
release-zh-001-ansible-ansible |
ansible/ansible | 3/10 | 2 | 30.0% |
release-zh-007-flipt-io-flipt |
flipt-io/flipt | 3/10 | 2 | 30.0% |
release-zh-015-internetarchive-openlibrary |
internetarchive/openlibrary | 0/10 | 0 | 0.0% |
release-zh-018-qutebrowser-qutebrowser |
qutebrowser/qutebrowser | 0/9 | 0 | 0.0% |
release-zh-016-internetarchive-openlibrary |
internetarchive/openlibrary | 0/5 | 0 | 0.0% |
release-zh-017-navidrome-navidrome |
navidrome/navidrome | 0/5 | 0 | 0.0% |
LongCat 2.0 Preview 是一个排名 #25 附近的探索型但波动较大结果。它的重点不只是 24.78 分,而是 33 道触达题和 17 道稳定题之间的差距。
最接近的同系参照是排名 #24 的 MiMo v2.5。和它相比,这一行最终分低 0.46 分,触达题持平,稳定题少 1 个。
这个结果最容易读成三点形状:数量在Ansible 自动化 · release 003,5/10(50.0%),效率在vuls 漏洞扫描器 · release 012,3/4(75.0%),阻力在Open Library · release 015,0/10(0.0%)。因为 agent shell 是 OpenCode,这个结果更直接暴露底层模型的规划习惯,而不是强工作流包装后的表现。
对这一行来说,suite 柱更像对比工具。Ansible 自动化 · release 003,5/10(50.0%)和Open Library · release 015,0/10(0.0%)之间的距离,就是模型的实用边界。
这些案例让中段模型画像更具体:将 build_marc() 重构为 expand_record() 并迁移至 catalog/utils 以提升清晰度与复用性 是上限,在 ansible-config command 中支持 Galaxy Server Configuration 是失败面,这页应该在两者之间读。
复核从 LongCat 2.0 Preview 中剔除了 6 次成功,但仍保留 92% 的成功集合,因此即便个别胜利有争议,suite 形状仍然有参考价值。
The available audit keeps 69 of 75 initial solved attempts. Read this as a robustness check, especially when the audit sample is smaller than 453 attempts.当前可用复核保留了 75 次初始成功中的 69 次。这更适合作为稳健性检查,特别是在复核样本小于 453 次尝试时。
实际选择时,更应该通过Ansible 自动化 · release 003,5/10(50.0%)和Open Library · release 015,0/10(0.0%)之间的落差来读它。这个落差比分数排名更可操作,因为它说明哪类代码库更容易得到连贯补丁。75/453 的单次尝试成功数只是背景;上面的文章重点是哪些部分足够可重复、值得当成能力看。
支撑这个判断的 suite 表
| Suite | Repo | 解出 | Pass^3 | 通过率 |
|---|---|---|---|---|
release-zh-012-future-architect-vuls |
future-architect/vuls | 3/4 | 3 | 75.0% |
release-zh-003-ansible-ansible |
ansible/ansible | 5/10 | 3 | 50.0% |
release-zh-002-ansible-ansible |
ansible/ansible | 4/10 | 1 | 40.0% |
release-zh-004-ansible-ansible |
ansible/ansible | 1/3 | 0 | 33.3% |
release-zh-001-ansible-ansible |
ansible/ansible | 3/10 | 2 | 30.0% |
release-zh-007-flipt-io-flipt |
flipt-io/flipt | 3/10 | 2 | 30.0% |
release-zh-015-internetarchive-openlibrary |
internetarchive/openlibrary | 0/10 | 0 | 0.0% |
release-zh-018-qutebrowser-qutebrowser |
qutebrowser/qutebrowser | 0/9 | 0 | 0.0% |
release-zh-016-internetarchive-openlibrary |
internetarchive/openlibrary | 0/5 | 0 | 0.0% |
release-zh-017-navidrome-navidrome |
navidrome/navidrome | 0/5 | 0 | 0.0% |