MiniMax M2.7 highspeed
A lower-table result with a few useful bright spots: 38/151 tasks solved at least once, 15/151 solved in all three attempts, with the clearest wins around large Python/Django application repairs plus automation and configuration-management work. 这是一个排名靠后但仍有局部亮点的结果:151 题中至少一次解出 38 题,三次都解出 15 题;强项主要落在大型 Python/Django 应用修复以及自动化和配置管理类改动。
How to read this result可以这样读
- MiniMax M2.7 highspeed is best read as wide but retry-sensitive: rank #23, 38 reached tasks, 15 stable solves.MiniMax M2.7 highspeed 更适合读成覆盖不窄但依赖重试:排名 #23,触达 38 题,稳定解出 15 题。
- Best suite signal: Open Library · release 013 at 6/10 (60.0%).最强 suite 信号:Open Library · release 013,6/10(60.0%)。
- Weakest visible area: Flipt · release 006 at 0/10 (0.0%).最弱可见区域:Flipt feature flag 服务 · release 006,0/10(0.0%)。
- Because the agent shell is OpenCode, the result mostly exposes the underlying model's planning habits rather than a heavily opinionated workflow.因为 agent shell 是 OpenCode,这个结果更直接暴露底层模型的规划习惯,而不是强工作流包装后的表现。
MiniMax M2.7 highspeed is broad but volatile. It can touch 38/151 tasks, yet only 15 become 3/3 solves, so much of its value comes from retrying the same benchmark surface.
The closest family reference is MiniMax M3 at rank #6. Compared with that row, this one is 8.67 points behind, with 18 fewer reached tasks and 10 fewer stable solves.
The profile has one obvious anchor: Open Library · release 013 at 6/10 (60.0%). That anchor matters because Flipt · release 006 at 0/10 (0.0%) shows the score does not generalize evenly across the benchmark. Because the agent shell is OpenCode, the result mostly exposes the underlying model’s planning habits rather than a heavily opinionated workflow.
Read the bars as a volatility chart: Open Library · release 013 at 6/10 (60.0%) shows the upside, while the Pass^3 gap explains why the same row can feel much weaker on a single rerun.
The useful contrast is between Inconsistent return type of update_key in Solr updaters (internetarchive/openlibrary · solved 3/3) and Identify CentOS Stream from CentOS to prevent incorrect EOL status and inaccurate vulnerability lookups (future-architect/vuls · solved 2/3). The model reaches both kinds of problems, but only one becomes dependable.
The audit trims 2 solved attempts from MiniMax M2.7 highspeed but still keeps 97% of the solved set, so the suite shape remains useful even where individual wins are debatable.
The available audit keeps 74 of 76 initial solved attempts. Read this as a robustness check, especially when the audit sample is smaller than 453 attempts.当前可用复核保留了 76 次初始成功中的 74 次。这更适合作为稳健性检查,特别是在复核样本小于 453 次尝试时。
Use it when breadth matters more than deterministic replay. It can find openings around Open Library · release 013 at 6/10 (60.0%), but the 23-task reach gap says a second or third run may tell a different story. The 76/453 attempt score is best read as exploration bandwidth: 38 tasks are reachable, but many need retry luck.
Supporting suite table
| Suite | Repo | Solved | Pass^3 | Rate |
|---|---|---|---|---|
release-zh-013-internetarchive-openlibrary |
internetarchive/openlibrary | 6/10 | 5 | 60.0% |
release-zh-010-future-architect-vuls |
future-architect/vuls | 5/10 | 2 | 50.0% |
release-zh-015-internetarchive-openlibrary |
internetarchive/openlibrary | 5/10 | 1 | 50.0% |
release-zh-014-internetarchive-openlibrary |
internetarchive/openlibrary | 4/10 | 2 | 40.0% |
release-zh-016-internetarchive-openlibrary |
internetarchive/openlibrary | 2/5 | 1 | 40.0% |
release-zh-004-ansible-ansible |
ansible/ansible | 1/3 | 1 | 33.3% |
release-zh-006-flipt-io-flipt |
flipt-io/flipt | 0/10 | 0 | 0.0% |
release-zh-008-flipt-io-flipt |
flipt-io/flipt | 0/10 | 0 | 0.0% |
release-zh-018-qutebrowser-qutebrowser |
qutebrowser/qutebrowser | 0/9 | 0 | 0.0% |
release-zh-017-navidrome-navidrome |
navidrome/navidrome | 0/5 | 0 | 0.0% |
MiniMax M2.7 highspeed 的特点是覆盖不窄但波动较大。它能至少一次摸到 38/151 题,但只有 15 题能做到 3/3,因此很大一部分价值来自重试。
最接近的同系参照是排名 #6 的 MiniMax M3。和它相比,这一行最终分低 8.67 分,触达题少 18 个,稳定题少 10 个。
这组画像有一个明显锚点:Open Library · release 013,6/10(60.0%)。这个锚点重要,是因为Flipt feature flag 服务 · release 006,0/10(0.0%)说明分数没有均匀迁移到整套 benchmark。因为 agent shell 是 OpenCode,这个结果更直接暴露底层模型的规划习惯,而不是强工作流包装后的表现。
这张图更像波动率图:Open Library · release 013,6/10(60.0%)展示上限,而 Pass^3 落差解释了为什么单次重跑会显得弱很多。
最有用的对比是 Solr updaters 中 update_key 返回类型不一致(internetarchive/openlibrary · 3 次中成功 3 次)和 从 CentOS 中识别 CentOS Stream,以防止 EOL 状态错误和漏洞查询不准确(future-architect/vuls · 3 次中成功 2 次):模型都能触达,但只有前者变成可靠结果。
复核从 MiniMax M2.7 highspeed 中剔除了 2 次成功,但仍保留 97% 的成功集合,因此即便个别胜利有争议,suite 形状仍然有参考价值。
The available audit keeps 74 of 76 initial solved attempts. Read this as a robustness check, especially when the audit sample is smaller than 453 attempts.当前可用复核保留了 76 次初始成功中的 74 次。这更适合作为稳健性检查,特别是在复核样本小于 453 次尝试时。
当你更看重覆盖面而不是确定复现时,它更合适。它能在Open Library · release 013,6/10(60.0%)附近找到入口,但 23 题的覆盖-稳定差说明第二、第三次运行可能给出不同结果。76/453 的单次尝试成功数更像探索带宽:38 道题能触达,但很多仍需要重试运气。
支撑这个判断的 suite 表
| Suite | Repo | 解出 | Pass^3 | 通过率 |
|---|---|---|---|---|
release-zh-013-internetarchive-openlibrary |
internetarchive/openlibrary | 6/10 | 5 | 60.0% |
release-zh-010-future-architect-vuls |
future-architect/vuls | 5/10 | 2 | 50.0% |
release-zh-015-internetarchive-openlibrary |
internetarchive/openlibrary | 5/10 | 1 | 50.0% |
release-zh-014-internetarchive-openlibrary |
internetarchive/openlibrary | 4/10 | 2 | 40.0% |
release-zh-016-internetarchive-openlibrary |
internetarchive/openlibrary | 2/5 | 1 | 40.0% |
release-zh-004-ansible-ansible |
ansible/ansible | 1/3 | 1 | 33.3% |
release-zh-006-flipt-io-flipt |
flipt-io/flipt | 0/10 | 0 | 0.0% |
release-zh-008-flipt-io-flipt |
flipt-io/flipt | 0/10 | 0 | 0.0% |
release-zh-018-qutebrowser-qutebrowser |
qutebrowser/qutebrowser | 0/9 | 0 | 0.0% |
release-zh-017-navidrome-navidrome |
navidrome/navidrome | 0/5 | 0 | 0.0% |