GPT 5.5 (xhigh)
A top-tier result with 51/151 tasks solved at least once and 40/151 solved in all three attempts; strongest around large Python/Django application repairs plus localized Go security-scanner changes. 这是一个第一梯队结果:151 题中至少一次解出 51 题,三次都解出 40 题;强项主要落在大型 Python/Django 应用修复以及边界相对清楚的 Go 漏洞扫描器改动。
How to read this result可以这样读
- GPT 5.5 (xhigh) is best read as repeatability-first: rank #2, 51 reached tasks, 40 stable solves.GPT 5.5 (xhigh) 更适合读成稳定性优先:排名 #2,触达 51 题,稳定解出 40 题。
- Best suite signal: Open Library · release 013 at 9/10 (90.0%).最强 suite 信号:Open Library · release 013,9/10(90.0%)。
- Weakest visible area: Flipt · release 008 at 0/10 (0.0%).最弱可见区域:Flipt feature flag 服务 · release 008,0/10(0.0%)。
- The Codex shell is doing what it should here: fewer lucky one-offs, more repeated verifier-backed patches.Codex shell 在这里体现出的不是偶然命中,而是更多可重复的 verifier-backed patch。
GPT 5.5 (xhigh) stands out for repeatability. The reach number is 51/151, and 40 tasks survive all three attempts, giving it a 78% repeatability ratio among reached tasks.
The closest family reference is GPT 5.4 (xhigh) at rank #3. Compared with that row, this one is 0.09 points ahead, with 4 fewer reached tasks and 5 more stable solves.
Most of the positive signal concentrates in Open Library · release 013 at 9/10 (90.0%). The opposing read is Flipt · release 008 at 0/10 (0.0%), which keeps the row from looking like a generalist. The Codex shell is doing what it should here: fewer lucky one-offs, more repeated verifier-backed patches.
The chart matters here because it separates repeatable skill from accidental reach. Open Library · release 013 at 9/10 (90.0%) is not just a high bar; it is the area where this row most often turns a found fix into a repeatable one.
The examples reinforce the repeatability story: Inconsistent Edition Matching and Record Expansion is internetarchive/openlibrary · solved 3/3, while Forked output from ‘Display.display’ is unreliable and exposes shutdown deadlock risk shows the kind of task that still needs retry luck.
The verifier audit keeps 136/136 solved attempts for GPT 5.5 (xhigh), so the interesting question is not score inflation; it is where the model repeatedly finds the same kind of patch.
The available audit keeps 136 of 136 initial solved attempts. Read this as a robustness check, especially when the audit sample is smaller than 453 attempts.当前可用复核保留了 136 次初始成功中的 136 次。这更适合作为稳健性检查,特别是在复核样本小于 453 次尝试时。
If you are choosing it for production-style agent work, the argument is consistency: start with tasks that resemble Open Library · release 013 at 9/10 (90.0%) and expect fewer lucky-only wins. The caution is Flipt · release 008 at 0/10 (0.0%), where even this stable profile does not transfer cleanly. Across 453 attempts, the important number is not only 136 successes; it is that 40 tasks repeat cleanly.
Supporting suite table
| Suite | Repo | Solved | Pass^3 | Rate |
|---|---|---|---|---|
release-zh-013-internetarchive-openlibrary |
internetarchive/openlibrary | 9/10 | 8 | 90.0% |
release-zh-012-future-architect-vuls |
future-architect/vuls | 3/4 | 3 | 75.0% |
release-zh-014-internetarchive-openlibrary |
internetarchive/openlibrary | 7/10 | 4 | 70.0% |
release-zh-015-internetarchive-openlibrary |
internetarchive/openlibrary | 6/10 | 6 | 60.0% |
release-zh-001-ansible-ansible |
ansible/ansible | 5/10 | 3 | 50.0% |
release-zh-003-ansible-ansible |
ansible/ansible | 4/10 | 3 | 40.0% |
release-zh-008-flipt-io-flipt |
flipt-io/flipt | 0/10 | 0 | 0.0% |
release-zh-018-qutebrowser-qutebrowser |
qutebrowser/qutebrowser | 0/9 | 0 | 0.0% |
release-zh-009-flipt-io-flipt |
flipt-io/flipt | 0/5 | 0 | 0.0% |
release-zh-017-navidrome-navidrome |
navidrome/navidrome | 0/5 | 0 | 0.0% |
GPT 5.5 (xhigh) 最突出的地方是重复稳定性。它至少一次解出 51/151 题,其中 40 题三次都过,在已触达题目里的稳定比例约为 78%。
最接近的同系参照是排名 #3 的 GPT 5.4 (xhigh)。和它相比,这一行最终分高 0.09 分,触达题少 4 个,稳定题多 5 个。
正面信号大多集中在Open Library · release 013,9/10(90.0%)。反向读法是Flipt feature flag 服务 · release 008,0/10(0.0%),它让这一行看起来不像通用型。Codex shell 在这里体现出的不是偶然命中,而是更多可重复的 verifier-backed patch。
这里看图的重点不是谁最高,而是区分“稳定能力”和“偶然触达”。Open Library · release 013,9/10(90.0%) 不只是高柱子,它也是这一行最容易把解法变成稳定补丁的区域。
案例进一步说明了稳定性:版本匹配不一致与记录扩展 是internetarchive/openlibrary · 3 次中成功 3 次,而 "# 从 fork 进程调用 Display.display 的输出不可靠,并暴露 shutdown 死锁风险 则代表仍然需要重试运气的任务。
GPT 5.5 (xhigh) 的复核保留了 136 次成功中的 136 次,所以重点不是分数膨胀,而是模型在哪些地方能反复找到同类补丁。
The available audit keeps 136 of 136 initial solved attempts. Read this as a robustness check, especially when the audit sample is smaller than 453 attempts.当前可用复核保留了 136 次初始成功中的 136 次。这更适合作为稳健性检查,特别是在复核样本小于 453 次尝试时。
如果把它用于偏生产的 agent 工作,核心理由是稳定性:优先放在接近Open Library · release 013,9/10(90.0%)的任务上,不要只期待偶然命中。需要避开的参照是Flipt feature flag 服务 · release 008,0/10(0.0%),这里即使稳定型画像也不能顺利迁移。在 453 次尝试里,重要的不只是 136 次成功,而是有 40 道题可以稳定复现。
支撑这个判断的 suite 表
| Suite | Repo | 解出 | Pass^3 | 通过率 |
|---|---|---|---|---|
release-zh-013-internetarchive-openlibrary |
internetarchive/openlibrary | 9/10 | 8 | 90.0% |
release-zh-012-future-architect-vuls |
future-architect/vuls | 3/4 | 3 | 75.0% |
release-zh-014-internetarchive-openlibrary |
internetarchive/openlibrary | 7/10 | 4 | 70.0% |
release-zh-015-internetarchive-openlibrary |
internetarchive/openlibrary | 6/10 | 6 | 60.0% |
release-zh-001-ansible-ansible |
ansible/ansible | 5/10 | 3 | 50.0% |
release-zh-003-ansible-ansible |
ansible/ansible | 4/10 | 3 | 40.0% |
release-zh-008-flipt-io-flipt |
flipt-io/flipt | 0/10 | 0 | 0.0% |
release-zh-018-qutebrowser-qutebrowser |
qutebrowser/qutebrowser | 0/9 | 0 | 0.0% |
release-zh-009-flipt-io-flipt |
flipt-io/flipt | 0/5 | 0 | 0.0% |
release-zh-017-navidrome-navidrome |
navidrome/navidrome | 0/5 | 0 | 0.0% |