Rank排名 2 Codex OpenAI

GPT 5.5 (xhigh)

A top-tier result with 51/151 tasks solved at least once and 40/151 solved in all three attempts; strongest around large Python/Django application repairs plus localized Go security-scanner changes. 这是一个第一梯队结果:151 题中至少一次解出 51 题,三次都解出 40 题;强项主要落在大型 Python/Django 应用修复以及边界相对清楚的 Go 漏洞扫描器改动。

codex-cli 0.135.0 gpt-5.5#effort=xhigh Updated更新 2026-06-18

How to read this result可以这样读

  • GPT 5.5 (xhigh) is best read as repeatability-first: rank #2, 51 reached tasks, 40 stable solves.GPT 5.5 (xhigh) 更适合读成稳定性优先:排名 #2,触达 51 题,稳定解出 40 题。
  • Best suite signal: Open Library · release 013 at 9/10 (90.0%).最强 suite 信号:Open Library · release 013,9/10(90.0%)。
  • Weakest visible area: Flipt · release 008 at 0/10 (0.0%).最弱可见区域:Flipt feature flag 服务 · release 008,0/10(0.0%)。
  • The Codex shell is doing what it should here: fewer lucky one-offs, more repeated verifier-backed patches.Codex shell 在这里体现出的不是偶然命中,而是更多可重复的 verifier-backed patch。

GPT 5.5 (xhigh) stands out for repeatability. The reach number is 51/151, and 40 tasks survive all three attempts, giving it a 78% repeatability ratio among reached tasks.

The closest family reference is GPT 5.4 (xhigh) at rank #3. Compared with that row, this one is 0.09 points ahead, with 4 fewer reached tasks and 5 more stable solves.

Most of the positive signal concentrates in Open Library · release 013 at 9/10 (90.0%). The opposing read is Flipt · release 008 at 0/10 (0.0%), which keeps the row from looking like a generalist. The Codex shell is doing what it should here: fewer lucky one-offs, more repeated verifier-backed patches.

Selected high and low suites, grouped by pass-at-least-once rate.选取高分和低分 suite,按三次尝试至少解出一次的比例展示。
Open Library · release 013Open Library · release 013 9/10 · 90.0%

Best visible cluster for this row: 9/10 tasks reached.这一行最明显的强项簇:10 题中解出 9 题。

vuls · release 012vuls 漏洞扫描器 · release 012 3/4 · 75.0%
Open Library · release 014Open Library · release 014 7/10 · 70.0%
Open Library · release 015Open Library · release 015 6/10 · 60.0%
Ansible · release 001Ansible 自动化 · release 001 5/10 · 50.0%
Ansible · release 003Ansible 自动化 · release 003 4/10 · 40.0%
Flipt · release 008Flipt feature flag 服务 · release 008 0/10 · 0.0%

Weak cluster: Go product plumbing across configuration, storage, and service APIs resisted this model-agent pairing.弱项簇:横跨配置、存储和服务 API 的 Go 产品工程对这个模型-agent 组合不友好。

qutebrowser · release 018qutebrowser 浏览器 · release 018 0/9 · 0.0%

Weak cluster: browser/runtime integration around QtWebEngine behavior resisted this model-agent pairing.弱项簇:围绕 QtWebEngine 行为的浏览器/runtime 集成对这个模型-agent 组合不友好。

Flipt · release 009Flipt feature flag 服务 · release 009 0/5 · 0.0%

Weak cluster: Go product plumbing across configuration, storage, and service APIs resisted this model-agent pairing.弱项簇:横跨配置、存储和服务 API 的 Go 产品工程对这个模型-agent 组合不友好。

Navidrome · release 017Navidrome 音乐服务 · release 017 0/5 · 0.0%

Weak cluster: Go service work with persistence and API behavior resisted this model-agent pairing.弱项簇:涉及持久化和 API 行为的 Go 服务改动对这个模型-agent 组合不友好。

The chart matters here because it separates repeatable skill from accidental reach. Open Library · release 013 at 9/10 (90.0%) is not just a high bar; it is the area where this row most often turns a found fix into a repeatable one.

The examples reinforce the repeatability story: Inconsistent Edition Matching and Record Expansion is internetarchive/openlibrary · solved 3/3, while Forked output from ‘Display.display’ is unreliable and exposes shutdown deadlock risk shows the kind of task that still needs retry luck.

The verifier audit keeps 136/136 solved attempts for GPT 5.5 (xhigh), so the interesting question is not score inflation; it is where the model repeatedly finds the same kind of patch.

Original harness result vs verifier-backed audit sample原始 harness 结果 vs verifier-backed 复核样本
136 of 136 headline successes survived strict re-verification. 136 次初始成功里,136 次通过了更严格的复核。

The available audit keeps 136 of 136 initial solved attempts. Read this as a robustness check, especially when the audit sample is smaller than 453 attempts.当前可用复核保留了 136 次初始成功中的 136 次。这更适合作为稳健性检查,特别是在复核样本小于 453 次尝试时。

136 verifier-backed复核通过 0 strict rejected严格拒绝
36.88 36.88 +0.00 points+0.00 分

If you are choosing it for production-style agent work, the argument is consistency: start with tasks that resemble Open Library · release 013 at 9/10 (90.0%) and expect fewer lucky-only wins. The caution is Flipt · release 008 at 0/10 (0.0%), where even this stable profile does not transfer cleanly. Across 453 attempts, the important number is not only 136 successes; it is that 40 tasks repeat cleanly.

Supporting suite table
Suite Repo Solved Pass^3 Rate
release-zh-013-internetarchive-openlibrary internetarchive/openlibrary 9/10 8 90.0%
release-zh-012-future-architect-vuls future-architect/vuls 3/4 3 75.0%
release-zh-014-internetarchive-openlibrary internetarchive/openlibrary 7/10 4 70.0%
release-zh-015-internetarchive-openlibrary internetarchive/openlibrary 6/10 6 60.0%
release-zh-001-ansible-ansible ansible/ansible 5/10 3 50.0%
release-zh-003-ansible-ansible ansible/ansible 4/10 3 40.0%
release-zh-008-flipt-io-flipt flipt-io/flipt 0/10 0 0.0%
release-zh-018-qutebrowser-qutebrowser qutebrowser/qutebrowser 0/9 0 0.0%
release-zh-009-flipt-io-flipt flipt-io/flipt 0/5 0 0.0%
release-zh-017-navidrome-navidrome navidrome/navidrome 0/5 0 0.0%

GPT 5.5 (xhigh) 最突出的地方是重复稳定性。它至少一次解出 51/151 题,其中 40 题三次都过,在已触达题目里的稳定比例约为 78%。

最接近的同系参照是排名 #3 的 GPT 5.4 (xhigh)。和它相比,这一行最终分高 0.09 分,触达题少 4 个,稳定题多 5 个。

正面信号大多集中在Open Library · release 013,9/10(90.0%)。反向读法是Flipt feature flag 服务 · release 008,0/10(0.0%),它让这一行看起来不像通用型。Codex shell 在这里体现出的不是偶然命中,而是更多可重复的 verifier-backed patch。

Selected high and low suites, grouped by pass-at-least-once rate.选取高分和低分 suite,按三次尝试至少解出一次的比例展示。
Open Library · release 013Open Library · release 013 9/10 · 90.0%

Best visible cluster for this row: 9/10 tasks reached.这一行最明显的强项簇:10 题中解出 9 题。

vuls · release 012vuls 漏洞扫描器 · release 012 3/4 · 75.0%
Open Library · release 014Open Library · release 014 7/10 · 70.0%
Open Library · release 015Open Library · release 015 6/10 · 60.0%
Ansible · release 001Ansible 自动化 · release 001 5/10 · 50.0%
Ansible · release 003Ansible 自动化 · release 003 4/10 · 40.0%
Flipt · release 008Flipt feature flag 服务 · release 008 0/10 · 0.0%

Weak cluster: Go product plumbing across configuration, storage, and service APIs resisted this model-agent pairing.弱项簇:横跨配置、存储和服务 API 的 Go 产品工程对这个模型-agent 组合不友好。

qutebrowser · release 018qutebrowser 浏览器 · release 018 0/9 · 0.0%

Weak cluster: browser/runtime integration around QtWebEngine behavior resisted this model-agent pairing.弱项簇:围绕 QtWebEngine 行为的浏览器/runtime 集成对这个模型-agent 组合不友好。

Flipt · release 009Flipt feature flag 服务 · release 009 0/5 · 0.0%

Weak cluster: Go product plumbing across configuration, storage, and service APIs resisted this model-agent pairing.弱项簇:横跨配置、存储和服务 API 的 Go 产品工程对这个模型-agent 组合不友好。

Navidrome · release 017Navidrome 音乐服务 · release 017 0/5 · 0.0%

Weak cluster: Go service work with persistence and API behavior resisted this model-agent pairing.弱项簇:涉及持久化和 API 行为的 Go 服务改动对这个模型-agent 组合不友好。

这里看图的重点不是谁最高,而是区分“稳定能力”和“偶然触达”。Open Library · release 013,9/10(90.0%) 不只是高柱子,它也是这一行最容易把解法变成稳定补丁的区域。

案例进一步说明了稳定性:版本匹配不一致与记录扩展 是internetarchive/openlibrary · 3 次中成功 3 次,而 "# 从 fork 进程调用 Display.display 的输出不可靠,并暴露 shutdown 死锁风险 则代表仍然需要重试运气的任务。

GPT 5.5 (xhigh) 的复核保留了 136 次成功中的 136 次,所以重点不是分数膨胀,而是模型在哪些地方能反复找到同类补丁。

Original harness result vs verifier-backed audit sample原始 harness 结果 vs verifier-backed 复核样本
136 of 136 headline successes survived strict re-verification. 136 次初始成功里,136 次通过了更严格的复核。

The available audit keeps 136 of 136 initial solved attempts. Read this as a robustness check, especially when the audit sample is smaller than 453 attempts.当前可用复核保留了 136 次初始成功中的 136 次。这更适合作为稳健性检查,特别是在复核样本小于 453 次尝试时。

136 verifier-backed复核通过 0 strict rejected严格拒绝
36.88 36.88 +0.00 points+0.00 分

如果把它用于偏生产的 agent 工作,核心理由是稳定性:优先放在接近Open Library · release 013,9/10(90.0%)的任务上,不要只期待偶然命中。需要避开的参照是Flipt feature flag 服务 · release 008,0/10(0.0%),这里即使稳定型画像也不能顺利迁移。在 453 次尝试里,重要的不只是 136 次成功,而是有 40 道题可以稳定复现。

支撑这个判断的 suite 表
Suite Repo 解出 Pass^3 通过率
release-zh-013-internetarchive-openlibrary internetarchive/openlibrary 9/10 8 90.0%
release-zh-012-future-architect-vuls future-architect/vuls 3/4 3 75.0%
release-zh-014-internetarchive-openlibrary internetarchive/openlibrary 7/10 4 70.0%
release-zh-015-internetarchive-openlibrary internetarchive/openlibrary 6/10 6 60.0%
release-zh-001-ansible-ansible ansible/ansible 5/10 3 50.0%
release-zh-003-ansible-ansible ansible/ansible 4/10 3 40.0%
release-zh-008-flipt-io-flipt flipt-io/flipt 0/10 0 0.0%
release-zh-018-qutebrowser-qutebrowser qutebrowser/qutebrowser 0/9 0 0.0%
release-zh-009-flipt-io-flipt flipt-io/flipt 0/5 0 0.0%
release-zh-017-navidrome-navidrome navidrome/navidrome 0/5 0 0.0%