Rank排名 3 Codex OpenAI

GPT 5.4 (xhigh)

A top-tier result with 55/151 tasks solved at least once and 35/151 solved in all three attempts; strongest around large Python/Django application repairs plus localized Go security-scanner changes. 这是一个第一梯队结果:151 题中至少一次解出 55 题,三次都解出 35 题;强项主要落在大型 Python/Django 应用修复以及边界相对清楚的 Go 漏洞扫描器改动。

codex-cli 0.135.0 gpt-5.4#effort=xhigh Updated更新 2026-06-18

How to read this result可以这样读

  • GPT 5.4 (xhigh) is best read as front-runner with a balanced profile: rank #3, 55 reached tasks, 35 stable solves.GPT 5.4 (xhigh) 更适合读成第一梯队里的均衡型:排名 #3,触达 55 题,稳定解出 35 题。
  • Best suite signal: Open Library · release 013 at 9/10 (90.0%).最强 suite 信号:Open Library · release 013,9/10(90.0%)。
  • Weakest visible area: qutebrowser · release 018 at 0/9 (0.0%).最弱可见区域:qutebrowser 浏览器 · release 018,0/9(0.0%)。
  • This Codex row searches broadly, but the lower repeatability says several wins still depend on one successful trajectory.这一行 Codex 搜索面很宽,但较低的重复稳定性说明不少胜利仍依赖某一次成功轨迹。

GPT 5.4 (xhigh) belongs in the leading cluster because it keeps both breadth and stability in play: 55 reached tasks, 35 stable solves, and a 36.79 Final Score.

The closest family reference is GPT 5.5 (xhigh) at rank #2. Compared with that row, this one is 0.09 points behind, with 4 more reached tasks and 5 fewer stable solves.

The profile has one obvious anchor: Open Library · release 013 at 9/10 (90.0%). That anchor matters because qutebrowser · release 018 at 0/9 (0.0%) shows the score does not generalize evenly across the benchmark. This Codex row searches broadly, but the lower repeatability says several wins still depend on one successful trajectory.

Selected high and low suites, grouped by pass-at-least-once rate.选取高分和低分 suite,按三次尝试至少解出一次的比例展示。
Open Library · release 013Open Library · release 013 9/10 · 90.0%

Best visible cluster for this row: 9/10 tasks reached.这一行最明显的强项簇:10 题中解出 9 题。

vuls · release 012vuls 漏洞扫描器 · release 012 3/4 · 75.0%
vuls · release 011vuls 漏洞扫描器 · release 011 6/10 · 60.0%
Open Library · release 014Open Library · release 014 6/10 · 60.0%
Open Library · release 015Open Library · release 015 6/10 · 60.0%
Ansible · release 001Ansible 自动化 · release 001 4/10 · 40.0%
qutebrowser · release 018qutebrowser 浏览器 · release 018 0/9 · 0.0%

Weak cluster: browser/runtime integration around QtWebEngine behavior resisted this model-agent pairing.弱项簇:围绕 QtWebEngine 行为的浏览器/runtime 集成对这个模型-agent 组合不友好。

Navidrome · release 017Navidrome 音乐服务 · release 017 0/5 · 0.0%

Weak cluster: Go service work with persistence and API behavior resisted this model-agent pairing.弱项簇:涉及持久化和 API 行为的 Go 服务改动对这个模型-agent 组合不友好。

Flipt · release 005Flipt feature flag 服务 · release 005 1/10 · 10.0%
Flipt · release 008Flipt feature flag 服务 · release 008 1/10 · 10.0%

At the front of the board, the chart is a fingerprint. The score is close to peers, so the repo distribution says more than the rank delta.

The cases are useful because top rows can look similar in aggregate. Refactor build_marc() into expand_record() and relocate to catalog/utils for clarity and reuse shows the reliable core; Rollout audit logs lack necessary fields for segment information shows the remaining edge of variance.

The verifier audit keeps 136/136 solved attempts for GPT 5.4 (xhigh), so the interesting question is not score inflation; it is where the model repeatedly finds the same kind of patch.

Original harness result vs verifier-backed audit sample原始 harness 结果 vs verifier-backed 复核样本
136 of 136 headline successes survived strict re-verification. 136 次初始成功里,136 次通过了更严格的复核。

The available audit keeps 136 of 136 initial solved attempts. Read this as a robustness check, especially when the audit sample is smaller than 453 attempts.当前可用复核保留了 136 次初始成功中的 136 次。这更适合作为稳健性检查,特别是在复核样本小于 453 次尝试时。

136 verifier-backed复核通过 0 strict rejected严格拒绝
36.79 36.79 +0.00 points+0.00 分

In practice, read it through the gap between Open Library · release 013 at 9/10 (90.0%) and qutebrowser · release 018 at 0/9 (0.0%). That gap is more actionable than the rank because it says which repo shape gets coherent patches. The 136/453 attempt score is the backdrop; the article above is about which parts of that score are repeatable enough to matter.

Supporting suite table
Suite Repo Solved Pass^3 Rate
release-zh-013-internetarchive-openlibrary internetarchive/openlibrary 9/10 8 90.0%
release-zh-012-future-architect-vuls future-architect/vuls 3/4 2 75.0%
release-zh-011-future-architect-vuls future-architect/vuls 6/10 2 60.0%
release-zh-014-internetarchive-openlibrary internetarchive/openlibrary 6/10 5 60.0%
release-zh-015-internetarchive-openlibrary internetarchive/openlibrary 6/10 5 60.0%
release-zh-001-ansible-ansible ansible/ansible 4/10 1 40.0%
release-zh-018-qutebrowser-qutebrowser qutebrowser/qutebrowser 0/9 0 0.0%
release-zh-017-navidrome-navidrome navidrome/navidrome 0/5 0 0.0%
release-zh-005-flipt-io-flipt flipt-io/flipt 1/10 0 10.0%
release-zh-008-flipt-io-flipt flipt-io/flipt 1/10 0 10.0%

GPT 5.4 (xhigh) 能进入第一梯队,是因为覆盖和稳定性都没有掉队:至少一次解出 55 题,稳定解出 35 题,Final Score 36.79。

最接近的同系参照是排名 #2 的 GPT 5.5 (xhigh)。和它相比,这一行最终分低 0.09 分,触达题多 4 个,稳定题少 5 个。

这组画像有一个明显锚点:Open Library · release 013,9/10(90.0%)。这个锚点重要,是因为qutebrowser 浏览器 · release 018,0/9(0.0%)说明分数没有均匀迁移到整套 benchmark。这一行 Codex 搜索面很宽,但较低的重复稳定性说明不少胜利仍依赖某一次成功轨迹。

Selected high and low suites, grouped by pass-at-least-once rate.选取高分和低分 suite,按三次尝试至少解出一次的比例展示。
Open Library · release 013Open Library · release 013 9/10 · 90.0%

Best visible cluster for this row: 9/10 tasks reached.这一行最明显的强项簇:10 题中解出 9 题。

vuls · release 012vuls 漏洞扫描器 · release 012 3/4 · 75.0%
vuls · release 011vuls 漏洞扫描器 · release 011 6/10 · 60.0%
Open Library · release 014Open Library · release 014 6/10 · 60.0%
Open Library · release 015Open Library · release 015 6/10 · 60.0%
Ansible · release 001Ansible 自动化 · release 001 4/10 · 40.0%
qutebrowser · release 018qutebrowser 浏览器 · release 018 0/9 · 0.0%

Weak cluster: browser/runtime integration around QtWebEngine behavior resisted this model-agent pairing.弱项簇:围绕 QtWebEngine 行为的浏览器/runtime 集成对这个模型-agent 组合不友好。

Navidrome · release 017Navidrome 音乐服务 · release 017 0/5 · 0.0%

Weak cluster: Go service work with persistence and API behavior resisted this model-agent pairing.弱项簇:涉及持久化和 API 行为的 Go 服务改动对这个模型-agent 组合不友好。

Flipt · release 005Flipt feature flag 服务 · release 005 1/10 · 10.0%
Flipt · release 008Flipt feature flag 服务 · release 008 1/10 · 10.0%

在榜单前排,这张图更像指纹。分数和相邻模型很接近,因此代码库分布比分差更说明问题。

前排模型在总分上容易看起来相似,所以案例很关键:将 build_marc() 重构为 expand_record() 并迁移至 catalog/utils 以提升清晰度与复用性 展示可靠核心,Rollout 审计日志缺少表示 segment 信息所需的字段 展示剩余波动边界。

GPT 5.4 (xhigh) 的复核保留了 136 次成功中的 136 次,所以重点不是分数膨胀,而是模型在哪些地方能反复找到同类补丁。

Original harness result vs verifier-backed audit sample原始 harness 结果 vs verifier-backed 复核样本
136 of 136 headline successes survived strict re-verification. 136 次初始成功里,136 次通过了更严格的复核。

The available audit keeps 136 of 136 initial solved attempts. Read this as a robustness check, especially when the audit sample is smaller than 453 attempts.当前可用复核保留了 136 次初始成功中的 136 次。这更适合作为稳健性检查,特别是在复核样本小于 453 次尝试时。

136 verifier-backed复核通过 0 strict rejected严格拒绝
36.79 36.79 +0.00 points+0.00 分

实际选择时,更应该通过Open Library · release 013,9/10(90.0%)和qutebrowser 浏览器 · release 018,0/9(0.0%)之间的落差来读它。这个落差比分数排名更可操作,因为它说明哪类代码库更容易得到连贯补丁。136/453 的单次尝试成功数只是背景;上面的文章重点是哪些部分足够可重复、值得当成能力看。

支撑这个判断的 suite 表
Suite Repo 解出 Pass^3 通过率
release-zh-013-internetarchive-openlibrary internetarchive/openlibrary 9/10 8 90.0%
release-zh-012-future-architect-vuls future-architect/vuls 3/4 2 75.0%
release-zh-011-future-architect-vuls future-architect/vuls 6/10 2 60.0%
release-zh-014-internetarchive-openlibrary internetarchive/openlibrary 6/10 5 60.0%
release-zh-015-internetarchive-openlibrary internetarchive/openlibrary 6/10 5 60.0%
release-zh-001-ansible-ansible ansible/ansible 4/10 1 40.0%
release-zh-018-qutebrowser-qutebrowser qutebrowser/qutebrowser 0/9 0 0.0%
release-zh-017-navidrome-navidrome navidrome/navidrome 0/5 0 0.0%
release-zh-005-flipt-io-flipt flipt-io/flipt 1/10 0 10.0%
release-zh-008-flipt-io-flipt flipt-io/flipt 1/10 0 10.0%