Will GPT-5.4 outperform Claude Opus 4.6 at METR 50% time horizon?
41
Ṁ1kṀ3kApr 4
36%
chance
1H
6H
1D
1W
1M
ALL
This question is managed and resolved by Manifold.
Market context
Get
1,000 to start trading!
People are also trading
Sort by:
Do you mean the initial Claude 4.6 ~14.5h time horizon or the revised ~ 12h ?
Betting NO. Opus 4.6 scored ~14.5h on METR 50% time horizon. GPT-5.3 Codex scored ~5.8h. GPT-5.4 would need a >2.5x improvement over 5.3 to beat Opus 4.6, but GPT-5.2→5.3 showed essentially zero METR improvement despite being a different model. GPT-5.4 is a bigger capability jump (native computer use, strong agentic benchmarks), but the multi-choice METR market for GPT-5.4 puts the median expectation around 10-12h — still below 14.5h. The market here is pricing ~50% YES, while the multi-choice market implies ~35-38% for scores ≥14h. I see ~32% YES.