After OpenAI launched GPT-5.5 on April 23, in its announcement, the company called it the "smartest and most intuitive to use model yet."
However, going through the full benchmark tables OpenAI published might tell a different story. GPT-5.5 is behind its rival in some categories. Anthropic Claude Opus 4.7 model was released on April 16. Google Gemini 3.1 Pro launched back in February. Both models beat GPT-5.5 at specific tasks that actually matter for real work.
Matthew Berman, an AI engineer and CEO at ForwardFuture, got early access and tested GPT-5.5 for two weeks. His verdict, which he posted on X, was, "It is as good as any Opus model and oftentimes better at certain tasks…It's better than Opus at the backend, but it's still not as good at the front-end design."
Here are six (6) things GPT-5.5 still can’t do better compared to Claude Opus 4.7 and Gemini 3.1 based on publicly available benchmarks:
1. GPT-5.5 can't solve real GitHub issues better than Claude Opus 4.7
On the SWE-Bench Pro, which tests whether AI models can actually fix real GitHub issues end-to-end, Claude Opus 4.7 scored 64.3% on SWE-Bench Pro. GPT-5.5 scored 58.6%. That 5.7-point gap represents hundreds of coding tasks where Opus ships working code and GPT-5.5 won't.
Subscribe for free to continue reading this article
Subscribe SubscribeAlready Have an Account? Log In