Human
ChatGPT 5.5 scored 87 where the next best model scored 67. Here's what that gap looks like in real work.
Using GPT-5.5 for the first time was the most blown-away I have felt about a model release in a while, and the reason is not benchmark scores. The reason is that I handed it the kind of work that breaks models, the kind with messy files and legal risk and 23 deliverables that have to open in the right format, and it came back with something close to a real executive handoff. That has not happened before.
4 months ago










