The Opus 5 release was a perfect example of how useless these benchmarks are for a head to head model comparison. Anthropic published a post showing Opus 5 beating Fable in almost every eval but then added a disclaimer that it was still a tier below Fable in intelligence (and thus pricing). So then what did all the numbers represent exactly?
My experience has been a difference between "applied intelligence" and "breadth of intelligence".
Fable is the theoretical computer scientist while Opus is the Staff engineer who will implement it.
I find that Opus has continually done better on tasks mechanically but if it misunderstands even one thing -- it might waste your time doing the wrong task well.
I've found Fable to be the better thinker, filling it the gaps in your spec, and having a common sense understanding of what you likely meant.