Ha! So we were not looking at the same chart. That makes more sense.
> Anthropic did post an official explanation, stating the original chart used a "simpler methodology" that "underestimated Sonnet 5's performance." The new chart supposedly uses their "standard methodology."
The explanation Anthropic gave for the update doesn't address how the x-axis needed to range up to $50 previously and only $10 now. In any case the pass rates are also lower.
Probably the difference between whatever it is people notice when they say models become "nerfed".
You have to test each task obviously but it is not a bad model on its face.