There’s not a single thing out there that even comes close to GPT-4. Not one that I’ve used anyway. Benchmarks be damned, it’s the experience that matters and I’ve yet to have an LLM blow my mind the way GPT-4 does.
Have you used Claude? I regularly use them instead of GPT4 because of the larger context window. 4 is still useful when I want to give instructions to the system (Claude will refuse to do a lot), but generally the responses seem on par with each other.