These benchmark numbers are insane. The days when China was 6 months behind are over? How are they doing this with so much less resources than the US??? I have so much respect for the researchers there
Mythos/Fable-class models have been around for at least 4 months internally in the US, and Kimi still isn't quite there, so I'd say the 6-months is still about right.
Initial testing for Mythos was in April 2026, right? Sure, they had the model internally before that when they were working on it, but the same is true for Moonshot and K3.
This is fair, with the caveat that we don't know for how long this model has been around internally in China either. So we can only go about appearance / releases.
I'm not sure where "so much less resources" comes from. Training the best model has nothing to do with having the most NVIDIA GPUs around. If that were true then xAI would have the best model. It comes down to the quality of data, research, and financial backing.
Obviously it comes down to that, but you can't make the claim that GPUs aren't a huge part of it. Otherwise, billions wouldn't be getting invested into them in the west, no? And "financial backing" essentially boils down to the researchers, which boils down to the quality of the data and research, and to compute. I do think they have really smart researchers, obviously — otherwise this wouldn't be possible.
To summarise the full results table further down the page (which doesn't render on the page for me!):
Kimi K3 beats each model (out of 35 benchmarks, excluding missing):
vs Fable 5 : 12/35 (34%) (ties: 1)
vs GPT 5.6 Sol : 19/34 (56%) (ties: 1)
vs Opus 4.8 : 30/35 (86%)
vs GPT 5.5 : 30/34 (88%) (ties: 2)
vs GLM-5.2 : 19/19 (100%)
Beats Opus 4.8 and GPT 5.5 on all programming and agentic programming benchmarks except Toolathlon-Verified, often by a lot!
Astonishing. Considering none of the BigTech except Google (Microsoft, Apple, Meta, Amazon, Nvidia, SpaceX) have managed to challenge OpenAI & Anthropic frontier models, such achievements are scarcely believable.
Re: GLM-5.2: For a ~750b model, it holds up pretty good against models 3x its size (and ~10x the cost). Same goes for Tencent Hy3 and MiniMax M3, which almost match Opus 4.6 levels with ~295b params.
Because, realistically that's all programmers ever need, would be my guess. I do think linear algebra is an extremely interesting topic in its own right / outside, but yeah.
In the past, they just ran Deepseek OCR on your image and extracted the text, then gave it to a language only model. I believe now there is a model that actually takes images as input directly.
I don't think this is private knowledge guessing from when and how I was told, so I feel comfortable sharing it.
When I talked to some Huawei representatives, I was told DeepSeek V4 was trained entirely on Huawei chips. It's up to you whether you believe it or not, and while I see the incentives in faking these news, the blow if not true would be so massive that I don't think their representatives at large venues would be making these claims without thinking it's truly correct.
reply