> In our evaluations, Kimi K3 delivers frontier-level performance. Among the models tested, its overall intelligence ranks second only to Claude Fable 5 and GPT-5.6 Sol. For the complete benchmark results, see our tech blog. The full model weights of Kimi K3 will be released in the coming days. More details on the architecture, training, and evaluation will be published together with the Kimi K3 technical report.
> K3 pushes the boundary of end-to-end knowledge work. On the GDPval-AA v2 leaderboard, Kimi K3 scores 1687. The benchmark evaluates AI models on real-world tasks across 44 occupations and 9 major industries; Kimi K3 ranks behind only Claude Fable 5 Max and GPT-5.6 Sol Max, and ahead of Claude Opus 4.8 Max at 1600.
> On AA-Briefcase, Kimi K3 scores 1527, ranking second among all models — behind only Claude Fable 5 Max and ahead of GPT-5.6 Sol Max (1495). AA-Briefcase is a private agentic knowledge-work benchmark developed by Artificial Analysis to evaluate frontier agentic capability in long-horizon knowledge work.
Really good benchmark score it seems. Maybe another DeepSeek moment right here.
France’s football team is second only to England’s and Argentina’s.
It’s a miracle that in language same words have different meanings depending on context. If this wouldn’t be the case we could have hardcoded NLP algorithmically without inventing these expensive LLMs!
That’s not what second means in this context in English, and it’s incorrect to use it that way. This is because for something to be second there must have been something in first and only first, and so on; in this case there was a first and a second already, and you cannot amalgamate then because they didn’t tie (and even if they did, they’d be 1 and 2). Both logically and grammatically, it’s incorrect.
Either think and write for yourself or stay silent next time. It'd be infinitely better than telling another person to use an LLM to understand something you yourself don't understand and are too lazy to try to figure out.
Hah, I had expected this knee-jerk response, but kinda hoped you'd avoid this pitfall. Alas.
See, I could tell you that in English, "second to" is a construct that usually means "next to" or "inferior to" and has nothing to do with "being in second place", and that if it did, it would make the popular construct "second only to" completely redundant. But others already did that in sibling comments before me, and you could just respond with "you're wrong" anyway, so what's the point? Pointing to an LLM is, of course, often a lazy and unhelpful cop out from the discussion, but in this particular case it's pointing you to a dataset that's explicitly about extracting meaning and finding relationships between phrases in languages - so you don't have to trust me or anyone else that this phrase is actually being used in this particular way, you can find it out yourself based on enormous training datasets illegally collected from all over the Internet.
If what you are saying were true, I could rank anything high by simply putting everything actually higher than it in rank into one named set and then turn around and say that the thing I want to highly rank comes in 2nd only to that first set of n items.
okay, so let’s get this straight: even though there are seemingly quite a few people here that clearly understand what is being said. However, the fact that YOU specifically either genuinely do not understand this / have never come across this before, or are being intentionally difficult because of some philosophical disagreement, feel that you can unilaterally assert that they’re “redefining words”?
I don’t know if this is genuinely your first day on Earth or something, but if you’re trying to parse English like a programming language then you’re not only making things hard on yourself, but also 99% of people you’ll ever speak to.
Which is still great because it means neither of the two best financed labs in the world manage to produce even two models themselves that would beat Kimi K3.
> > K3 pushes the boundary of end-to-end knowledge work. On the GDPval-AA v2 leaderboard, Kimi K3 scores 1687. The benchmark evaluates AI models on real-world tasks across 44 occupations and 9 major industries; Kimi K3 ranks behind only Claude Fable 5 Max and GPT-5.6 Sol Max, and ahead of Claude Opus 4.8 Max at 1600.
This is the same benchmark where Sonnet 5 outperforms Opus 4.8 max.
Like all model releases, the benchmarks aren't going to tell the whole story. All of the open weight models come with amazing benchmark results now. It's hard to believe anything other than that the benchmarks are leaking into (or intentionally included) into training data.
Possible, but pay-as-you-go Hy3 / DeepSeek v4 Pro / MiMo v2.5 Pro (from respective vendors) are genuinely good enough as daily drivers, given the costs (especially, low prices for input cache, which usually makes up 70%+ of total input for agentic workflows). I put in $10 in DeepSeek & Xiaomi MiMo, and I've barely used $1 each, in a week of coding work.
Coding Plans by MiniMax ($20/mo for 1.7b tokens) and Z.ai (~$30/week use for $17/mo) are also tremendous value for money.
It was also disruptive because it was open weight, meaning anyone and their dog could theoretically compete with the frontier labs for their inference revenue.
The frontier labs need to recoup a huge amount of cash to cover their model development costs, and justify their valuations. That’s plausible when they’re only ones capable of selling inference on these models, it a lot less plausible when models themselves become cheap commodities, and you’re just competing on your ability to provide compute. Anthropic and OpenAI can’t compete with people like AWS on that front.
It's different, but similar. If they release the weights, then we have a Fable / frontier model people can tinker with. Either way, it's still quite impressive and knocked a US company out of the top three (google). How long before China dominates the top-10 (if they don't already) or the #1 model?
cost has nothing to do with why deepseek was disruptive, the fact that it means there is zero moat around anthropic or openai is what's disruptive about it. it means in the mid-term LLMs will be commoditized and customers will flock to the cheapest inference wherever they can find it. there's no reason to stick to the "frontier" labs
if deepseek cost twice as much to train it would prove the same thing: the american companies have no monopoly on state of the art llms, and commoditization is happening
If AA is to be believed then per-task it is about the same cost as Sol. Agree that it's very different from DeepSeek v4 Pro, which is ~15x cheaper than K3.
DeepSeek didn’t really change any trends though, unless you count the stock market.
It was impressive work, but models were commoditizing and inference costs were dropping rapidly already. They were neither the first nor the last 10x optimization, from what I’ve seen.
If you know of any other 10x optimisations currently, please let me know! I'm in the market for a model that's a tenth the price of a frontier model at the same level of quality.
OK, let me be more precise: If you know of a frontier model that's ten times cheaper than the previous frontier model at that level of intelligence, please let me know, I'm in the market for one.
Its just a single benchmark, but Luna 5.6 xhigh scores within the margin of error the same as Opus 4.8 max on DeepSWE for 8x cheaper. Luna max is quite a bit higher than Opus and still 4x cheaper
In my experience, the Chinese models are much more benchmaxxed than their frontier lab competitors, so I'm taking these results with a fairly large helping of salt.
> K3 pushes the boundary of end-to-end knowledge work. On the GDPval-AA v2 leaderboard, Kimi K3 scores 1687. The benchmark evaluates AI models on real-world tasks across 44 occupations and 9 major industries; Kimi K3 ranks behind only Claude Fable 5 Max and GPT-5.6 Sol Max, and ahead of Claude Opus 4.8 Max at 1600.
> On AA-Briefcase, Kimi K3 scores 1527, ranking second among all models — behind only Claude Fable 5 Max and ahead of GPT-5.6 Sol Max (1495). AA-Briefcase is a private agentic knowledge-work benchmark developed by Artificial Analysis to evaluate frontier agentic capability in long-horizon knowledge work.
Really good benchmark score it seems. Maybe another DeepSeek moment right here.