Good breakdown of KDA and LatentMoE. Clearest explanation I've read of what's actually new in K3.
The "cheap" framing is where I get stuck. Your own numbers put it at 83 turns, 120K output tokens, 56 minutes, and $10.57 per task on AA-Briefcase. That isn't cheap from where I sit as a user. Efficient per active parameter, sure. But those are two different claims wearing the same word, and I'm not sure which one the reader is meant to walk away with. Worth flagging too that Opus 5 has since passed Fable 5 on that benchmark while cutting cost per task around 20%, so the thing K3 is being measured against already moved.
Related question on the 104B active. Doesn't the full 2.8T still have to sit in GPU memory across the cluster regardless of how little fires per token? If so, sparsity lowers compute cost per token but not the cost of hosting the thing at all. Community estimates I've seen put a real self-hosted setup north of $1M. Feels like it deserves a line given how much of the framing rests on "open weights."
The part I keep circling is GLM-5.3. Same 60 on the Intelligence Index, reached through post-training RL on an older architecture. Does that complicate the case that KDA and LatentMoE are what's driving the gains? Or is the payoff somewhere the top-line number doesn't capture?
One more and I'll get out of your comments. You mention in passing that the execution-heavy RL makes K3 assumption-prone and brittle when history isn't preserved. That reads like a deployment risk, not a footnote. Have you hit it in practice?
If you want it shorter still, the GLM-5.3 question is the one most likely to get a real answer out of him. The other three he can deflect.
I am sorry, been a follower for a while but the title is just wrong/clickbait. Lots to admire about K3 and my experience with it has been great but it is not in the same league as Fable.
I'm quoting both benchmarks and people's subjective opinions (many of whom have made the claim publicly that Kimi is close).
Because evals is still such a fuzzy field and evaluations can change a lot based on various factors (including user preferences), the standard behcmkars are the best we can base any claims on. It's not "wrong" or "clickbait" to point to industry standards in that case. You can disagree with the AA benchmark (and the others including Harvey Bench which operates in a completely different domain, where Kimi also is close to Fable), but in which case you're going to have to give a viable alternative that I can use.
Not entirety sure why you are being so defensive. To your point, I agreed that Opus 5 is terrible ( I stick to Opus 4.7). With all the fuzziness that you point to around benchmarks (that I agree with), the title of your article is very categorical which is what I took issue with.
Gave you a viable alternative in the linked article. It introduces a benchmark called TB-fn which has tasks similar to TB-2.1 but not exactly the same. Notice the difference in scores and how some models (notably Sol, Opus, Fable) are stable across both while others like Kimi K3, GLM 5.2/5.3 have noticeably reduced scores.
Most regular frontier model users would agree that K3 is not in the same class. In the same way that they agree that Opus 5 tops all benchmarks but is a horrible model to talk to/get work done.
By regular, I mean folks using hundreds of thousands of tokens per day which would likely be enterprise users with few/no token limits and all model choices available in BigTech.
You're sayijg I'm clickbaiting for referring to the industry standard benchmark and your solution is for me to refer to a new benchmark that's more niche. I could point to issues there as well (for instance opus 5 has had massive issues and many Claude code power users such as myself default to 4.6 or that their setup should also be testing cross validation, minor perfurbances and a myriad of other factors that would define intelligence or be useful for users in prod with agentic systems.).
Any benchmark/use case is an imperfect signal given how much variance there is in the space and in usage. Which is why we stack multiple use cases and user reports. That's still a very fuzzy signal but it's the best we can do while makijg more general statements.
Good breakdown of KDA and LatentMoE. Clearest explanation I've read of what's actually new in K3.
The "cheap" framing is where I get stuck. Your own numbers put it at 83 turns, 120K output tokens, 56 minutes, and $10.57 per task on AA-Briefcase. That isn't cheap from where I sit as a user. Efficient per active parameter, sure. But those are two different claims wearing the same word, and I'm not sure which one the reader is meant to walk away with. Worth flagging too that Opus 5 has since passed Fable 5 on that benchmark while cutting cost per task around 20%, so the thing K3 is being measured against already moved.
Related question on the 104B active. Doesn't the full 2.8T still have to sit in GPU memory across the cluster regardless of how little fires per token? If so, sparsity lowers compute cost per token but not the cost of hosting the thing at all. Community estimates I've seen put a real self-hosted setup north of $1M. Feels like it deserves a line given how much of the framing rests on "open weights."
The part I keep circling is GLM-5.3. Same 60 on the Intelligence Index, reached through post-training RL on an older architecture. Does that complicate the case that KDA and LatentMoE are what's driving the gains? Or is the payoff somewhere the top-line number doesn't capture?
One more and I'll get out of your comments. You mention in passing that the execution-heavy RL makes K3 assumption-prone and brittle when history isn't preserved. That reads like a deployment risk, not a footnote. Have you hit it in practice?
If you want it shorter still, the GLM-5.3 question is the one most likely to get a real answer out of him. The other three he can deflect.
I am sorry, been a follower for a while but the title is just wrong/clickbait. Lots to admire about K3 and my experience with it has been great but it is not in the same league as Fable.
It is benchmaxxed wrt benchmarks.
Interesting article: https://fidian.ai/blog/tb-fn-benchmark-results/
I'm quoting both benchmarks and people's subjective opinions (many of whom have made the claim publicly that Kimi is close).
Because evals is still such a fuzzy field and evaluations can change a lot based on various factors (including user preferences), the standard behcmkars are the best we can base any claims on. It's not "wrong" or "clickbait" to point to industry standards in that case. You can disagree with the AA benchmark (and the others including Harvey Bench which operates in a completely different domain, where Kimi also is close to Fable), but in which case you're going to have to give a viable alternative that I can use.
Not entirety sure why you are being so defensive. To your point, I agreed that Opus 5 is terrible ( I stick to Opus 4.7). With all the fuzziness that you point to around benchmarks (that I agree with), the title of your article is very categorical which is what I took issue with.
Gave you a viable alternative in the linked article. It introduces a benchmark called TB-fn which has tasks similar to TB-2.1 but not exactly the same. Notice the difference in scores and how some models (notably Sol, Opus, Fable) are stable across both while others like Kimi K3, GLM 5.2/5.3 have noticeably reduced scores.
Most regular frontier model users would agree that K3 is not in the same class. In the same way that they agree that Opus 5 tops all benchmarks but is a horrible model to talk to/get work done.
By regular, I mean folks using hundreds of thousands of tokens per day which would likely be enterprise users with few/no token limits and all model choices available in BigTech.
You're sayijg I'm clickbaiting for referring to the industry standard benchmark and your solution is for me to refer to a new benchmark that's more niche. I could point to issues there as well (for instance opus 5 has had massive issues and many Claude code power users such as myself default to 4.6 or that their setup should also be testing cross validation, minor perfurbances and a myriad of other factors that would define intelligence or be useful for users in prod with agentic systems.).
Any benchmark/use case is an imperfect signal given how much variance there is in the space and in usage. Which is why we stack multiple use cases and user reports. That's still a very fuzzy signal but it's the best we can do while makijg more general statements.