Discussion about this post

User's avatar
Brian Carter's avatar

Good breakdown of KDA and LatentMoE. Clearest explanation I've read of what's actually new in K3.

The "cheap" framing is where I get stuck. Your own numbers put it at 83 turns, 120K output tokens, 56 minutes, and $10.57 per task on AA-Briefcase. That isn't cheap from where I sit as a user. Efficient per active parameter, sure. But those are two different claims wearing the same word, and I'm not sure which one the reader is meant to walk away with. Worth flagging too that Opus 5 has since passed Fable 5 on that benchmark while cutting cost per task around 20%, so the thing K3 is being measured against already moved.

Related question on the 104B active. Doesn't the full 2.8T still have to sit in GPU memory across the cluster regardless of how little fires per token? If so, sparsity lowers compute cost per token but not the cost of hosting the thing at all. Community estimates I've seen put a real self-hosted setup north of $1M. Feels like it deserves a line given how much of the framing rests on "open weights."

The part I keep circling is GLM-5.3. Same 60 on the Intelligence Index, reached through post-training RL on an older architecture. Does that complicate the case that KDA and LatentMoE are what's driving the gains? Or is the payoff somewhere the top-line number doesn't capture?

One more and I'll get out of your comments. You mention in passing that the execution-heavy RL makes K3 assumption-prone and brittle when history isn't preserved. That reads like a deployment risk, not a footnote. Have you hit it in practice?

If you want it shorter still, the GLM-5.3 question is the one most likely to get a real answer out of him. The other three he can deflect.

Think AI's avatar

Open weight models like Kimi K3 show that frontier level AI is becoming less about access and more about cost control and execution.

6 more comments...

No posts

Ready for more?