4 comments

  • senko 1 minute ago
    [delayed]
  • delichon 2 minutes ago
    If you want to believe that the success of Kimi is about distillation attacks, ignore this.
  • jasonjmcghee 1 hour ago
    (2025)

    As it's 9 months old and they just had a major model release

    • throwa356262 48 minutes ago
      For K3 read this instead: https://arxiv.org/abs/2607.24653

      The main contribution of the K3 paper is Stable LatentMoE. Like some other models it compresses data sent between layers, which puts certain requirements on the router. K3 improves performance by using a more balanced expert selection strategy.

      • mcbuilder 32 minutes ago
        Compared to the Opus 5 "model card", which read like a standard Anthropic set of alignment principles and safety concerns, this presents a plethora of useful technical details that advances the state of the art.
        • throwa356262 28 minutes ago
          Same with DeepSeek papers, they are a joy to read.
    • cptcobalt 45 minutes ago
      Rather under-discussed back then: https://news.ycombinator.com/item?id=45766937
    • GaggiX 1 hour ago
      I believe OP posted it because the new Kimi K3 has 69 KDA layers (the rest are 24 Gated MLA), I think previous large Kimi models had only MLA layers.
      • yorwba 23 minutes ago
        It's not the same KDA as used in Kimi Linear, though.
        • GaggiX 21 minutes ago
          What's the difference? They are both called Kimi Delta Attention.
  • Topology1 16 minutes ago
    Another banger from Zhang et. al