1 Matching Annotations
  1. Last 7 days
    1. GLM-5.3-Flash features a mixture of experts architecture with 320 billion parameters. It activates 18 billion parameters to answer prompts.

      320B 总参只激活 18B,激活率不到 6%,比上一代激进得多。这个配比说明中国厂商已把稀疏度当成主要降本旋钮,而不是缩小模型。后果是部署门槛卡在显存容量而非算力,反而更依赖大内存硬件。