Kimi K2.6 surpassed Opus 4.5 with a score of 56.3 in 4.8 months, and GLM-5.2 cleared GPT-5.2 with a score of 72.4 in 6 months.
这两个数字是判断模型层会否商品化的关键。但分子分母都由作者自选:换一组基准、换一个时代起点模型,追赶期就会变。当成方向性信号可以,当成定量结论会踩坑。
Kimi K2.6 surpassed Opus 4.5 with a score of 56.3 in 4.8 months, and GLM-5.2 cleared GPT-5.2 with a score of 72.4 in 6 months.
这两个数字是判断模型层会否商品化的关键。但分子分母都由作者自选:换一组基准、换一个时代起点模型,追赶期就会变。当成方向性信号可以,当成定量结论会踩坑。
GLM-5.3-Flash features a mixture of experts architecture with 320 billion parameters. It activates 18 billion parameters to answer prompts.
320B 总参只激活 18B,激活率不到 6%,比上一代激进得多。这个配比说明中国厂商已把稀疏度当成主要降本旋钮,而不是缩小模型。后果是部署门槛卡在显存容量而非算力,反而更依赖大内存硬件。
Amazon kept shutting down my tablet, so I spent $266 on four AI models to own it
Problem & Context:
com.amazon.device.software.ota, etc.) that could not be disabled without root access.The Experiment & Financials:
Model Contributions & Breakthroughs:
0x5C000) compared to the reference OTA image.selinux_enforcing to permissive, obtain a root shell, and safely remove over 100 Amazon packages (pm uninstall --user 0) without bricking the device.Key Insights & Takeaways:
Autonomous Reverse Engineering:
Device Ownership & Rights:
AI Policy & Safeguard Disparity:
AI Writing Style Debates:
GLM-5.2 vs Claude Opus
Qwen3.5 397B A17B: 15.3%, DeepSeek V3.2: 14.5%, GLM-5: 14.5%, Kimi K2.5: 11.5%, MiniMax-M2.7: 10.6%
中美专业服务 Agent 的差距在这里变得具体可见:顶级美国模型 33%,中国最强开源模型(Qwen3.5、DeepSeek、GLM-5)约 14-15%,差距超过 2 倍。更值得注意的是智谱 AI 的 GLM-5 与 DeepSeek V3.2 并列,说明在专业服务 Agent 这个维度,国内头部玩家的能力相当接近。对于智谱的战略意义:这个 2 倍差距是否可以通过领域专精(比如专注于中国本土金融场景)来弥补?
【洞察】Mythos 发布的同一天(2026年4月7日),Z.ai 发布了 GLM-5.1——一个 744B 参数的 MIT 开源模型,在 SWE-bench Pro 上甚至以 58.4% 超越了 Opus 4.6 的 57.3%。这个时间巧合揭示了一个无法回避的张力:Anthropic 试图通过限制访问来防止 AI 网络武器扩散,但开源生态系统正在以同样的速度追赶闭源前沿——Glasswing 的「防御窗口」可能比预期短得多。