Kimi K2.6 surpassed Opus 4.5 with a score of 56.3 in 4.8 months, and GLM-5.2 cleared GPT-5.2 with a score of 72.4 in 6 months.
这两个数字是判断模型层会否商品化的关键。但分子分母都由作者自选:换一组基准、换一个时代起点模型,追赶期就会变。当成方向性信号可以,当成定量结论会踩坑。
Kimi K2.6 surpassed Opus 4.5 with a score of 56.3 in 4.8 months, and GLM-5.2 cleared GPT-5.2 with a score of 72.4 in 6 months.
这两个数字是判断模型层会否商品化的关键。但分子分母都由作者自选:换一组基准、换一个时代起点模型,追赶期就会变。当成方向性信号可以,当成定量结论会踩坑。
getting an LLM to give you a good benchmark score is fairly easy
EP.99 故事线C: 好的 benchmark 分数很容易得到,好的真实性能很难得到——这个分离正在系统性地污染 AI 能力的公开叙事。Kimi K3 在 benchmark 上接近 GPT-5.6,但实际使用中差距显著,正是这种现象的体现。
OpenClaw, like many other open-source tools, allows users to connect to different AI models via an application programming interface, or API. Within days of OpenClaw’s release, the team revealed that Kimi’s K2.5 had surpassed Claude Opus and became the most used AI model—by token count, meaning it was handling more total text processed across user prompts and model responses.
Wow, I had no idea that Kimi 2.5 had subbed in for Claude Opus so quickly.