getting an LLM to give you a good benchmark score is fairly easy
EP.99 故事线C: 好的 benchmark 分数很容易得到,好的真实性能很难得到——这个分离正在系统性地污染 AI 能力的公开叙事。Kimi K3 在 benchmark 上接近 GPT-5.6,但实际使用中差距显著,正是这种现象的体现。
getting an LLM to give you a good benchmark score is fairly easy
EP.99 故事线C: 好的 benchmark 分数很容易得到,好的真实性能很难得到——这个分离正在系统性地污染 AI 能力的公开叙事。Kimi K3 在 benchmark 上接近 GPT-5.6,但实际使用中差距显著,正是这种现象的体现。
OpenClaw, like many other open-source tools, allows users to connect to different AI models via an application programming interface, or API. Within days of OpenClaw’s release, the team revealed that Kimi’s K2.5 had surpassed Claude Opus and became the most used AI model—by token count, meaning it was handling more total text processed across user prompts and model responses.
Wow, I had no idea that Kimi 2.5 had subbed in for Claude Opus so quickly.