Z.ai compared GLM-5.3-Flash against Claude Opus 4.8, GPT-5.6 Terra and Gemini 3.7 Flash. The former model achieved the highest score on GDPval-AA v2, an evaluation that measures LLMs’ ability to perform knowledge work.
对照组混搭了各家不同档位(Flash 级与 Terra 级),是同价位比而非同能力比,榜首要打折看。真正可信的信号是十倍成本下降这个自比数据,而不是跨厂商名次。这与本期 Jalapeño 挑对照组的毛病同源。