3 Matching Annotations
  1. Last 7 days
    1. Gemini 3.8 Live has shown a high preference among users, securing a second place in the Speech Agent Arena.

      【非共识】文章声称Gemini 3.8 Live在用户偏好方面表现出色,但未提供用户测试的具体方法、样本大小或参与用户的人口统计信息。这种'用户偏好'的主观性很强,可能受到测试设计、用户群体特征等多种因素影响,这一说法需要更多客观证据支持。

  2. Sep 2026
    1. Astra is our most aligned model, with substantial improvements in understanding user intent and model behavior—you can delegate tasks with greater confidence in Astra's judgment.

      【非共识】这一关于模型对齐的声明是主观的,没有提供客观衡量标准。'对齐'是一个复杂且多方面的概念,仅凭OpenAI的自我评估不足以证明这一点。需要第三方验证和更具体的评估方法来支持这一说法。

  3. Sep 2020