Cloud
Hi localhost community!
Cloud
Hi localhost community!
METR pays human programmers a minimum of $50 per hour, so getting a baseline for a single 160-hour task would cost at least $8,000.
一道测试题的人类基准成本高达 8000 美元——这个数字揭示了 AI 评测的一个被严重低估的物理限制:测量 AI 能力需要大量人类劳动,而随着 AI 能力向「月级任务」延伸,建立可靠基准的成本将呈超线性增长。更根本的问题是:你很难让一个有能力的程序员花数周时间做一个「测试任务」,即便报酬丰厚。人类评测员的可获得性,将成为 AI 能力评估的真正天花板。
Portable Typewriters Today - February 2015<br /> by [[Will Davis]] on 2015-02-10<br /> accessed on 2025-08-05T16:35:48
The learned Dr. Guy Patin says: “On dit que M. de Meziriac avoit corrigé dans son Amyot huit mille fautes, et qu’Amyot n’avoit pas de bons exemplaires, ou qu’il n’avoit pas bien entendu le Grec de Plutarque.”3
Translation: It is said that M. de Meziriac had corrected eight thousand mistakes in his Amyot, and that Amyot did not have good copies, or that he had not understood Plutarch's Greek well.