two major conditions of the validity of benchmark results
Less reletad to the paper random thought: These evaluaiton practices are great but I wonder if what are the most popular ways to evaluate generalization of AI agents. Feels like interval evals (with non-contaminated data and complex problems from diverse domains) of frontier labs is still the standard. But this may become less and less scalable as people put everything into model training data, where data contamination is kind of unavoidable, semantically speaking.