GPT-5 (high) from 0.0% to 34.4%.
Ok so clearly some tasks need models to act fast. However, im curious why we dont just have a super cheap subagent or some filtering method figure out if this is a time sensitive task or a reasoning task and delegate. For example a mother agent should be able to identify a time strict task and open a fast subagent to wait / execute it. I suppose these arent eval based considerations but certainly an implication to keep in mind for realworld harness deisng and IMO should be more clearly stated