4,407 Matching Annotations
  1. Jul 2025
  2. Mar 2025
  3. Nov 2024
  4. Sep 2024
    1. consistently improves with more reinforcement learning (train-time compute) and with more time spent thinking (test-time compute)

      RL for post-train, time spent thinking for inference? How?