Self-Trained Verification for Training- and Test-Time Self-Improvement

ArXi:2605.30290v1 Announce Type: cross Self-improvement at scale has been a longstanding goal for reasoning models, and there are two natural places to do it: at test time, through verification-refinement (V-R) loops; and at