Soft-SVeRL: Self-Verified Reinforcement Learning with Soft Rewards

ArXi:2605.28561v1 Announce Type: cross Reinforcement Learning from Verifiable Rewards (RLVR) has improved language models in domains such as mathematics and code, where correctness can be checked automatically. However, many important tasks are only partially verifiable: prompts contain multiple requirements, responses may satisfy some but not all of them, or no single reference answer might exist. We