From RLVR to RLSVR Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement paper: | HACKOBAR_