
Research3w ago
AI method reports benchmark gains over GRPO with far less compute
A preprint introduces SRPO, a reflection-based training method that achieves benchmark gains across mathematics and agent tasks using significantly less compute…
#AI#Machine Learning#Research#SRPO#GRPO#AI Agents#Benchmarks