Can GRPO be 10x Efficient? Kwai AI’s SRPO Suggests Yes with SRPO
The remarkable success of OpenAI’s o1 series and DeepSeek-R1 has unequivocally demonstrated the power of large-scale reinforcement learning (RL) in el
Synced
The remarkable success of OpenAI’s o1 series and DeepSeek-R1 has unequivocally demonstrated the power of large-scale reinforcement learning (RL) in el
Synced
DeepSeek AI has announced the release of DeepSeek-Prover-V2, a groundbreaking open-source large language model specifically designed for formal theore
Synced
A newly released 14-page technical paper from the team behind DeepSeek-V3, with DeepSeek CEO Wenfeng Liang as a co-author, sheds light on the “Scaling
Synced
Video world models, which predict future frames conditioned on actions, hold immense promise for artificial intelligence, enabling agents to plan and