Start / Large Language Model (LLM) Talk / Grpo group relative policy optimization

GRPO (Group Relative Policy Optimization)

13 min • 5 februari 2025

Group Relative Policy Optimization (GRPO) is a reinforcement learning algorithm that enhances mathematical reasoning in large language models (LLMs). It is like training students in a study group, where they learn by comparing answers without a tutor. GRPO eliminates the need for a critic model, unlike Proximal Policy Optimization (PPO), making it more resource efficient. It calculates advantages based on relative rewards within the group and directly adds KL divergence to the loss function. GRPO uses both outcome and process supervision, and can be applied iteratively, further enhancing performance. This approach is effective at improving LLMs' math skills with reduced training resources.

Kategorier

Poddar Teknologi

Förekommer på

Teknik

00:00 -00:00