Mathematical reasoning is a hallmark of human intelligence and a cornerstone of complex problem-solving. In this paper, we introduce DeepSeekMath 7B, designed to push the limits of mathematical reasoning in open language models. We explore the potential of self-improving mathematical reasoning through Large Language Models (LLMs) and Group Relative Policy Optimization (GRPO), demonstrating state-of-the-art performance on open mathematical benchmarks.
Le raisonnement mathématique constitue une marque fondamentale de l’intelligence humaine et le pilier de la résolution de problèmes complexes. Dans cet article, nous présentons DeepSeekMath 7B, un modèle conçu pour repousser les limites du raisonnement mathématique dans les modèles de langage ouverts. Nous explorons le potentiel d’auto-amélioration du raisonnement mathématique grâce aux grands modèles de langage (LLM) et à l’optimisation de politique relative de groupe (GRPO).
Share with my friend :





Avis
Il n’y a pas encore d’avis.