with-RL
자연어 처리 / LLM2025년 2월 9일

DeepSeek-R1 논문리뷰

느낀점

논문 주소

Abstract

1. Introduction

1.1. Contributions

1.2. Summary of Evaluation Results

2. Approach

2.1. Overview

2.2. DeepSeek-R1-Zero: Reinforcement Learning on the Base Model

2.2.1. Reinforcement Learning Algorithm

2.2.2. Reward Modeling

2.2.3. Training Template

2.2.4. Performance, Self-evolution Process and Aha Moment of DeepSeek-R1-Zero

2.3. DeepSeek-R1: Reinforcement Learning with Cold Start

2.3.1. Cold Start

2.3.2. Reasoning-oriented Reinforcement Learning

2.3.3. Rejection Sampling and Supervised Fine-Tuning

2.3.4. Reinforcement Learning for all Scenarios

2.4. Distillation: Empower Small Models with Reasoning Capability

3. Experiment

3.1. DeepSeek-R1 Evaluation

3.2. Distilled Model Evaluation

DeepSeek-R1 논문리뷰 | with-RL