with-RL
강화 학습2024년 1월 18일

바닥부터 배우는 강화 학습 | 09. 정책 기반 에이전트

'바닥부터 배우는 강화 학습' 9장에는 정책 기반 에이전트를 학습하는 방법에 대해서 설명하고 있습니다. 아래 내용은 공부하면서 핵심 내용을 정리한 것입니다.

참고자료

9.1 Policy Gradient

◈ 목적 함수 정하기

◈ 1-Step MDP

$$\begin{equation}
\begin{split}
J(\theta) &= \sum_{s \in S} d(s) * v_{\pi_{\theta}}(s) \\
&= \sum_{s \in S} d(s) \sum_{a \in A} \pi_{\theta}(s, a) * R_{s, a} \\
\nabla_{\theta}J(\theta) &= \nabla_{\theta} \sum_{s \in S} d(s) \sum_{a \in A} \pi_{\theta}(s, a) * R_{s, a}
\end{split}
\end{equation}$$

$$\begin{equation}
\begin{split}
\nabla_{\theta}J(\theta) &= \nabla_{\theta} \sum_{s \in S} d(s) \sum_{a \in A} \pi_{\theta}(s, a) * R_{s, a} \\
&= \sum_{s \in S} d(s) \sum_{a \in A} \nabla_{\theta} \pi_{\theta}(s, a) * R_{s, a} \\
&= \sum_{s \in S} d(s) \sum_{a \in A} {\pi_{\theta}(s, a) \over \pi_{\theta}(s, a)} \nabla_{\theta} \pi_{\theta}(s, a) * R_{s, a} \\
&= \sum_{s \in S} d(s) \sum_{a \in A} \pi_{\theta}(s, a) {\nabla_{\theta} \pi_{\theta}(s, a) \over \pi_{\theta}(s, a)} * R_{s, a} \\
&= \sum_{s \in S} d(s) \sum_{a \in A} \pi_{\theta}(s, a) \nabla_{\theta} \log \pi_{\theta}(s, a) * R_{s, a} \\
\end{split}
\end{equation}$$

$$\begin{equation}
\begin{split}
\nabla_{\theta}J(\theta) &= \sum_{s \in S} d(s) \sum_{a \in A} \pi_{\theta}(s, a) \nabla_{\theta} \log \pi_{\theta}(s, a) * R_{s, a} \\
&= \mathbb{E}_{\pi_{\theta}} \left [ \nabla_{\theta} \log \pi_{\theta}(s, a) * R_{s, a} \right ]
\end{split}
\end{equation}$$

◈ 일반적 MDP에서의 Policy Gradient

9.2 REINFORCE 알고리즘

◈ 이론적 배경

REINFORCE pseudo code

◈ REINFORCE 구현

9.3 액터-크리틱

◈ Q 액터-크리틱

Q Actor-Critic pseudo code

◈ 어드밴티지 액터-크리틱

$$\nabla_{\theta}J(\theta) = \mathbb{E}_{\pi_{\theta}} \left [ \nabla_{\theta} \log \pi_{\theta}(s, a) * \{Q_{\pi_{\theta}} (s, a) -V_{\pi_{\theta}}(s)\} \right ]$$

$$\begin{equation}
\begin{split}
\mathbb{E}_{\pi_{\theta}} \left [ \nabla_{\theta} \log \pi_{\theta}(s, a) * Q_{\pi_{\theta}} (s, a) \right ] &= \mathbb{E}_{\pi_{\theta}} \left [ \nabla_{\theta} \log \pi_{\theta}(s, a) * \{Q_{\pi_{\theta}} (s, a) -V_{\pi_{\theta}}(s)\} \right ] \\
&= \mathbb{E}_{\pi_{\theta}} \left [ \nabla_{\theta} \log \pi_{\theta}(s, a) * Q_{\pi_{\theta}} (s, a) \right ] - \mathbb{E}_{\pi_{\theta}} \left [ \nabla_{\theta} \log \pi_{\theta}(s, a) * V_{\pi_{\theta}} (s) \right ] \\
\text{즉},&\;\;\;\; \mathbb{E}_{\pi_{\theta}} \left [ \nabla_{\theta} \log \pi_{\theta}(s, a) * V_{\pi_{\theta}} (s) \right ] = 0 \\
\end{split}
\end{equation}$$

◈ 증명

$$\mathbb{E}_{\pi_{\theta}} \left [ \nabla_{\theta} \log \pi_{\theta}(s, a) * B(s) \right ] = \sum_{s \in S} d_{\pi_{\theta}}(s) \sum_{a \in A} \pi_{\theta}(s, a) \nabla_{\theta} \log \pi_{\theta} (s, a) * B(s)$$

$$\begin{equation}
\begin{split}
\sum_{s \in S} d_{\pi_{\theta}}(s) \sum_{a \in A} \pi_{\theta}(s, a) \nabla_{\theta} \log \pi_{\theta} (s, a) * B(s) &= \sum_{s \in S} d_{\pi_{\theta}}(s) \sum_{a \in A} \pi_{\theta}(s, a) {\nabla_{\theta}\pi_{\theta}(s, a) \over \pi_{\theta}(s, a)} * B(s) \\
&= \sum_{s \in S} d_{\pi_{\theta}}(s) \sum_{a \in A} \nabla_{\theta}\pi_{\theta}(s, a) * B(s) \\
&= \sum_{s \in S} d_{\pi_{\theta}}(s)B(s) \sum_{a \in A} \nabla_{\theta}\pi_{\theta}(s, a) \\
&= \sum_{s \in S} d_{\pi_{\theta}}(s)B(s) \nabla_{\theta} \sum_{a \in A} \pi_{\theta}(s, a) \\
&= \sum_{s \in S} d_{\pi_{\theta}}(s)B(s) \nabla_{\theta} 1 = 0 \\
\end{split}
\end{equation}$$

$$\therefore \mathbb{E}_{\pi_{\theta}} \left [ \nabla_{\theta} \log \pi_{\theta}(s, a) * B(s) \right ] = 0$$

$$\begin{equation}
\begin{split}
\nabla_{\theta}J(\theta) &= \mathbb{E}_{\pi_{\theta}} \left [ \nabla_{\theta} \log \pi_{\theta}(s, a) * A_{\pi_{\theta}}(s, a) \right ] \\
A_{\pi_{\theta}}(s, a) &= Q_{\pi_{\theta}}(s, a) - V_{\pi_{\theta}}(s)
\end{split}
\end{equation}$$

어드밴티지 Actor-Critic pseudo code

◈ TD 액터-크리틱

$$\begin{equation}
\begin{split}
\mathbb{E}_{\pi}[\delta|s, a] &= \mathbb{E}_{\pi} \left [ r + \gamma V(s') - V(s)|s, a \right ] \\
&= \mathbb{E}_{\pi} \left [ r + \gamma V(s')|s, a \right ] - V(s) \\
&= Q(s, a) - V(s) = A(s, a)\\
\end{split}
\end{equation}$$

TD Actor-Critic pseudo code

◈ TD Actor-Critic 구현