> For the complete documentation index, see [llms.txt](https://drdh.gitbook.io/rl/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://drdh.gitbook.io/rl/deep-rl-course/intro-to-rl/types-of-rl-algorithms.md).

# Types of RL algorithms

We knew that the final objective is

$$
\theta^\star=\arg\max\_{\theta} E\_{\tau\sim p\_\theta(\tau)}\left\[\sum\_t r(s\_t,a\_t)\right]
$$

There are various methods to accomplish this objective.

## Model-free RL

No need to estimate the transition model.

### Policy gradient

Directly differentiate the above objective. If the problem has huge possibles, this solution will be intractable, but we can use approximation.

![Policy gradient](https://4133958719-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-LigLKy0c06y4iTEtrkI%2F-LjUunVO2ZsQQIPz9tvA%2F-LjUuykcKnBP2HI1rHew%2F1562829673011.png?generation=1562829910081093\&alt=media)

### Value-based

Estimate value function or Q-function of the optimal policy with **arg max trick**, but no explicit policy.

![Value-based](https://4133958719-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-LigLKy0c06y4iTEtrkI%2F-LjUunVO2ZsQQIPz9tvA%2F-LjUuykeKi5bd70xCP6P%2F1562829549875.png?generation=1562829910120610\&alt=media)

Use samples to fit Q-function -- supervised learning, then policy set to the best.

### Actor-critic

Estimate value function or Q-function of the current policy, the use it to improve policy. The combination of *policy gradient* and *Value-based*.

![Actor-critic](https://4133958719-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-LigLKy0c06y4iTEtrkI%2F-LjUunVO2ZsQQIPz9tvA%2F-LjUuykgKJ9r_Qo9ptP9%2F1562829738655.png?generation=1562829910069651\&alt=media)

## Model-based RL

Estimate the transition model and then...

* Use it for planning (no explicit policy)
  * Trajectory optimization/Optimal control (primarily in continuous spaces)

    Essentially backpropagation to optimize over actions
  * Discrete planning in discrete action spaces, e.g. Monte Carlo tree search
* Use it to improve policy
  * Backpropagate gradients into the policy

    requires some tricks to make it work
* Use the model to learn a value function
  * Dynamic programming
  * Generate simulated experience for model-free learner (Dyna)

![Model-based RL](https://4133958719-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-LigLKy0c06y4iTEtrkI%2F-LjUunVO2ZsQQIPz9tvA%2F-LjUuykibqt3ORfMLlBj%2F1562829437133.png?generation=1562829910138328\&alt=media)

## Examples of specific algorithms

* Policy gradient methods
  * REINFORCE
  * Natural policy gradient
  * Trust region policy optimization
* Value function fitting methods
  * Q-learning, DQN
  * Temporal difference learning
  * Fitted value iteration
* Actor-critic algorithms
  * Asynchronous advantage actor-critic (A3C)
  * Soft actor-critic (SAC)
* Model-based RL
  * Dyna
  * Guided policy search
