> For the complete documentation index, see [llms.txt](https://drdh.gitbook.io/rl/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://drdh.gitbook.io/rl/deep-rl-course/value-function-methods/value-iteration.md).

# Value iteration

## Tabular value iteration

Considering that $$\arg \max\_{a\_t}A^\pi(s\_t,a\_t)=\arg \max\_{a\_t}V^\pi(s\_t,a\_t)$$, we can just use $$Q^\pi$$ instead of $$A^\pi$$.

$$
Q^\pi(s,a)=r(s,a)+\gamma\mathbb{E}\_{s'\sim p(s'|s,a)}\[V^\pi(s')]
$$

Apart from that, the policy can be calculated by $$\arg\max\_{a}Q(s,a)$$. So skip the policy and compute values directly:

![Value Iteration](https://4133958719-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-LigLKy0c06y4iTEtrkI%2F-Lo5YRZanJ0iJTZaeTl8%2F-Lo5YSjwRnGBbiZCOlRS%2F1567772160319.png?generation=1567773025095450\&alt=media)

> value iteration algorithm:
>
> repeat until converge:
>
> \====1: set $$Q(s,a)\leftarrow r(s,a)+\gamma \mathbb{E}\_{s'\sim p(s'|s,a)}\[V(s')]$$
>
> \====2: set $$V(s)\leftarrow \max\_a Q(s,a)$$

## Fitted value iteration

The question is how do we represent $$V(s)$$. For small cases, we can use a big table, one entry for each discrete $$s$$, but it is not appropriate for real world, especially for image inputs, due to the curse of dimensionality. In this case, neural net function can be used: $$V: \mathcal{S}\to \mathbb{R}$$

$$
\mathcal{L}(\phi)=\frac{1}{2}\left|V\_\phi(s)-\max\_{a}Q^\pi(s,a) \right|^2
$$

> Fitted value iteration
>
> repeat until converge
>
> \====1: set $$y\_i\leftarrow \max\_{a\_i}(r(s\_i,a\_i)+\gamma \mathbb{E}\[V\_\phi(s'\_i)])$$
>
> \====2: set $$\phi\leftarrow \arg\min\_\phi\frac{1}{2}\sum\_i|V\_\phi(s\_i)-y\_i|^2$$
