Notes from: Understanding Deep Learning Basics RL Basics Markov Process Policy Value Function Bellman Equations