Variance Penalized On-Policy and Off-Policy Actor-Critic.
Arushi JainGandharv PatilAyush JainKhimya KhetarpalDoina PrecupPublished in: AAAI (2021)
Keyphrases
- actor critic
- policy gradient
- reinforcement learning
- variance reduction
- approximate dynamic programming
- temporal difference
- neuro fuzzy
- optimal control
- gradient method
- policy gradient methods
- policy iteration
- reinforcement learning algorithms
- function approximation
- average reward
- least squares
- markov decision processes
- monte carlo
- natural actor critic
- state space
- linear program
- optimal policy
- multi agent systems
- state action
- action selection
- partially observable markov decision processes
- approximation methods
- function approximators
- machine learning
- model free
- convergence rate
- reinforcement learning problems
- model selection
- decision making