Event Details
Bandit-Based Planning in Continuous Action Markov Decision Processes
- Event Date: October 24, 2011
- Event Start Time: 12:00 PM
- Event End Time: 7:00 PM
- Event Location: Rutgers Computer Science, Rutgers Perceptual Science
- Event Type: Human and Computer Vision Series
- Event Semester: Fall 2011
- Event Contact: Ari Weinstein
- Event Extra info: <a href="http://aresearch.wordpress.com/">Ari Weinstein</a>
In reinforcement learning, algorithms traditionally are concerned with finding a policy (an optimal mapping of all states to actions) in domains that have a finite state and action space. Extending this approach to spaces with continuous state and action spaces, however, is difficult because methods such as coarse discretization or function approximation can provably cause failure to converge to optimal values in many cases. In this talk, I will discuss a planning algorithm that functions natively in continuous action spaces and is agnostic to state during planning. As such, it does not suffer from problems which arise when trying to represent a global policy. Empirical results demonstrate that the algorithm outperforms current state of the art methods for planning in continuous state and action domains.