Dynamic equilibrium through reinforcement learning

Faustino, Paulo Fernando Pinho

Publicação

Dynamic equilibrium through reinforcement learning

2011-09Dissertação de mestrado

dc.contributor.advisor	Morgado, Luís Filipe Graça
dc.contributor.author	Faustino, Paulo Fernando Pinho
dc.date.accessioned	2012-02-24T14:27:36Z
dc.date.available	2012-02-24T14:27:36Z
dc.date.issued	2011-09
dc.description.abstract	Reinforcement Learning is an area of Machine Learning that deals with how an agent should take actions in an environment such as to maximize the notion of accumulated reward. This type of learning is inspired by the way humans learn and has led to the creation of various algorithms for reinforcement learning. These algorithms focus on the way in which an agent’s behaviour can be improved, assuming independence as to their surroundings. The current work studies the application of reinforcement learning methods to solve the inverted pendulum problem. The importance of the variability of the environment (factors that are external to the agent) on the execution of reinforcement learning agents is studied by using a model that seeks to obtain equilibrium (stability) through dynamism – a Cart-Pole system or inverted pendulum. We sought to improve the behaviour of the autonomous agents by changing the information passed to them, while maintaining the agent’s internal parameters constant (learning rate, discount factors, decay rate, etc.), instead of the classical approach of tuning the agent’s internal parameters. The influence of changes on the state set and the action set on an agent’s capability to solve the Cart-pole problem was studied. We have studied typical behaviour of reinforcement learning agents applied to the classic BOXES model and a new form of characterizing the environment was proposed using the notion of convergence towards a reference value. We demonstrate the gain in performance of this new method applied to a Q-Learning agent.	en
dc.description.abstract	A Aprendizagem por Reforço é uma área da Aprendizagem Automática que se preocupa com a forma como um agente deve tomar acções num ambiente de modo a maximizar a noção de recompensa acumulada. Esta forma de aprendizagem é inspirada na forma como os humanos aprendem e tem levado à criação de diversos algoritmos de aprendizagem por reforço. Estes algoritmos focam a forma de melhorar o comportamento do agente, assumindo uma independência em relação ao meio que os rodeia. O presente trabalho estuda a aplicação de métodos de aprendizagem por reforço na resolução do problema do pêndulo invertido. Neste contexto é estudado a importância da variabilidade do ambiente (factores externos ao agente) na execução de agentes de aprendizagem por reforço utilizando um modelo que tenta obter equilíbrio (estabilidade) através de dinamismo – o sistema Cart-Pole ou pêndulo invertido. Procurou-se melhorar o comportamento dos agentes autónomos alterando a informação passada a estes, mantendo constantes os parâmetros internos dos agentes (ritmo ou taxa de aprendizagem, factores de desconto, ritmo ou taxa de decaimento, etc.), em vez da vertente clássica de afinar os parâmetros internos dos agentes. Estudaram-se as influências nas alterações no conjunto de estados e no conjunto de acções na capacidade de um agente de resolver o problema do pêndulo invertido. Estudou-se o comportamento típico dos agentes de aprendizagem por reforço aplicado ao modelo clássico BOXES, sendo proposto uma nova forma de caracterizar o ambiente utilizando a noção de convergência para um valor de referência. Demonstrou-se o ganho em desempenho deste novo método aplicado a um agente Q-Learning.	por
dc.identifier.citation	Faustino, Paulo Fernando Pinho - Dynamic equilibrium through reinforcement learning. Lisboa: Instituto Superior de Engenharia de Lisboa, 2011. Dissertação de mestrado.
dc.identifier.uri	http://hdl.handle.net/10400.21/1144
dc.language.iso	eng	por
dc.peerreviewed	yes	por
dc.subject	Dynamic equilibrium	por
dc.subject	Equilíbrio dinâmico	por
dc.subject	Reinforcement learning	por
dc.subject	Aprendizagem por reforço	por
dc.subject	Autonomous agents	por
dc.subject	Agentes autónomos	por
dc.subject	Inverted pendulum	por
dc.subject	Pêndulo invertido	por
dc.title	Dynamic equilibrium through reinforcement learning	por
dc.type	master thesis
dspace.entity.type	Publication
rcaap.rights	openAccess	por
rcaap.type	masterThesis	por

Ficheiros

Principais

A mostrar 1 - 2 de 2

Nome:: Dissertação Inglês
Tamanho:: 2.49 MB
Formato:: Adobe Portable Document Format

Ver/Abrir

Nome:: Anexo_plano
Tamanho:: 476.77 KB
Formato:: Adobe Portable Document Format

Ver/Abrir

Licença

A mostrar 1 - 1 de 1

Nome:: license.txt
Tamanho:: 1.71 KB
Formato:: Item-specific license agreed upon to submission
Descrição:

Ver/Abrir

Coleções

ISEL - Eng. Elect. Tel. Comp. - Dissertações de Mestrado