개미 집단 시스템에서 TD-오류를 이용한 강화학습 기법

Seung Hwan Lee ,Tae Sub Chung

doi:10.3745/kipstb.2004.11b.1.077

Abstract

강화학습에서 temporal-credit 할당 문제 즉, 에이전트가 현재 상태에서 어떤 행동을 선택하여 상태전이를 하였을 때 에이전트가 선택한 행동에 대해 어떻게 보상(reward)할 것인가는 강화학습에서 중요한 과제라 할 수 있다. 본 논문에서는 조합최적화(hard combinational optimization) 문제를 해결하기 위한 새로운 메타 휴리스틱(meta heuristic) 방법으로, greedy search뿐만 아니라 긍정적 반응의 탐색을 사용한 모집단에 근거한 접근법으로 Traveling Salesman Problem(TSP)를 풀기 위해 제안된 Ant Colony System(ACS) Algorithms에 Q-학습을 적용한 기존의 Ant-Q 학습방범을 살펴보고 이 학습 기법에 다양화 전략을 통한 상태전이와 TD-오류를 적용한 학습방법인 Ant-TD 강화학습 방법을 제안한다. 제안한 강화학습은 기존의 ACS, Ant-Q학습보다 최적해에 더 빠르게 수렴할 수 있음을 실험을 통해 알 수 있었다. Reinforcement learning takes reward about selecting action when agent chooses some action and did state transition in Present state. this can be the important subject in reinforcement learning as temporal-credit assignment problems. In this paper, by new meta heuristic method to solve hard combinational optimization problem, examine Ant-Q learning method that is proposed to solve Traveling Salesman Problem (TSP) to approach that is based for population that use positive feedback as well as greedy search. And, suggest Ant-TD reinforcement learning method that apply state transition through diversification strategy to this method and TD-error. We can show through experiments that the reinforcement learning method proposed in this Paper can find out an optimal solution faster than other reinforcement learning method like ACS and Ant-Q learning.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

개미 집단 시스템에서 TD-오류를 이용한 강화학습 기법

Abstract

Talk to us

Similar Papers

More From: The KIPS Transactions:PartB

Lead the way for us

Similar Papers

Multiagent Reinforcement Learning Algorithm Using Temporal Difference Error
Seunggwan Lee
-
Seunggwan LeeSeunggwan Lee
01 Jan 2004
01 Jan 2004

Ant Colony System에서 효율적 경로 탐색을 위한 지역갱신과 전역갱신에서의 추가 강화에 관한 연구
...
The KIPS Transactions:PartB | VOL. 10B
, et. al. ...
01 Jun 2003
The KIPS Transactions:PartB | VOL. 10B

Improved ant agents system by the dynamic parameter decision
Seunggwan Lee ... Taechoong Chung
-
Seunggwan Lee, et. al. Seunggwan Lee ... Taechoong Chung
02 Dec 2001
02 Dec 2001

An effective dynamic weighted rule for ant colony system optimization
Seung Gwan Lee ... Tae Ung Jung
-
Seung Gwan Lee, et. al. Seung Gwan Lee ... Tae Ung Jung
27 May 2001
27 May 2001

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

개미 집단 시스템에서 TD-오류를 이용한 강화학습 기법

Abstract

Talk to us

Similar Papers

More From: The KIPS Transactions:PartB