Preference Learning for Move Prediction and Evaluation Function Approximation in Othello

Thomas Philip Runarsson,Simon M Lucas

doi:10.1109/tciaig.2014.2307272

Thomas Philip Runarsson, Simon M Lucas

Open Access

https://doi.org/10.1109/tciaig.2014.2307272

Copy DOI

Abstract

This paper investigates the use of preference learning as an approach to move prediction and evaluation function ap- proximation, using the game of Othello as a test domain. Using the same sets of features, we compare our approach with least squares temporal difference learning, direct classification, and with the Bradley-Terry model, fitted using minorization-maximization (MM). The results show that the exact way in which preference learning is applied is critical to achieving high performance. Best results were obtained using a combination of board inversion and pair-wise preference learning. This combination significantly out- performed the others under test, both in terms of move prediction accuracy, and in the level of play achieved when using the learned evaluation function as a move selector during game play.

Full Text