한국어 명사의 지식기반 의미중의성 해소를 위한 효과적인 품사집합

Chul-Heon Kwak,Chung-Hee Lee,Young-Hoon Seo

doi:10.5392/jkca.2016.16.04.418

Abstract

본 논문에서는 지식기반 기법에서 한국어 명사의 의미중의성 해소에 유용한 품사집합을 제시한다. 세종 형태의미분석 말뭉치에서 174,000 문장을 추출하여 테스트 셋으로 이용하고, 표준국어대사전의 뜻풀이와 용례를 이용하여 각 문장의 의미중의성을 해소하였다. 그 결과 전체 테스트 셋의 성능을 가장 좋게하는 15개의 품사집합과 단어별 평균을 가장 높게 하는 17 개의 품사집합이 제시되었다. 실험결과 45 개의 전체 품사집합을 이용하는 것보다 정확도가 최대 12%까지 향상되었다. This paper presents the part-of-speech set which is highly efficient at knowledge-based word sense disambiguation for Korean nouns. 174,000 sentences extracted for test set from Sejong semantic tagged corpus whose sense is based on Standard korean dictionary. We disambiguate selected nouns in test set using glosses and examples in Standard Korean dictionary. 15 part-of-speeches which give the best performance for all test set and 17 part-of-speeches which give the best performance for accuracy average of selected nouns are selected. We obtain 12% more performance by those part-of-speech sets than by full 45 part-of-speech set.

Full Text