영한 기계 번역에서 미가공 텍스트 데이터를 이용한 대역어 선택 중의성 해소

Yuseop Kim ,Jeong-Ho Chang

doi:10.3745/kipstb.2004.11b.6.749

Abstract

본 논문에서는 미가공 말뭉치 데이터를 활용하여 영한 기계번역 시스템의 대역어 선택 시 발생하는 중의성을 해소하는 방법을 제안한다. 이를 위하여 은닉 의미 분석(Latent Semantic Analysis : LSA)과 확률적 은닉 의미 분석(Probabilistic LSA : PLSA)을 적용한다. 이 두 기법은 텍스트 문단과 같은 문맥 정보가 주어졌을 때, 이 문맥이 내포하고 있는 복잡한 의미 구조를 표현할 수 있다 본 논문에서는 이들을 사용하여 언어적인 의미 지식(Semantic Knowledge)을 구축하였으며 이 지식은 결국 영한 기계번역에서의 대역어 선택 시 발생하는 중의성을 해소하기 위하여 단어간 의미 유사도를 추정하는데 사용된다. 또한 대역어 선택을 위해서는 미리 사전에 저장된 문법 관계를 활용하여야 한다. 본 논문에서는 이러한 대역어 선택 시 발생하는 데이터 희소성 문제를 해소하기 위하여 k-최근점 학습 알고리즘을 사용한다. 그리고 위의 두 모델을 활용하여 k-최근점 학습에서 필요한 예제 간 거리를 추정하였다. 실험에서는, 두 기법에서의 은닉 의미 공간을 구성하기 위하여 TREC 데이터(AP news)론 활용하였고, 대역어 선택의 정확도를 평가하기 위하여 Wall Street Journal 말뭉치를 사용하였다. 그리고 은닉 의미 분석을 통하여 대역어 선택의 정확성이 디폴트 의미 선택과 비교하여 약 10% 향상되었으며 PLSA가 LSA보다 근소하게 더 좋은 성능을 보였다. 또한 은닉 공간에서의 축소된 벡터의 차원수와 k-최근점 학습에서의 k값이 대역어 선택의 정확도에 미치는 영향을 대역어 선택 정확도와의 상관관계를 계산함으로써 검증하였다.젝트의 성격에 맞도록 필요한 조정만을 통하여 품질보증 프로세스를 확립할 수 있다. 개발 된 패키지의 효율적인 활용이 내조직의 소프트웨어 품질보증 구축에 투입되는 공수 및 어려움을 줄일 것으로 기대된다.도가 증가할 때 구기자 열수 추출 농축액은 <TEX>$1.6182{\sim}2.0543$</TEX>, 혼합구기자 열수 추출 농축액은 <TEX>$1.7057{\sim}2.1462{\times}10^7\;J/kg{\cdot}mol$</TEX>로 증가하였다. 이와 같이 구기자 열수 추출 농축액과 혼합구기자 열수 추출 농축액의 리올리지적 특성에 큰 차이를 나타내지는 않았다. security simultaneously.% 첨가시 pH 5.0, 7.0 및 8.0에서 각각 대조구의 57, 413 및 315% 증진되었다. 거품의 열안정성은 15분 whipping시, pH 4.0(대조구, 30.2%) 및 5.0(대조구, 23.7%)에서 각각 <TEX>$0{\sim}38.0$</TEX> 및 <TEX>$0{\sim}57.0%$</TEX>이었고 pH 7.0(대조구, 39.6%) 및 8.0(대조구, 43.6%)에서 각각 <TEX>$0{\sim}59.4$</TEX> 및 <TEX>$36.6{\sim}58.4%$</TEX>이었으며 sodium alginate 첨가시가 가장 양호하였다. 전체적으로 보아 거품안정성이 높은 것은 열안정성도 높은 경향이며, 표면장력이 낮으면 거품형성능이 높아지고, 비점도가 높으면 거품안정성 및 열안정성이 높아지는 경향이 있었다.protocol.eractions between application agents that are developed using different In this paper, we propose a new method utilizing only raw corpus without additional human effort for disambiguation of target word selection in English-Korean machine translation. We use two data-driven techniques; one is the Latent Semantic Analysis(LSA) and the other the Probabilistic Latent Semantic Analysis(PLSA). These two techniques can represent complex semantic structures in given contexts like text passages. We construct linguistic semantic knowledge by using the two techniques and use the knowledge for target word selection in English-Korean machine translation. For target word selection, we utilize a grammatical relationship stored in a dictionary. We use k- nearest neighbor learning algorithm for the resolution of data sparseness Problem in target word selection and estimate the distance between instances based on these models. In experiments, we use TREC data of AP news for construction of latent semantic space and Wail Street Journal corpus for evaluation of target word selection. Through the Latent Semantic Analysis methods, the accuracy of target word selection has improved over 10% and PLSA has showed better accuracy than LSA method. finally we have showed the relatedness between the accuracy and two important factors ; one is dimensionality of latent space and k value of k-NT learning by using correlation calculation.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

영한 기계 번역에서 미가공 텍스트 데이터를 이용한 대역어 선택 중의성 해소

Abstract

Talk to us

Similar Papers

More From: The KIPS Transactions:PartB

Lead the way for us

Similar Papers

A comparative evaluation of data-driven models in translation selection of machine translation
Yu-Seop Kim ... Byoung-Tak Zhang
-
Yu-Seop Kim, et. al.Yu-Seop Kim ... Byoung-Tak Zhang
01 Jan 2002
01 Jan 2002

Target Word Selection Using WordNet and Data-Driven Models in Machine Translation
Yuseop Kim ... Byoung-Tak Zhang
-
Yuseop Kim, et. al.Yuseop Kim ... Byoung-Tak Zhang
01 Jan 2002
01 Jan 2002

An Empirical Study on Dimensionality Optimization in Text Mining for Linguistic Knowledge Acquisition
Yu-Seop Kim ... Jeong-Ho Chang
-
Yu-Seop Kim, et. al.Yu-Seop Kim ... Jeong-Ho Chang
01 Jan 2003
01 Jan 2003

Comparing the Performance of Latent Semantic Analysis and Probability Latent Semantic Analysis Models on Autoscoring Essay Tasks
Xiaohua Ke ... Haijiao Luo
-
Xiaohua Ke, et. al.Xiaohua Ke ... Haijiao Luo
01 Jan 2017
01 Jan 2017

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

영한 기계 번역에서 미가공 텍스트 데이터를 이용한 대역어 선택 중의성 해소

Abstract

Talk to us

Similar Papers

More From: The KIPS Transactions:PartB