Abstract

Despite their educational level and professional qualifications, an important percentage of highly-skilled migrants and refugees find employment in low-skill vocations throughout the world. Typical vocational domains include agriculture, cooking, crafting, construction, and hospitality. As a first step towards developing an educational tool for helping such underprivileged communities become acquainted with the sublanguage of their vocational domain in their host country, automatic domain identification among the aforementioned domains was attempted in this paper, using domain-specific textual data. Wikis and social networks provide a valuable data source for data mining, Natural Language Processing and machine learning tasks. Wikipedia articles, in regard to these domains, were collected and processed in order to create a novel text data set. Extracted linguistic features were used in the experiments with Random Forest combined with Adaboost, and Gradient Boosted Trees. The machine learning models achieved high performance in vocational domain identification (up to 99.93% accuracy).

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Disclaimer: All third-party content on this website/platform is and will remain the property of their respective owners and is provided on "as is" basis without any warranties, express or implied. Use of third-party content does not indicate any affiliation, sponsorship with or endorsement by them. Any references to third-party content is to identify the corresponding services and shall be considered fair use under The CopyrightLaw.