Annif: DIY automated subject indexing using multiple algorithms

Osma Suominen

doi:10.18352/lq.10285

Abstract

Manually indexing documents for subject-based access is a labour-intensive process. We propose using metadata gathered from bibliographic databases to train algorithms that assist librarians in that work. We have developed Annif, an open source tool and microservice for automated subject indexing. After training it with a subject vocabulary and existing metadata, Annif can be used to assign subject headings for new documents. We have tested Annif with different document collections including scientific papers, old scanned books and contemporary e-books, Q&A pairs from an “ask a librarian” service, Finnish Wikipedia, and the archives of a local newspaper. The results of analysing scientific papers and current books have been reassuring, while other types of documents have proved to be more challenging. The current version is based on a combination of existing natural language processing and machine learning tools. By combining multiple approaches and existing open source algorithms, Annif can build on the strengths of individual algorithms and adapt to different settings. With Annif, we expect to improve subject indexing and classification processes especially for electronic documents as well as collections that otherwise would not be indexed at all.

Full Text

Paper version not known

Open DOI Link

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: LIBER Quarterly: The Journal of the Association of European Research Libraries	Publication Date: Jul 29, 2019
Citations: 18	License type: CC BY 4.0

R Discovery Prime

R Discovery Prime

Annif: DIY automated subject indexing using multiple algorithms

Abstract

Talk to us

Similar Papers

More From: LIBER Quarterly: The Journal of the Association of European Research Libraries

Lead the way for us

Similar Papers

Using supervised machine learning for large‐scale classification in management research: The case for identifying artificial intelligence patents
Milan Miric ... Kenneth G Huang
Strategic Management Journal | VOL. 44
Milan Miric, et. al.Milan Miric ... Kenneth G Huang
11 Jul 2022
Strategic Management Journal | VOL. 44

Derivation and Validation of Natural Language Processing Algorithms to Identify and Classify Venous Thrombotic Events from Lower Extremity Duplex Ultrasound Reports
Abdi Abud ... Damon E Houghton
Blood | VOL. 138
Abdi Abud, et. al.Abdi Abud ... Damon E Houghton
05 Nov 2021
Blood | VOL. 138

A Review of Natural Language Processing and Machine Learning Tools Used to Analyze Arabic Social Media
Tarek Kanan ... Hanadi Alshwabka
-
Tarek Kanan, et. al.Tarek Kanan ... Hanadi Alshwabka
01 Apr 2019
01 Apr 2019

Natural Language Processing Tools
Justin F Brunelle ... Chutima Boonthum-Denecke
-
Justin F Brunelle, et. al.Justin F Brunelle ... Chutima Boonthum-Denecke
01 Jan 2012
01 Jan 2012

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Annif: DIY automated subject indexing using multiple algorithms

Abstract

Talk to us

Similar Papers

More From: LIBER Quarterly: The Journal of the Association of European Research Libraries