DocILE 2023 Teaser: Document Information Localization and Extraction

Štěpán Šimsa,Yash Patel,Milan Šulc,Matyáš Skalický,Ahmed Hamdi

doi:10.1007/978-3-031-28241-6_69

Abstract

The lack of data for information extraction (IE) from semi-structured business documents is a real problem for the IE community. Publications relying on large-scale datasets use only proprietary, unpublished data due to the sensitive nature of such documents. Publicly available datasets are mostly small and domain-specific. The absence of a large-scale public dataset or benchmark hinders the reproducibility and cross-evaluation of published methods. The DocILE 2023 competition, hosted as a lab at the CLEF 2023 conference and as an ICDAR 2023 competition, will run the first major benchmark for the tasks of Key Information Localization and Extraction (KILE) and Line Item Recognition (LIR) from business documents. With thousands of annotated real documents from open sources, a hundred thousand of generated synthetic documents, and nearly a million unlabeled documents, the DocILE lab comes with the largest publicly available dataset for KILE and LIR. We are looking forward to contributions from the Computer Vision, Natural Language Processing, Information Retrieval, and other communities. The data, baselines, code and up-to-date information about the lab and competition are available at https://docile.rossum.ai/ .

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

DocILE 2023 Teaser: Document Information Localization and Extraction

Abstract

Talk to us

Similar Papers

Lead the way for us

Similar Papers

Proceedings of the ACL-2000 workshop on Recent advances in natural language processing and information retrieval held in conjunction with the 38th Annual Meeting of the Association for Computational Linguistics -
-
-
--
01 Jan 1999
01 Jan 1999

Graph-based Natural Language Processing and Information Retrieval
Rada Mihalcea ... Dragomir Radev
-
Rada Mihalcea, et. al.Rada Mihalcea ... Dragomir Radev
11 Apr 2011
11 Apr 2011

A Survey of Information Extraction Based on Deep Learning
Yang Yang ... Zhiwei Wang
Applied Sciences | VOL. 12
Yang Yang, et. al.Yang Yang ... Zhiwei Wang
27 Sep 2022
Applied Sciences | VOL. 12

Information Extraction from Text
Jing Jiang
-
Jing JiangJing Jiang
01 Jan 2012
01 Jan 2012

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

DocILE 2023 Teaser: Document Information Localization and Extraction

Abstract

Talk to us

Similar Papers