Contextual Embeddings-Based Web Page Categorization Using the Fine-Tune BERT Model

Amit Kumar Nandanwar,Jaytrilok Choudhary

doi:10.3390/sym15020395

Amit Kumar Nandanwar, Jaytrilok Choudhary

Open Access

https://doi.org/10.3390/sym15020395

Copy DOI

Journal: Symmetry	Publication Date: Feb 2, 2023
Citations: 5	License type: CC BY 4.0

Affiliation: Maulana Azad National Institute of Technology

Abstract

The World Wide Web has revolutionized the way we live, causing the number of web pages to increase exponentially. The web provides access to a tremendous amount of information, so it is difficult for internet users to locate accurate and useful information on the web. In order to categorize pages accurately based on the queries of users, methods of categorizing web pages need to be developed. The text content of web pages plays a significant role in the categorization of web pages. If a word’s position is altered within a sentence, causing a change in the interpretation of that sentence, this phenomenon is called polysemy. In web page categorization, the polysemy property causes ambiguity and is referred to as the polysemy problem. This paper proposes a fine-tuned model to solve the polysemy problem, using contextual embeddings created by the symmetry multi-head encoder layer of the Bidirectional Encoder Representations from Transformers (BERT). The effectiveness of the proposed model was evaluated by using the benchmark datasets for web page categorization, i.e., WebKB and DMOZ. Furthermore, the experiment series also fine-tuned the proposed model’s hyperparameters to achieve 96.00% and 84.00% F1-Scores, respectively, demonstrating the proposed model’s importance compared to baseline approaches based on machine learning and deep learning.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Contextual Embeddings-Based Web Page Categorization Using the Fine-Tune BERT Model

Abstract

Talk to us

Similar Papers

More From: Symmetry

Lead the way for us

Similar Papers

Using BERT for Multi-Label Multi-Language Web Page Classification
Codrut-Georgian Artene ... Marius Nicolae Tibeica
-
Codrut-Georgian Artene, et. al.Codrut-Georgian Artene ... Marius Nicolae Tibeica
28 Oct 2021
28 Oct 2021

Automated Credibility Assessment of Web-Based Health Information Considering Health on the Net Foundation Code of Conduct (HONcode): Model Development and Validation Study.
Azadeh Bayani ... Alexandre Ayotte
JMIR formative research | VOL. 7
Azadeh Bayani, et. al.Azadeh Bayani ... Alexandre Ayotte
22 Dec 2023
JMIR formative research | VOL. 7

Bidirectional encoders to state-of-the-art: a review of BERT and its transformative impact on natural language processing
Rajesh Gupta
Информатика. Экономика. Управление - Informatics. Economics. Management | VOL. 3
Rajesh GuptaRajesh Gupta
02 Mar 2024
Информатика. Экономика. Управление - Informatics. Economics. Management | VOL. 3

Automatic Topic-Based Web Page Classification Using Deep Learning
Siti Hawa Apandi ... Norkhairi Ahmad
JOIV : International Journal on Informatics Visualization | VOL. 7
Siti Hawa Apandi, et. al.Siti Hawa Apandi ... Norkhairi Ahmad
30 Nov 2023
JOIV : International Journal on Informatics Visualization | VOL. 7

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Contextual Embeddings-Based Web Page Categorization Using the Fine-Tune BERT Model

Abstract

Talk to us

Similar Papers

More From: Symmetry