Автоматическое определение сходства Javadoc-комментариев

Dmitry Vladimirovich Koznov,Pavel Isaakovich Braslavski,Ekaterina Iurevna Ledeneva,Dmitry Vadimovich Luciv

doi:10.15514/ispras-2023-35(4)-10

Dmitry Vladimirovich Koznov, Pavel Isaakovich Braslavski + Show 2 more

Open Access

https://doi.org/10.15514/ispras-2023-35(4)-10

Copy DOI

Abstract

Code comments are an essential part of software documentation. Many software projects suffer the problem of low-quality comments that are often produced by copy-paste. In case of similar methods, classes, etc. copy-pasted comments with minor modifications are justified. However, in many cases this approach leads to degraded documentation quality and, subsequently, to problematic maintenance and development of the project. In this study, we address the problem of near-duplicate code comments detection, which can potentially improve software documentation. We have conducted a thorough evaluation of traditional string similarity metrics and modern machine learning methods. In our experiment, we use a collection of Javadoc comments from four industrial open-source Java projects. We have found out that LCS (Longest Common Subsequence) is the best similarity algorithm taking into account both quality (Precision 94%, Recall 74%) and performance.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Автоматическое определение сходства Javadoc-комментариев

Abstract

Talk to us

Similar Papers

More From: Proceedings of the Institute for System Programming of the RAS

Lead the way for us

Journal: Proceedings of the Institute for System Programming of the RAS	Publication Date: Jan 1, 2023
License type: cc-by

Similar Papers

Benchmarking machine learning methods for modeling physical properties of ionic liquids
Igor Baskin ... Yair Ein-Eli
Journal of Molecular Liquids | VOL. 351
Igor Baskin, et. al.Igor Baskin ... Yair Ein-Eli
29 Jan 2022
Journal of Molecular Liquids | VOL. 351

Role of Artificial Intelligence and Machine Learning in Nanosafety.
David A Winkler
Small | VOL. 16
David A WinklerDavid A Winkler
15 Jun 2020
Small | VOL. 16

A String Similarity Evaluation for Healthcare Ontologies Alignment to HL7 FHIR Resources
Athanasios Kiourtis ... Dimosthenis Kyriazis
-
Athanasios Kiourtis, et. al.Athanasios Kiourtis ... Dimosthenis Kyriazis
01 Jan 2019
01 Jan 2019

Estimation of Paddy Rice Nitrogen Content and Accumulation Both at Leaf and Plant Levels from UAV Hyperspectral Imagery
Li Wang ... Qiong Zheng
Remote Sensing | VOL. 13
Li Wang, et. al.Li Wang ... Qiong Zheng
27 Jul 2021
Remote Sensing | VOL. 13

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Автоматическое определение сходства Javadoc-комментариев

Abstract

Talk to us

Similar Papers

More From: Proceedings of the Institute for System Programming of the RAS