On the Effectiveness of Transfer Learning for Code Search

Pasquale Salza,Jian Gu,Harald C Gall,Christoph Schwizer

doi:10.1109/tse.2022.3192755

Abstract

The Transformer architecture and transfer learning have marked a quantum leap in natural language processing, improving the state of the art across a range of text-based tasks. This paper examines how these advancements can be applied to and improve code search. To this end, we pre-train a BERT-based model on combinations of natural language and source code data and fine-tune it on pairs of StackOverflow question titles and code answers. Our results show that the pre-trained models consistently outperform the models that were not pre-trained. In cases where the model was pre-trained on natural language “and” source code data, it also outperforms an information retrieval baseline based on Lucene. Also, we demonstrated that the combined use of an information retrieval-based approach followed by a Transformer leads to the best results overall, especially when searching into a large search pool. Transfer learning is particularly effective when much pre-training data is available and fine-tuning data is limited. We demonstrate that natural language processing models based on the Transformer architecture can be directly applied to source code analysis tasks, such as code search. With the development of Transformer models designed more specifically for dealing with source code data, we believe the results of source code analysis tasks can be further improved.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

On the Effectiveness of Transfer Learning for Code Search

Abstract

Talk to us

Similar Papers

More From: IEEE Transactions on Software Engineering

Lead the way for us

Journal: IEEE Transactions on Software Engineering	Publication Date: Apr 1, 2023
Citations: 10

Similar Papers

Assessing Generalizability of CodeBERT
Xin Zhou ... David Lo
-
Xin Zhou, et. al.Xin Zhou ... David Lo
01 Sep 2021
01 Sep 2021

Text2PyCode: Machine Translation of Natural Language Intent to Python Source Code
Sridevi Bonthu ... M H M Krishna Prasad
-
Sridevi Bonthu, et. al.Sridevi Bonthu ... M H M Krishna Prasad
01 Jan 2020
01 Jan 2020

A time-sensitive historical thesaurus-based semantic tagger for deep semantic annotation
Scott Piao ... Marc Alexander
Computer Speech & Language | VOL. 46
Scott Piao, et. al.Scott Piao ... Marc Alexander
17 May 2017
Computer Speech & Language | VOL. 46

Enriching query semantics for code search with reinforcement learning
Chaozheng Wang ... Yang Liu
Neural Networks | VOL. 145
Chaozheng Wang, et. al.Chaozheng Wang ... Yang Liu
11 Oct 2021
Neural Networks | VOL. 145

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

On the Effectiveness of Transfer Learning for Code Search

Abstract

Talk to us

Similar Papers

More From: IEEE Transactions on Software Engineering