Semi-supervised inverted file index approach for approximate nearest neighbor search

Anton Bazdyrev

doi:10.20535/srit.2308-8893.2023.4.05

Abstract

This paper introduces a novel modification to the Inverted File (IVF) index approach for approximate nearest neighbor search, incorporating supervised learning techniques to enhance the efficacy of intermediate clustering and achieve more balanced cluster sizes. The proposed method involves creating clusters using a neural network by solving a task to classify query vectors into the same bucket as their corresponding nearest neighbor vectors in the original dataset. When combined with minimizing the standard deviation of the bucket sizes, the indexing process becomes more efficient and accurate during the approximate nearest neighbor search. Through empirical evaluation on a test dataset, we demonstrate that the proposed semi-supervised IVF index approach outperforms the industry-standard IVF implementation with fixed parameters, including the total number of clusters and the number of clusters allocated to queries. This novel approach has promising implications for enhancing nearest-neighbor search efficiency in high-dimensional datasets across various applications, including information retrieval, natural language search, recommendation systems, etc.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Semi-supervised inverted file index approach for approximate nearest neighbor search

Abstract

Talk to us

Similar Papers

More From: System research and information technologies

Lead the way for us

Journal: System research and information technologies	Publication Date: Dec 26, 2023
License type: cc-by-nd

Similar Papers

Approximate nearest neighbor search on HDD
Noritaka Himei ... Toshikazu Wada
-
Noritaka Himei, et. al.Noritaka Himei ... Toshikazu Wada
01 Sep 2009
01 Sep 2009

A New Cell-Level Search Based Non-Exhaustive Approximate Nearest Neighbor (ANN) Search Algorithm in the Framework of Product Quantization
Yang Wang ... Zhibin Pan
IEEE Access | VOL. 7
Yang Wang, et. al.Yang Wang ... Zhibin Pan
01 Jan 2019
IEEE Access | VOL. 7

Asymmetric Mapping Quantization for Nearest Neighbor Search.
Weixiang Hong ... Xueyan Tang
IEEE transactions on pattern analysis and machine intelligence | VOL. 42
Weixiang Hong, et. al.Weixiang Hong ... Xueyan Tang
27 Jun 2019
IEEE transactions on pattern analysis and machine intelligence | VOL. 42

Effect of Neighborhood Approximation on Downstream Analytics
Saranya Soundar Rajan
-
Saranya Soundar RajanSaranya Soundar Rajan
01 Jan 2019
01 Jan 2019

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Semi-supervised inverted file index approach for approximate nearest neighbor search

Abstract

Talk to us

Similar Papers

More From: System research and information technologies