Multimedia Search Research Articles

Over the past decade, deep learning has emerged as a revolutionary technology for the analysis of visual data, particularly images. This master's thesis focuses on deep learning approaches to image classification, which is a key task in many applications using visual data analysis. A state-of-the-art deep learning model, namely the Vision Transformer (ViT), is explored for image classification. ViT is trained using transfer-learning techniques on a new dataset of over 350,000 photographs of European buildings in eight cities, obtained across two separate flights from a drone-mounted camera. Initial results demonstrate that models pre-trained on large datasets such as JFT-300M can achieve performance competitively with the fine-tuning of models trained from scratch on smaller datasets and that ViT outperforms convolutional neural networks for drone-captured images. Further, the prospects of deep learning for image classification are discussed, highlighting the potential impact of new research directions within the architectural vision transformer domain (e.g., Swin-Transformer, CrossViTs, T2T-vision Transformer) and new training techniques (e.g., Vision-Language Pre-training models, multi-modality input). The exponential increase in data generated by cameras, mobile devices, and Internet-of-Things (IoT) sensors has escalated the need for automated processing and analysis of visual data. Furthermore, images and video frames are a popular medium for data collection across various domains, including commercial and industrial. Image classification, or finding the most relevant label for a given photograph, is one key task in many applications using visual data analysis. Popular applications include multimedia search engines, mobile applications navigating to points of interest (POI), and anomaly detection in industrial cameras. As a consequence, many datasets have been assembled, containing millions of photographs collected and labeled according to city, object, or scene. Deep neural networks trained end-to-end directly on pixels have become state-of-the-art image classification technology. More recently, architectures based solely on attention mechanisms, eschewing convolutions, have challenged the long-standing dominance of convolutional neural networks.

Read full abstract

In this paper, a method for knowledge expansion of metadata using script mining analysis for multimedia recommendation systems is proposed. The method allows the extraction of new metadata and knowledge expansion through the mining analysis of multimedia scripts, which include a large amount of information. The scripts are collected by a Web crawler based on Python. From the collected scripts, hidden information is extracted through keyword analysis and sentiment analysis. In keyword analysis, scripts, unlike general documents, show a high frequency of names of characters or proper nouns. Such names or proper nouns are not frequently used in other media content, and therefore, their importance is high. Frequently, they are already offered in the conventional metadata, and consequently cause information duplication. Accordingly, term frequency–inverse document and metadata frequency (TF–IDMF), which considers the frequency of metadata in general term frequency–inverse document frequency (TF–IDF), is used. Thus, the importance of the names of characters or proper nouns in scripts can be decreased. Because the keywords for the extracted scripts are in fact included in the scripts, they can be used for precise multimedia search and recommendation. In sentiment analysis, the AFINN lexicon and the Bing lexicon are utilized to scan words in a script. The Bing lexicon is used to examine whether the words in the entire script are positive or negative. Then, the total numbers of positive words and negative words are used to calculate the representative sentiment of the script. The AFINN lexicon includes approximately 170 sentiment words, the negative or positive sentiment of which is presented in the range − 5 to +5. One script is divided into 100 sentences, and then, the representative sentiment in each sentence is evaluated as either positive or negative. Through script scanning, the flow of sentiment in multimedia streams can be discovered. The Bing lexicon categorizes words into positive, negative, and neutral sentiments. Through script scanning, the words included in each category can be quantified. Depending on the result of the script sentiment analysis, a different sentence embedding method based on inter-sentence similarity is used to cluster similar media. The results of the keyword analysis and sentiment analysis of a script are added to the metadata in a new column in a knowledge base to expand knowledge. To evaluate the significance of multimedia recommendations, keywords and sentiment information are used, and then, the similarity and clustering of the extracted media are assessed. As a result, script mining analysis based on the attributes that include actual information of media is considerably better than that based on types or a range of metadata attributes. Therefore, the proposed knowledge expansion method achieves significant results and shows an excellent performance in multimedia recommendation.

Read full abstract

Multimedia Search Research Articles

Related Topics

Articles published on Multimedia Search

Improving semantic video retrieval models by training with a relevance-aware online mining strategy

Implementation of multimedia search & management system based on remote education

ViSTORY: Effective Video Storyboard Generation with Visual Keyframes using Discrete Cosine Transform

Deep Learning Approaches To Image Classification: Exploring The Future Of Visual Data Analysis

Intelligent Auxiliary Artificial Wood Plank Pattern Design Based on the Subject Search Algorithm of Multimedia Resources

Robust Video Hashing Based on Multidimensional Scaling and Ordinal Measures

An architecture for non-linear discovery of aggregated multimedia document web search results.

An Effectual Video Indexing and Retrieval Model: A Comparative Study

ASTS: attention based spatio-temporal sequential framework for movie trailer genre classification

Multimedia Intelligence: When Multimedia Meets Artificial Intelligence

Evaluating Multimedia and Language Tasks.

Knowledge expansion of metadata using script mining analysis in multimedia recommendation

Multi-hash chain based multimedia search technology optimized for distributed environments using IoT devices

Artificial Intelligence Recognition Simulation of 3D Multimedia Visual Image Based on Sparse Representation Algorithm

ПОШУК МУЛЬТИМЕДІЙНОЇ ІНФОРМАЦІЇ НА ОСНОВІ НЕЙРОННИХ МЕРЕЖ

Automatic text location of multimedia video for subtitle frame

Efficient Supervised Discrete Multi-View Hashing for Large-Scale Multimedia Search

On Machine Learning and Knowledge Organization in Multimedia Information Retrieval

A personalized cloud engine for multimedia search based on binary ant colony algorithm

Deep Attention-Guided Hashing

Lead the way for us

Editage

Paperpal

R Discovery

Mind the Graph

Multimedia Search Research Articles

Related Topics

Articles published on Multimedia Search

Improving semantic video retrieval models by training with a relevance-aware online mining strategy

Implementation of multimedia search &amp; management system based on remote education

ViSTORY: Effective Video Storyboard Generation with Visual Keyframes using Discrete Cosine Transform

Deep Learning Approaches To Image Classification: Exploring The Future Of Visual Data Analysis

Intelligent Auxiliary Artificial Wood Plank Pattern Design Based on the Subject Search Algorithm of Multimedia Resources

Robust Video Hashing Based on Multidimensional Scaling and Ordinal Measures

An architecture for non-linear discovery of aggregated multimedia document web search results.

An Effectual Video Indexing and Retrieval Model: A Comparative Study

ASTS: attention based spatio-temporal sequential framework for movie trailer genre classification

Multimedia Intelligence: When Multimedia Meets Artificial Intelligence

Evaluating Multimedia and Language Tasks.

Knowledge expansion of metadata using script mining analysis in multimedia recommendation

Multi-hash chain based multimedia search technology optimized for distributed environments using IoT devices

Artificial Intelligence Recognition Simulation of 3D Multimedia Visual Image Based on Sparse Representation Algorithm

ПОШУК МУЛЬТИМЕДІЙНОЇ ІНФОРМАЦІЇ НА ОСНОВІ НЕЙРОННИХ МЕРЕЖ

Automatic text location of multimedia video for subtitle frame

Efficient Supervised Discrete Multi-View Hashing for Large-Scale Multimedia Search

On Machine Learning and Knowledge Organization in Multimedia Information Retrieval

A personalized cloud engine for multimedia search based on binary ant colony algorithm

Deep Attention-Guided Hashing

Implementation of multimedia search & management system based on remote education