MAlign: Explainable static raw-byte based malware family classification using sequence alignment

Shoumik Saha,Sadia Afroz,Atif Hasan Rahman

doi:10.1016/j.cose.2024.103714

Shoumik Saha, Sadia Afroz + Show 1 more

Open Access

https://doi.org/10.1016/j.cose.2024.103714

Copy DOI

Abstract

For a long time, malware classification and analysis have been an arms-race between antivirus systems and malware authors. Though static analysis is vulnerable to evasion techniques, it is still popular as the first line of defense in antivirus systems. But most of the static analyzers failed to gain the trust of practitioners due to their black-box nature. We propose MAlign, a novel static malware family classification approach inspired by genome sequence alignment that can not only classify malware families but can also provide explanations for its decision. MAlign encodes raw bytes using nucleotides and adopts genome sequence alignment approaches to create a signature of a malware family based on the conserved code segments in that family, without any human labor or expertise. We evaluate MAlign on two malware datasets, and it outperforms other state-of-the-art machine learning-based malware classifiers (by 4.49%∼0.07%), especially on small datasets (by 19.48%∼1.2%). Furthermore, we explain the generated signatures by MAlign on different malware families illustrating the kinds of insights it can provide to analysts, and show its efficacy as an analysis tool. Additionally, we evaluate its theoretical and empirical robustness against some common attacks. In this paper, we approach static malware analysis from a unique perspective, aiming to strike a delicate balance among performance, interpretability, and robustness.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

MAlign: Explainable static raw-byte based malware family classification using sequence alignment

Abstract

Talk to us

Similar Papers

More From: Computers & Security

Lead the way for us

Journal: Computers & Security	Publication Date: Jan 17, 2024
Citations: 1

Similar Papers

Extracting the Representative API Call Patterns of Malware Families Using Recurrent Neural Network
Iltaek Kwon ... Eul Gyu Im
-
Iltaek Kwon, et. al.Iltaek Kwon ... Eul Gyu Im
20 Sep 2017
20 Sep 2017

Unsuccessful story about few shot malware family classification and siamese network to the rescue
Yude Bai ... Xiaohong Li
-
Yude Bai, et. al.Yude Bai ... Xiaohong Li
27 Jun 2020
27 Jun 2020

Attention-Based Cross-Modal CNN Using Non-Disassembled Files for Malware Classification
Jeongwoo Kim ... Joon-Young Paik
IEEE Access | VOL. 11
Jeongwoo Kim, et. al.Jeongwoo Kim ... Joon-Young Paik
01 Jan 2023
IEEE Access | VOL. 11

MalClassifier: Malware family classification using network flow sequence behaviour
Bushra A Alahmadi ... Ivan Martinovic
-
Bushra A Alahmadi, et. al.Bushra A Alahmadi ... Ivan Martinovic
01 May 2018
01 May 2018

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

MAlign: Explainable static raw-byte based malware family classification using sequence alignment

Abstract

Talk to us

Similar Papers

More From: Computers &amp; Security

More From: Computers & Security