Autopsy Ontology: Supporting Semantics-Based Automated Knowledge Discovery for Digital Forensic Artifacts
The paper introduces the Autopsy Ontology, a formal OWL-based knowledge organization system for digital forensic artifacts, enabling automated reasoning, advanced querying, and semantic knowledge graph generation to enhance analysis and understanding within the Autopsy forensic tool.
Recognizing a lack of formal knowledge organization systems for digital forensic artifacts, this paper proposes the Autopsy Ontology , which, for the first time, defines the de facto standard Autopsy tool’s terminology in OWL. This ontology was designed to be used for automated reasoning over, and advanced querying of, digital forensic artifacts analyzed in Autopsy, and generating semantic knowledge graphs of digital forensic datasets in resource description framework.
- Research Article
1
- 10.3897/biss.3.37412
- Jun 26, 2019
- Biodiversity Information Science and Standards
The landscape of currently existing repositories of specimen data consists of isolated islands, with each applying its own underlying data model. Using standardized protocols such as DarwinCore or ABCD, specimen data and metadata are exchanged and published on web portals such as GBIF. However, data models differ across repositories. This can lead to problems when comparing and integrating content from different systems. for example, in one system there is a field with the label 'determination', in another there is a field with the label 'taxonomic identification'. Both might refer to the same concepts of organism identification process (e.g., 'obi:organism identification assay'; http://purl.obolibrary.org/obo/OBI_0001624), but the intuitive meaning of the content is not clear and the understanding of the providers of the information might differ from that of the users. Without additional information, data integration across isolated repositories is thus difficult and error-prone. As a consequence, interoperability and retrievability of data across isolated repositories is difficult. Linked Open Data (LOD) promises an improvement. URIs can be used for concepts that are ideally created and accepted by a community and that provide machine-readable meanings. LOD thereby supports transfer of data into information and then into knowledge, thus making the data FAIR (Findable, Accessible, Interoperable, Reusable; Wilkinson et al. 2016). Annotating specimen associated data with LOD, therefore, seems to be a promising approach to guarantee interoperability across different repositories. However, all currently used specimen collection management systems are based on relational database systems, which lack semantic transparency and thus do not provide easily accessible, machine-readable meanings for the terms used in their data models. As a consequence, transferring their data contents into an LOD framework may lead to loss or misinterpretation of information. This discrepancy between LOD and relational databases results from the lack of semantic transparency and machine-readability of data in relational databases. Storing specimen collection data as semantic Knowledge Graphs provides semantic transparency and machine-readability of data. Semantic Knowledge Graphs are graphs that are based on the syntax of ‘Subject – Property – Object’ of the Resource Description Framework (RDF). The ‘Subject’ and ‘Property’ position is taken by URIs and the ‘Object’ position can be taken either by a URI or by a label or value. Since a given URI can take the ‘Subject’ position in one RDF statement and the ‘Object’ position in another RDF statement, several RDF statements can be connected to form a directed labeled graph, i.e. a semantic graph. Semantic Knowledge Graphs are graphs in which each described specimen and its parts and properties possess their own URI and thus can be individually referenced. These URIs are used to describe the respective specimen and its properties using the RDF syntax. Additional RDF statements specify the ontology class that each part and property instantiates. The reference to the URIs of the instantiated ontology classes guarantees the Findability, Interoperability, and Reusability of information contained in semantic Knowledge Graphs. Specimen collection data contained in semantic Knowledge Graphs can be made Accessible in a human-readable form through an interface and in a machine-readable form through a SPARQL endpoint (https://en.wikipedia.org/wiki/SPARQL). As a consequence, semantic Knowledge Graphs comply with the FAIR guiding principles. By using URIs for the semantic Knowledge Graph of each specimen in the collection, it is also available as LOD. With semantic Morph·D·Base, we have implemented a prototype to this approach that is based on Semantic Programming. We present the prototype and discuss different aspects of how specimen collection data are handled. By using community created terminologies and standardized methods for the contents created (e.g. species identification) as well as URIs for each expression, we make the data and metadata semantically transparent and communicable. The source code for Semantic Programming and for semantic Morph·D·Base is available from https://github.com/SemanticProgramming. The prototype of semantic Morph·D·Base can be accessed here: https://proto.morphdbase.de.
- Research Article
2
- 10.3897/biss.3.37205
- Jun 19, 2019
- Biodiversity Information Science and Standards
Currently, morphological data and metadata are still mostly published as unstructured free texts, which lack semantic transparency, cannot be parsed by computers, and do not comply with the FAIR (Findable, Accessible, Interoperable, Reusable; Wilkinson et al. (2016) data principles, thus hampering their reuse by non-experts and their integration across many fields in the life sciences. With an ever-increasing amount of available ontologies and the development of adequate semantic technology, however, a solution to this problem becomes available. Instead of free text descriptions, morphological data and metadata can be recorded, stored, and communicated through the Web in the form of Resource Description Framework (RDF) triple statements that use the ‘Subject – Property – Object’ syntax of RDF and URIs of ontology classes and properties as well as URIs for individual entities as terminology. Since a given URI can take the ‘Subject’ position in one and the ‘Object’ position in another RDF statement, several triples can be linked to form a highly formalized and structured directed graph (semantic graph). After introducing an instance-based approach of recording morphological descriptions and their accompanying metadata as semantic knowledge graphs (i.e. Anatomy Knowledge Graphs), we propose a knowledge graph template pattern for each type of anatomical observation and a pattern for documenting metadata. The use of template patterns for knowledge graphs provides Interoperability and Reusability of comparable anatomical observations and of their accompanying metadata and a means to meaningfully visualize information contained in semantic graphs in a user-friendly HTML representation. Stored in a tuple store, Anatomy Knowledge Graphs become Findable and Accessible through the store’s SPARQL endpoint. As a consequence, anatomy data and metadata documented as Anatomy Knowledge Graphs in a tuple store are FAIR. Finally, we suggest a general scheme of how to efficiently organize Anatomy Knowledge Graphs in a tuple store framework based on instances of named graphs, with each individual named graph instantiating an ontology class that relates to a particular type of observation (e.g., weight measurement named graph class). A named graph is a fourth element in an RDF statement (‘Subject – Property – Object – Named Graph’), turning the triple into a quadruple. All RDF statements that share the same URI in the ‘Named Graph’ position belong to the same named graph. The use of named graph resources allows meaningful fragmentation of the contents of an Anatomy Knowledge Graph (Fig. 1), which in turn enables subsequent specification of all kinds of data views for managing and accessing morphological data and metadata. This scheme has been implemented in the description module of the prototype for semantic Morph∙D∙Base.
- Conference Article
3
- 10.1109/compsac54236.2022.00025
- Jun 1, 2022
This paper presents a systematic approach to designing a series of digital forensics instructional materials to address the severe shortage of active learning materials in the digital forensics community. The materials include real-world scenario-based case studies, a set of hands-on problem-driven labs for each case study, and an integrated forensic investigation environment. In this paper, we first clarify some fundamental concepts related to digital forensics, such as digital forensic artifacts, artifact generators, and evidence. We then re-categorize knowledge units of digital forensics based on the artifact generators for measuring the coverage of learning outcomes and topics. Finally, we utilize a real-world cybercrime scenario to demonstrate how knowledge units, digital forensics topics, concepts, artifacts, and investigation tools can be infused into each lab through active learning. The repository of the instructional materials is publicly available on GitHub. It has gained nearly 600 stars and 22k views within several months.
- Research Article
1
- 10.1016/j.chiabu.2024.106908
- Jun 25, 2024
- Child Abuse & Neglect
Testing a hybrid risk assessment model: Predicting CSAM offender risk from digital forensic artifacts
- Research Article
6
- 10.1016/j.fsidi.2023.301624
- Sep 18, 2023
- Forensic Science International: Digital Investigation
Post-mortem digital forensic analysis of the Garmin Connect application for Android
- Research Article
- 10.22042/isecure.2013.4.2.5
- Jul 1, 2012
The products of graphic design applications, leave behind traces of digital information which can be used during a digital forensic investigation in cases where counterfeit documents have been created. This paper analyzes the digital forensics involved in the creation of counterfeit documents. This is achieved by first recognizing the digital forensic artifacts left behind from the use of graphic design applications, and then analyzing the files associated with these applications. When analyzing digital forensic artifacts generated by an application, the specific focus is on determining whether the graphic design application was installed, whether the application was used, and determining whether an association can be made between the application’s actions and such a digital crime. This is accomplished by locating such information from the registry, log files and prefetch files. The file analysis involves analyzing files associated with these applications for file signatures and metadata. In the end it becomes possible to determine if a system has been used for creating counterfeit documents or not.
- Research Article
4
- 10.1016/j.fsidi.2023.301555
- May 11, 2023
- Forensic Science International: Digital Investigation
This paper studies the post-mortem digital forensic artifacts left by the Android ZeppLife (formerly MiFit) mobile application when used in conjunction with a Xiaomi MiBand 6. The MiBand 6 is a low-cost smart band device with several sensors that allow for health and activity monitoring, collecting metrics such as heart rate, blood oxygen saturation level, and step count. The device communicates via Bluetooth Low Energy with the ZeppLife application, which displays its data, provides some controls, and acts as a bridge to the Internet.We study, from a digital forensics perspective, the Android version of the mobile application in a rooted smartphone. For this purpose, we analyze the data repositories, namely its databases and XML files, and correlate the data on the smartphone with the corresponding usage of the Mi Band device. The paper also presents two open-source scripts we have developed to ease the task of forensic practitioners dealing with ZeppLife/MiBand 6: ZL_std and ZL_autopsy. The former refers to a Python 3 script that extracts high-level views of Zepp Life data through the command-line, whereas the latter is a module that integrates ZL_std functionalities within the popular open-source Autopsy digital forensic software. Data stored on the Android companion device of a MiBand 6 might include GPS coordinates, events and alarms, and biometric data such as heart rate, sleep time, and fitness activity, which can be valuable digital forensic artifacts.
- Book Chapter
- 10.4018/979-8-3373-1102-9.ch006
- Jan 17, 2025
The rapid rise of cybercrime, fueled by increased reliance on digital systems, underscores the importance of effective digital forensics. This research investigates forensic artifacts within two widely used operating systems, Windows and Linux, to offer a detailed comparison of their impact on forensic practices. Digital forensic artifacts, such as user activity logs, system logs, application data, and volatile memory artifacts, are key sources of evidence that investigators rely on to reconstruct events, identify threats, and secure convictions in cybercrime cases.
- Research Article
1
- 10.3390/electronics13142864
- Jul 20, 2024
- Electronics
Although Three-dimensional (3D) printers have legitimate applications in various fields, they also present opportunities for misuse by criminals who can infringe upon intellectual property rights, manufacture counterfeit medical products, or create unregulated and untraceable firearms. The rise of affordable 3D printers for general consumers has exacerbated these concerns, making it increasingly vital for digital forensics investigators to identify and analyze vital artifacts associated with 3D printing. In our study, we focus on the identification and analysis of digital forensic artifacts related to 3D printing stored in both Linux and Windows operating systems. We create five distinct scenarios and gather data, including random-access memory (RAM), configuration data, generated files, residual data, and network data, to identify when 3D printing occurs on a device. Furthermore, we utilize the 3D printing slicing software Ultimaker Cura version 5.7 and RepetierHost version 2.3.2 to complete our experiments. Additionally, we anticipate that criminals commonly engage in anti-forensics and recover valuable evidence after uninstalling the software and deleting all other evidence. Our analysis reveals that each data type we collect provides vital evidence relating to 3D printing forensics.
- Book Chapter
4
- 10.1016/b978-0-12-822468-7.00012-2
- Jan 1, 2021
- Web Semantics
Chapter 6 - Resource description framework based semantic knowledge graph for clinical decision support systems
- Research Article
- 10.24966/flis-733x/100106
- Jan 1, 2025
- Journal of Forensic, Legal & Investigative Sciences
Electronic evidence in cyber forensics can be any digital data that is useful to investigate cybercrimes, including logs, metadata, files, emails, and internet history. Computer crime evidence or digital evidence collection is important, and it has been utilized in investigating crimes. This paper deals with planning computer crime evidence collection. It includes crime evidence collection (approaches, actions, and steps), modeling methods for digital forensics, image forensics, forensic evidence collection using fingerprinting, data collection and rig derivation, digital forensic artifacts and artifact cataloging, anti-forensic types and techniques, cross-border electronic evidence, and cross-border criminal investigations. It is necessary for forensic methods and tools to keep up with new forms of digital evidence and cyberattacks.
- Book Chapter
13
- 10.1016/b978-0-12-805303-4.00014-9
- Jan 1, 2017
- Contemporary Digital Forensic Investigations of Cloud and Mobile Applications
Chapter 14 - Residual Cloud Forensics: CloudMe and 360Yunpan as Case Studies
- Research Article
47
- 10.14778/2824032.2824124
- Aug 1, 2015
- Proceedings of the VLDB Endowment
The Resource Description Framework (RDF) is a graph-based data model promoted by the W3C as the standard for Semantic Web applications. Its associated query language is SPARQL. RDF graphs are often large and varied , produced in a variety of contexts, e.g., scientific applications, social or online media, government data etc. They are heterogeneous , i.e., resources described in an RDF graph may have very different sets of properties. An RDF resource may have: no types, one or several types (which may or may not be related to each other). RDF Schema (RDFS) information may optionally be attached to an RDF graph, to enhance the description of its resources. Such statements also entail that in an RDF graph, some data is implicit. According to the W3C RDF and SPARQL specification, the semantics of an RDF graph comprises both its explicit and implicit data ; in particular, SPARQL query answers must be computed reflecting both the explicit and implicit data. These features make RDF graphs complex, both structurally and conceptually. It is intrinsically hard to get familiar with a new RDF dataset, especially if an RDF schema is sparse or not available at all.
- Research Article
1
- 10.3897/biss.3.37206
- Jun 19, 2019
- Biodiversity Information Science and Standards
We would like to present FAIR Research Data: Semantic Knowledge Graph Infrastructure for the Life Sciences (in short, FAIR.ReD), a project initiative that is currently being evaluated for funding. FAIR.ReD is a software environment for developing data management solutions according to the FAIR (Findable, Accessible, Interoperable, Reusable; Wilkinson et al. 2016) data principles. It utilizes what we call a Data Sea Storage, which employs the idea of Data Lakes to decouple data storage from data access but modifies it by storing data in a semantically structured format as either semantic graphs or semantic tables, instead of storing them in their native form. Storage follows a top-down approach, resulting in a standardized storage model, which allows sharing data across all FAIR.ReD Knowledge Graph Applications (KGAs) connected to the same Sea, with newly developed KGAs having automatically access to all contents in the Sea. In contrast access and export of data follows a bottom-up approach that allows the specification of additional data models to meet the varying domain-specific and programmatic needs for accessing structured data. The FAIR.ReD engine enables bidirectional data conversion between the two storage models and any additional data model, which will substantially reduce conversion workload for data-rich institutes (Fig. 1). Moreover, with the possibility to store data in semantic tables, FAIR.ReD provides high performance storage for incoming data streams such as sensory data. FAIR.ReD KGAs are modularly organized. Modules can be edited using the FAIR.ReD editor and combined to form coherent KGAs. The editor allows domain experts to develop their own modules and KGAs without any programming experience required, thus also allowing smaller projects and individual researchers to build their own FAIR data management solution. Contents from FAIR.ReD KGAs can be published under a Creative Commons license as documents, micropublications, or nanopublications, each receiving their own DOI. A publication-life-cycle is implemented in FAIR.ReD and allows updating published contents for corrections or additions without overwriting the originally published version. Together with the fact that data and metadata are semantically structured and machine-readable, all contents from FAIR.ReD KGAs will comply with the FAIR Guiding Principles. Due to all FAIR.Red KGAs providing access to semantic knowledge graphs in both a human-readable and a machine-readable version, FAIR.ReD seamlessly integrates the complex RDF (Resource Description Framework) world with a more intuitively comprehensible presentation of data in form of data entry forms, charts, and tables. Guided by use cases, the FAIR.ReD environment will be developed using semantic programming where the source code of an application is stored in its own ontology. The set of source code ontologies of a KGA and its modules provides the steering logic for running the KGA. With this clear separation of steering logic from interpretation logic, semantic programming follows the idea of separating main layers of an application, analog to the separation of interpretation logic and presentation logic. Each KGA and module is specified exactly in this way and their source code ontologies stored in the Data Sea. Thus, all data and metadata are semantically transparent and so is the data management application itself, which substantially improves their sustainability on all levels of data processing and storing.
- Research Article
7
- 10.1111/1556-4029.14820
- Jul 30, 2021
- Journal of Forensic Sciences
The prevalence of online child pornography is a major societal issue. The criminal justice system has struggled with assessing the risk of individuals involved in online sexual offenses against children, especially when it involves the possession of child pornography. Research suggests there are different categories of offenders involved in this type of behavior (e.g., Online Child Pornography Offenders, Dual Offenders, Contact Offenders), with each category having different motivations, contributing factors, and levels of risk to re-offend or escalate their criminal behavior to more serious offenses (i.e., collecting pictures to contact offending). Determining the risk that individuals involved in online sexual offenses against children pose to re-offend or escalate their criminal behavior has been problematic. Traditional sexual offender risk measures have lower predictive validity when dealing with online child pornography offenders. This article discusses the need for a formalized hybrid risk assessment model that combines the current online sex offenses against children risk measures with digital forensics artifact analysis. The evidence derived from digital forensics artifact analysis can supplement the predictive risk factors obtained from these risk assessment tools, thus increasing the reliability and validity of the risk assessment.