We know what you're doing! Application detection using thermal data
This study demonstrates that thermal sensor data from smartphones can be used to identify running applications with up to 90% accuracy using neural networks, revealing a security vulnerability where sensitive user information could be inferred through temperature readings, highlighting potential privacy risks.
Modern mobile and embedded devices have high computing power which allows them to be used for multiple purposes. Therefore, applications with low security restrictions may execute on the same device as applications handling highly sensitive information. In such a setup, a security risk occurs if it is possible that an application uses system characteristics to gather information about another application on the same device.In this work, we present a method to leak sensitive runtime information by just using temperature sensor readings of a mobile device. We employ a Convolutional-Neural-Network, Long Short-Term Memory units and subsequent label sequence processing to identify the sequence of executed applications over time. To test our hypothesis we collect data from two state-of-the-art smartphones and real user usage patterns. We show an extensive evaluation using laboratory data, where we achieve labelling accuracies up to 90% and negligible timing error. Based on our analysis we state that the thermal information can be used to compromise sensitive user data and increase the vulnerability of mobile devices. A study based on data collected outside of the laboratory opens up various future directions for research.
- Single Report
- 10.2172/963952
- Aug 1, 2009
The Department of Energy and its contractors store and process massive quantities of sensitive information to accomplish national security, energy, science, and environmental missions. Sensitive unclassified data, such as personally identifiable information (PII), official use only, and unclassified controlled nuclear information require special handling and protection to prevent misuse of the information for inappropriate purposes. Industry experts have reported that more than 203 million personal privacy records have been lost or stolen over the past three years, including information maintained by corporations, educational institutions, and Federal agencies. The loss of personal and other sensitive information can result in substantial financial harm, embarrassment, and inconvenience to individuals and organizations. Therefore, strong protective measures, including data encryption, help protect against the unauthorized disclosure of sensitive information. Prior reports involving the loss of sensitive information have highlighted weaknesses in the Department's ability to protect sensitive data. Our report on Security Over Personally Identifiable Information (DOE/IG-0771, July 2007) disclosed that the Department had not fully implemented all measures recommended by the Office of Management and Budget (OMB) and required by the National Institute of Standards and Technology (NIST) to protect PII, including failures to identify and encrypt PII maintained on information systems. Similarly, the Government Accountability Office recently reported that the Department had not yet installed encryption technology to protect sensitive data on the vast majority of laptop computers and handheld devices. Because of the potential for harm, we initiated this audit to determine whether the Department and its contractors adequately safeguarded sensitive electronic information. The Department had taken a number of steps to improve protection of PII. Our review, however, identified opportunities to strengthen the protection of all types of sensitive unclassified electronic information and reduce the risk that such data could fall into the hands of individuals with malicious intent. In particular, for the seven sites we reviewed: (1) Four sites had either not ensured that sensitive information maintained on mobile devices was encrypted. Or, they had improperly permitted sensitive unclassified information to be transmitted unencrypted through email or to offsite backup storage facilities; (2) One site had not ensured that laptops taken on foreign travel, including travel to sensitive countries, were protected against security threats; and, (3) Although required by the OMB since 2003, we learned that programs and sites were still working to complete Privacy Impact Assessments - analyses designed to examine the risks and ramifications of using information systems to collect, maintain, and disseminate personal information. Our testing revealed that the weaknesses identified were attributable, at least in part, to Headquarters programs and field sites that had not implemented existing policies and procedures requiring protection of sensitive electronic information. In addition, a lack of performance monitoring contributed to the inability of the Department and the National Nuclear Security Administration (NNSA) to ensure that measures were in place to fully protect sensitive information. As demonstrated by previous computer intrusion-related data losses throughout the Department, without improvements, the risk or vulnerability for future losses remains unacceptably high. In conducting this audit, we recognized that data encryption and related techniques do not provide absolute assurance that sensitive data is fully protected. For example, encryption will not necessarily protect data in circumstances where organizational access controls are weak or are circumvented through phishing or other malicious techniques. However, as noted by NIST, when used appropriately, encryption is an effective tool that can, as part of an overall risk-management strategy, enhance security over critical personal and other sensitive information. The audit disclosed that Sandia National Laboratories had instituted a comprehensive program to protect laptops taken on foreign travel. In addition, the Department issued policy after our field work was completed that should standardize the Privacy Impact Assessment process, and, in so doing, provide increased accountability. While these actions are positive steps, additional effort is needed to help ensure that the privacy of individuals is adequately protected and that sensitive operational data is not compromised. To that end, our report contains several recommendations to implement a risk-based protection scheme for the protection of sensitive electronic information.
- Research Article
1
- 10.1016/j.is.2026.102720
- Aug 1, 2026
- Information Systems
The process of linking databases that contain sensitive information about individuals across organisations is an increasingly common requirement in the health and social science research domains, as well as with governments and businesses. The lack of unique entity identifiers means that linking often has to rely on personal details such as names and addresses. Data linkage protocols have been proposed to limit the leakage of sensitive personal information, while privacy-preserving record linkage (PPRL) techniques have been developed to conduct linkage on encoded data. While PPRL techniques are now being employed in real-world applications, the focus of PPRL research has been on the technical aspects of linking sensitive data, such as encoding methods and cryptanalysis attacks. Organisational and human challenges when employing such techniques in practice, however, have not been studied adequately. In this paper, we describe the end-to-end data linkage process and formalise two fundamental types of linkage protocols. We describe the types of parties that participate in such a protocol, and analyse what sensitive information each party can learn from the data it obtains legitimately within the protocol. We also discuss the possible motivations and objectives of an adversary who aims to learn sensitive information from the databases being linked, and show that current PPRL protocols still result in the unintentional leakage of sensitive information. We provide recommendations to help data custodians and other parties involved in data linkage projects to identify and prevent vulnerabilities and make their projects more secure.
- Research Article
4
- 10.1108/k-05-2014-0106
- Jan 12, 2015
- Kybernetes
Purpose – The purpose of this paper is to provide a model for quantitatively analyzing the security profile of an organization’s IT environment. The model considers the security risks associated with stored data, as well as services and devices that can act as channels for data leakages. The authors propose a sensitive information (SI) leakage vulnerability model. Design/methodology/approach – Factors identified as having an impact on the security profile are identified, and scores are assigned based on detailed criteria. These scores are utilized by mathematical models that produce a vulnerability index, which indicates the overall security vulnerability of the organization. In this chapter, the authors verify the model result extracted from SI leakage vulnerability weak index by applying the proposed model to an actual incident that occurred in South Korea in January 2014. Findings – The paper provides vulnerability result and vulnerability index. They are depends on SI state in information systems. Originality/value – The authors identify and define four core variables related to SI leakage: SI, security policy, and leakage channel and value of SI. The authors simplify the SI leakage problem. The authors propose a SI leakage vulnerability model.
- Single Report
7
- 10.4095/288863
- Jan 1, 2010
With the advances in geomatics technologies (i.e. applications, data storage capacities and communication bandwidth), the extensive efforts expended in collecting geospatial data (field surveys, monitoring systems, imagery) and the pervasiveness of the Internet, geomatics users expect easy access to an unprecedented variety of geospatial datasets. This expectation is mirrored in the basic principles of most government data management organizations which encourage the sharing of data for the greater societal good. However, while access to data is increasing, there is also a growing recognition of the extent to which barriers exist in sharing or accessing of sensitive geospatial data. A 2006 Environics survey not only identified the barriers in sharing data (privacy and confidentiality issues, licensing and ownership issues, and liability issues and broader data sensitivities) but found that removing such barriers to sharing data were felt to be the most important data issue to mitigate. In a 2006 workshop facilitated by GeoConnections, communities engaged in land management, environmental impact assessments and sustainable development (collectively referred to as the Environmental and Sustainable Development (E&SD) communities) identified several data sharing issues including the need to develop data-sharing agreements, facilitate open access, conduct further investigations and provide guidance on how to share data of a sensitive nature. To respond to these needs, GeoConnections contracted AMEC Earth and Environmental to conduct research and stakeholder consultation in supported of the development these Best Practices. The purpose of these Best Practices is to educate Data Contributors, Owners, Custodians, Stewards and Consumers of the issues and concepts associated with protecting, sharing and utilizing sensitive geospatial data, with a focus on supporting programs, services, businesses and / or applications related to the Environment and Sustainable Development (E&SD) community. The intention is to provide practical guidance to those interested in developing their own sensitive environmental geospatial data sharing policies and protocols. In reviewing the literature, surveying organizations and practitioners, and through consultation by workshops, it was determined that perspectives range widely on what might be considered sensitive environmental geospatial data. It was also found that there is no consistent mechanism for assessing whether a dataset should be classified as sensitive or not. What was revealed was that the concept of sensitivity changes with context (time and recent events), an organization's regulatory environment (legislation, policy, competition, etc.), jurisdictions and the personal views of Data Contributors/Owners/Custodians and in actuality, there is considerable intertwining of these elements. Anyone who is assessing a dataset to determine whether it should be considered sensitive or not should be aware of these elements and the potential impact on the credibility of their organization if sensitive data is mistreated. The first significant question to be answered is mp;lt;"What is sensitive geospatial data?mp;gt;" and how is it determined to be sensitive or not. What defines data as sensitive is related to legislation, regulations and policies governing an organization as well as standards adopted by the organization. With the focus on the E&SD community, the emphasis is on sensitive mp;lt;"environmentalmp;gt;" geospatial data as a subcategory of sensitive geospatial data. The Guidelines consider environmental geospatial data to be thematic geospatial data that could be used for analysis in areas such as environmental impact assessments, land use planning, land management, sustainable development, resource management, airshed management, etc. Due to the diversity of what can make a dataset sensitive, these Guidelines propose a categorization of sensitivity to assist an assessor (typically the Data Custodian) in understanding which aspect of sensitivity may apply to the dataset they are reviewing. In addition, each organization has to establish and publish its own criteria that allow the assessor to determine whether the dataset being reviewed is sensitive and justify why or why not. These criteria have to be established on an organization by organization basis due to the diversity of the data organizations handle and the specifics of the regulatory environment under which each operates. It is prudent that each organization develops these criteria independent of any specific dataset, establish them in advance of any dataset assessment, document the criteria and have it vetted by an authorized organizational representative (legal or policy). This step is critical in establishing not only the process but provides a documented baseline for justifying the classifying of a dataset as sensitive if challenged at a later date. Understanding these categories will also assist in defining the metrics for establishing what is to be considered sensitive data. Data can generally be categorized as sensitive geospatial data if it meets any of the following criteria: 1. Legislation/Policies/Permits - the data is identified by legislation as requiring safeguarding. The most prominent legislation in this regard is the federal Privacy Act - safeguarding the data is required if an individual can be identified, either directly by georeferenced information (such as the geo-coordinates of an address) or indirectly through the amalgamation of geospatial data and related attributes; 2. Confidentiality - the data is considered confidential by an organization or its use can be economically detrimental to a commercial interest; 3. Natural Resource Protection - the use of the information can result in the degradation of an environmentally significant site or resource; 4. Cultural Protection - the use of the information can result in the degradation of an culturally significant site or resource; or 5. Safety and Security - the information can be used to endanger public health and safety. Numerous articles have identified the need for organizations to establish frameworks for identifying and sharing sensitive data. This need is driven by: The requirement to support open government by making data readily accessible unless there is a legitimate and documented reason not to; Be consistent within an organization and across jurisdictions so that the mechanisms required to share the data are also applied consistently; and Document criteria and processes so that users can search out that the data exists, be made aware of any decisions relating to the safeguarding of the data and know who to contact to request access to safeguarded data. The Best Practices identify basic principles that can be applied to assessing sensitive environmental geospatial datasets in order to classify sensitivity consistently: 1. Unless the dataset is classified as sensitive it can be provided free of restrictions; 2. Information can not be considered sensitive if it is readily available through other sources or if it is not unique; 3. The Data Custodian of the information is the only agency that can determine whether an environmental geospatial dataset is to be classified as sensitive; 4. Data Consumers of sensitive environmental geospatial datasets must honour the restrictions accompanying the information in the form of an agreement, license and/or metadata; and 5. Organizations should document and openly publish their process, criteria and decisions. These Best Practices also present an example decision framework for assessing whether an environmental geospatial dataset is to be classified as sensitive. This framework has been adapted from the US Federal Geographic Data Committee (FGDC) document Guidelines for Providing Appropriate Access to Geospatial Data in Response to Security Concerns. As the FGDC guidelines are primarily concerned with Security and Public Safety, this framework has been modified to accommodate sensitive environmental geospatial data. Once a dataset is defined as sensitive and the organization's regulatory environment is understood then the appropriate mechanisms for sharing become apparent. In most cases instruments such as agreements or licenses are sufficient, in other cases the sensitivity must be removed from the dataset before it is shared and in other cases approval is given to a Data Consumer on a case by case basis. Regardless of the mechanism put in place, it essentially comes down to the Data Custodian trusting that the data will be adequately safeguarded by the Data Consumer and the mechanisms they have in place sufficiently limit the risk of inappropriate treatment of the data. Furthermore, there are numerous examples of data collection programs where Data Contributors contribute sensitive data content on a regular basis to the Data Custodian and the Data Contributors trust that their contributions are adequately safeguarded otherwise future contributions may be terminated. At its core, the successful long term sharing of sensitive environmental geospatial information is about trust, risk management, the credibility of the participating organizations and their overriding desire to disseminate information. Successful sharing of sensitive geospatial information lies within the mechanisms used to: present the underlying knowledge yet remove the sensitivity; the instrument that defines the conditions of use and protection; and, training of participants to ensure that they are cognizant of their roles and responsibilities. Those sharing or accessing geospatial data may use a combination of mechanisms to ensure data is shared and used responsibly and that the credibility of the process is maintained. It is intended that these Best Practices provide the reader with sufficient insight and links to resources in order to assist them in implementing a consistent and document
- Book Chapter
1
- 10.1007/978-3-031-21333-5_50
- Nov 21, 2022
The vast majority of Australian children own a smartphone. High rates of smartphone ownership are associated with high rates of leakage of sensitive information. A child’s time and location patterns are enough to enable someone to build an accurate profile of the child. But children think that their devices already ensure that their sensitive information is secure. The aim of this study was to use off-the-shelf computing devices to educate school age children about leakage of sensitive information from IoT devices. A Distributed Sensor Network (DSN) was assembled and installed around a high school campus in Australia to measure leakage from IoT devices. Children were then informed of the results of the DSN monitoring, during an online safety lesson, which trained them on how to change common default device settings to reduce data leakage. The DSN then again measured the amount of leakage from IoT devices to see if children modified their device settings to reduce leakage of sensitive information. The results of the study revealed that the amount of data leaked from smartphones after the intervention was significantly less than the traffic captured before the intervention, thus confirming the intervention must have had an effect on changing children’s behaviour. It is recommended that this evidence-based program be expanded to other high schools in Australia to empower children to secure their sensitive information.
- Conference Article
17
- 10.1109/icdew.2013.6547438
- Apr 1, 2013
In recent years, there has been rapid growth in mobile devices such as smartphones, and a number of applications are developed specifically for the smartphone market. In particular, there are many applications that are "free" to the user, but depend on advertisement services for their revenue. Such applications include an advertisement module - a library provided by the advertisement service - that can collect a user's sensitive information and transmit it across the network. Such information is used for targeted advertisements, and user behavior statistics. Users accept this business model, but in most cases the applications do not require the user's acknowledgment in order to transmit sensitive information. Therefore, such applications' behavior becomes an invasion of privacy. In our analysis of 1,188 Android applications' network traffic and permissions, 93% of the applications we analyzed connected to multiple destinations when using the network. 61% required a permission combination that included both access to sensitive information and use of networking services. These applications have the potential to leak the user's sensitive information. Of the 107,859 HTTP packets from these applications, 23,309 (22%) contained sensitive information, such as device identification number and carrier name. In an effort to enable users to control the transmission of their private information, we propose a system which, using a novel clustering method based on the HTTP packet destination and content distances, generates signatures from the clustering result and uses them to detect sensitive information leakage from Android applications. Our system does not require an Android framework modification or any special privileges. Thus users can easily introduce our system to their devices, and manage suspicious applications' network behavior in a fine grained manner. Our system accurately detected 94% of the sensitive information leakage from the applications evaluated and produced only 5% false negative results, and less than 3% false positive results.
- Book Chapter
- 10.70593/978-81-988918-5-3_5
- Jun 6, 2025
Using Cloud Services to Store and Process Electronic Health Information Healthcare organizations increasingly are turning to encrypted cloud services that offer the convenience effect associated with the tradeoff between storing sensitive healthcare information on their own servers and utilizing a commercially operated cloud service. Concerns about strict, pervasive security protections in HIPAA and other healthcare regulations are being addressed by creative solutions and technical innovation. New cloud-based healthcare information technology services are helping organizations, physicians, and patients manage the massive volumes of electronic health information being generated and consumed every day (Dilsizian & Siegel, 2014; Lee & Yoon, 2017; Davenport & Kalakota, 2019). Cloud services address the often challenging issues of system accessibility and downtime by allowing electronic health information to be stored and retrieved from remote servers accessible through the Internet. Virtualization technology, coupled with a variety of increasingly affordable storage options, allows cloud service providers to store significant amounts of data in a cost efficient manner, keeping subscription fees lower and making affordability a non-issue for many organizations. Yet, providing security protections, user access controls, monitoring, and identification of data breaches, while normal responsibilities of any organization holding sensitive data, take on added complexity when the data resides on a third-party's server. Cloud service customers are responsible for protecting the security and integrity of their sensitive information and must take extra precautions. Both government and private sector cybersecurity agencies and organizations have issued recommendations on steps to take to ensure the security of sensitive healthcare information stored in the cloud, which are discussed in more detail later. What-if scenarios regarding vulnerability to breaches of sensitive information, the associated potential liability, and the value of reputation for excellent cybersecurity must weigh heavily into the decision whether to utilize cloud services to store and process electronic health information.That is why establishing best practices for managing sensitive healthcare information in the cloud can help organizations to develop effective security risk management processes. Understanding the paths by which security breaches may occur and the harm they could cause is a fundamental component of any risk assessment and decision-making, yet medical organizations may struggle even to identify their most sensitive data, let alone create specific methodologies or best practices for its protection. The available cloud security guidelines and standards may be well known or even widely adopted. However, little guidance is available to assist organizations through the process of protecting sensitive healthcare information stored in the cloud, especially when the information retains its sensitive status only while stored in the cloud. Without best practices, organizations may struggle even to identify their most sensitive data, let alone create specific procedures for its protection.
- Research Article
- 10.1109/icnisc.2015.56
- Jan 23, 2015
The security of computer always depends on ID-password pair to perform identity authentication for users, while users need to consider the mediation between security and memorability on password creation, at most time they prefer to the latter, thereby cause the leakage of sensitive information. In this paper we introduce a novel approach to explain how to analyze sensitive information in user password according to category attribute, at last a scheme is proposed to measure the propagation of sensitive information, mainly focus on individual sensitive information and behavior characteristic of users. The experiment results show that for Chinese users the transform of sensitive information can be evaluated precisely by this means.
- Research Article
- 10.1177/18333583251338406
- May 24, 2025
- Health information management : journal of the Health Information Management Association of Australia
Existing research has long established that direct exposure to patient trauma, such as severe injuries, chronic illnesses and end-of-life care, places clinical healthcare workers at heightened risk of secondary traumatic stress, compassion fatigue and burnout. However, comparatively little attention has been paid to the impact on non-clinical healthcare personnel, such as health information managers (HIMs) who, despite being removed from direct patient care, regularly handle distressing and sensitive patient information. This scoping review explores the literature concerning non-clinical healthcare professionals and the potential impact upon their biopsychosocial-spiritual (BPSS) well-being given prolonged exposure to medical and/or patient records. Arksey and O'Malley's five-stage scoping review strategy was utilised. An initial search of the literature yielded no results specific to HIMs and other non-clinical healthcare professionals. Therefore, the scope of the review was broadened, and a second search of the literature was conducted to explore comparable non-patient/client-facing populations such as transcriptionists. In total 1226 articles were initially identified and 13 articles revealed either a biological, psychological, social and/or spiritual impact when professionals were exposed to traumatic and/or sensitive data. Exploring the roles of comparable non-patient/client-facing populations provides insight into the potential impact that exposure to traumatic and/or sensitive information may have on the health and well-being of HIMs and other non-clinical health professionals.Implications for health information management practice:Further research is recommended to explore the potential BPSS impact that HIMs and other non-clinical health professionals experience due to the exposure of traumatic and/or sensitive information.
- Conference Article
15
- 10.1109/cscloud.2015.21
- Nov 1, 2015
Nowadays, the ubiquity of smart phones make them carry large amounts of personal sensitive information, but at the same time, there are also many Apps in Android APP market that target to collect users' sensitive data. So it becomes quite important to prevent users from the threat of privacy leakage. In this paper, we analyze the Android's privacy protection mechanism, and describe various threats to users' different types of privacy data. After that, we enumerate two ways that can leak sensitive information, and discuss the current solutions and techniques from aspects of privacy protection enhancement and privacy leakage detection. We also make a fine-grained classification for these two aspects, and study the difference between solutions in each category. Finally, we summarize the deficiency of existing research of Android privacy protection and propose the future research direction.
- Research Article
- 10.3390/e27111163
- Nov 15, 2025
- Entropy
Federated representation learning (FRL) is a promising technique for learning shared data representations that capture general features across decentralized clients without sharing raw data. However, there is a risk of sensitive information leakage from learned representations. The conventional differential privacy (DP) mechanism protects the privacy of the whole data by randomizing (adding noise or random response) at the cost of deteriorating learning performance. Inspired by the fact that some data information may be public or non-private and only sensitive information (e.g., race) should be protected, we investigate the information-theoretic protection on specific sensitive information for FRL. To characterize the trade-off between utility and sensitive information leakage, we adopt mutual information-based metrics to measure utility and sensitive information leakage, and propose a method that maximizes the utility performance, while restricting sensitive information leakage less than any positive value via the local DP mechanism. Simulation demonstrates that our scheme can achieve the best utility–leakage trade-off among baseline schemes, and more importantly can adjust the trade-off between leakage and utility by controlling the noise level in local DP.
- Research Article
2
- 10.1109/access.2021.3107601
- Jan 1, 2021
- IEEE Access
Sensitive information leakages from applications are a critical issue in the Android ecosystem. Despite the advance of techniques to secure applications such as packing and obfuscation, a lot of applications are still under the threat of repackaging attacks that inject malicious code and re-distribute applications. Also, as we are becoming more dependent on mobile technologies, more sensitive information is used on our mobile devices. Hence, it is of great importance to reduce the risk of such sensitive information leaks. In this paper, we first present a threat model that attempts to leak users’ sensitive information by using the repackaging attack, named ReMaCi attack. By analyzing the top 8,546 applications downloaded from Google Play Store, we show that 50% of them are really vulnerable to the ReMaCi attack. We, thus, propose a novel, automated static anti-analysis tool, called AmpDroid, for preventing sensitive information leaks. AmpDroid identifies sensitive dataflows and isolates the code that handles the sensitive data from an application. To demonstrate the effectiveness of AmpDroid, we perform the security and performance evaluation of AmpDroid, comparing it with other obfuscation tools.
- Research Article
13
- 10.3390/e22020192
- Feb 7, 2020
- Entropy (Basel, Switzerland)
With the advent of the information age, the effective identification of sensitive information and the leakage of sensitive information during the transmission process are becoming increasingly serious issues. We designed a sensitive information recognition and encryption transmission system based on a decision tree. By training sensitive data to build a decision tree, unknown data can be classified and identified. The identified sensitive information can be marked and encrypted to achieve intelligent recognition and protection of sensitive information. This lays the foundation for the development of an information recognition and encryption transmission system.
- Research Article
1
- 10.1504/ijsn.2023.131599
- Jan 1, 2023
- International Journal of Security and Networks
An information system stores outside data in the backend database to process them efficiently and protects sensitive data from illegitimate flow or unauthorised users. However, most information systems are made in such a way that the sensitive information stored in a database may be leaked explicitly or implicitly during data processing along with the control structure of the program to the output channels. Therefore, sensitive data leakage is one of the crucial security threat. In this paper, the main objective is to detect the illegitimate flow of confidential information in an information system. We propose a framework to detect sensitive information leakage through the data-flow paths of an information system. In particular, to compute the precise set of data-flow paths, we use the non-relational abstract property of the interval domain and the relational abstract property of the polyhedra domain that enables the framework to produce efficient security analysis results.
- Conference Article
4
- 10.1109/ics.2016.0073
- Dec 1, 2016
Opening data of plenty valuable information as public dataset provides great potential treasure to academy or industry. Despite of de-identification process that most of data owner will take before releasing those data, however, the more datasets are opened to public, the more likely personal privacy exposed will be. Previous studies have shown that personal identity and sensitive information might be re-identified by joining two or more de-identified data table with common attributes. According to previous real case studies, even though the personally identifiable information have been de-identified, sensitive personal information still could be uncovered by heterogeneous or cross-domain data joining operation. This kind of privacy re-identification are usually too complicated or obscure to be realized by data owner, not to mention that this problem will be more severe as the scale of data goes large. For the purpose of preventing damage of sensitive information leakage, this paper shows how to use a novel open data de-identification visualization analysis tool (ODD Visualizer) to verify whether there exists sensitive information leakage problem in the target datasets. The high effectiveness, that ODD Visualizer can provide, mainly comes from implementing scalable computing platform as well as developing efficient data visualization technique. Demonstration proves that ODD Visualizer indeed uncovered one real vulnerability of record linkage attack among open datasets available on the internet.