Data clustering: a fundamental method in data science and management
Data clustering: a fundamental method in data science and management
- Research Article
2
- 10.1002/bult.2010.1720360611
- Aug 1, 2010
- Bulletin of the American Society for Information Science and Technology
ASIS&T research data access and preservation summit: Conference summary
- Research Article
52
- 10.1007/s11214-014-0128-5
- Feb 4, 2015
- Space Science Reviews
The four Magnetospheric Multiscale (MMS) spacecraft will collect a combined volume of approximately 100 gigabits per day of particle and field data. On average, only 4 gigabits of that volume can be transmitted to the ground. To maximize the scientific value of each transmitted data segment, MMS has developed the Science Operations Center (SOC) to manage science operations, instrument operations, and selection, downlink, distribution, and archiving of MMS science data sets. The SOC is managed by the Laboratory for Atmospheric and Space Physics (LASP) in Boulder, Colorado and serves as the primary point of contact for community participation in the mission. MMS instrument teams conduct their operations through the SOC, and utilize the SOC's Science Data Center (SOC) for data management and distribution. The SOC provides a single mission data archive for the housekeeping and science data, calibration data, ephemerides, attitude and other ancillary data needed to support the scientific use and interpretation. All levels of data products will reside at and be publicly disseminated from the SDC. Documentation and metadata describing data products, algorithms, instrument calibrations, validation, and data quality will be provided. Arguably, the most important innovation developed by the SOC is the MMS burst data management and selection system. With nested automation and 'Scientist-in-the-Loop' (SITL) processes, these systems are designed to maximize the value of the burst data by prioritizing the data segments selected for transmission to the ground. This paper describes the MMS science operations approach, processes and data systems, including the burst system and the SITL concept.
- Book Chapter
3
- 10.1007/978-3-030-50399-4_60
- Jun 10, 2020
In the daily scientific research activities, the university will form a large number of files. The goal of data management of university scientific research activities is to ensure the availability, authenticity and validity of scientific research data, but the electronic data is more easily modified. There are problems in data security, storage security and utilization security of scientific research data. The blockchain technology is applied to the data management of scientific research activities, which can realize the management of the whole life cycle of scientific research data and ensure the effective use and security of scientific research data. SWOT analysis is helpful for university research managers to understand the real situation of research management. Therefore, this paper makes a detailed SWOT analysis of university scientific research data management, and proposes a data management system of university scientific research activities based on blockchain technology.
- Research Article
2
- 10.1360/tb-2023-0463
- Feb 1, 2024
- Chinese Science Bulletin (Chinese Version)
<p indent="0mm">Massive amounts of data, dramatically growing computing power, and the development of the digital economy continue to give rise to data science. The rapid development in technologies, such as machine learning, artificial intelligence, and blockchain, has enhanced this trend. Data have become a new factor of production, bringing revolutionary changes to people’s lives and production techniques. In this context, scientific research has also begun to shift towards a new data-driven paradigm, using massive amounts of data as a basis to reveal the correlations hidden behind them, thereby adding new dimensions and perspectives to existing conventional research. Recently, data-based technologies, such as artificial intelligence, which are based on Big Data, have been gradually involved in scientific research. These technologies are used for integrating theory, computation, and experimental measurements, injecting new direction and impetus into scientific research. For example, algorithms such as random forests and neural networks have become commonly used data processing and analysis tools in materials science and achieved remarkable results in predicting properties, phase diagrams, and structures. Traditional scientific research methods often rely on specific theoretical assumptions and experimental designs, while data-driven scientific research focuses on obtaining knowledge and insights from data. By mining hidden correlations and patterns from massive amounts of data, researchers can conduct exploratory studies, discover new research directions and questions, expand the boundaries of research fields, and bring novel insights and breakthroughs to their research. Due to the importance of data, higher requirements for data management and utilization are demanded in scientific research. Currently, scientific research data is hindered by diverse formats, fragmented distribution, and inconsistent quality, causing data to remain inside the laboratory; this results in the wastage of scientific research resources and declination in the development of scientific research. To effectively manage data, improve data utilization, and address the issue of “data islands”, researchers have developed an Electronic Laboratory for Material Science, a data management platform based on the actual needs of frontline researchers and difficulties associated with the traditional flow of scientific research data. This platform covers the entire lifecycle of data production, collection, storage, analysis, and sharing. Furthermore, it is committed to realizing planned information processing, automated information collection, intelligent data empowerment, and promoting the commencement of a new paradigm in scientific research. This article introduces the development of scientific research paradigms and current advancements in data management at home and abroad, emphasizing the importance of data standards in scientific research and highlighting the successful exploration of the Electronic Laboratory for Material Science for significantly improving the efficiency in developing amorphous alloy materials, enhancing the quality of functional material single-crystal films, intelligently denoising angle-resolved photoemission spectroscopy spectra, and assisting the management of large-scale experimental stations. With this data management platform, researchers can now more effectively organize and manage experimental data, ensuring its integrity, systematicity, and standardization, thereby improving the efficiency and accuracy of their scientific research work and providing feasible paths and exemplary cases for data management in materials science. Based on the successful application of the Electronic Laboratory for Material Science, the research and development team will further improve its function and build an independent data platform for scientific research in China. In the era of Big Data, open sharing data and establishing an information-based ecosystem will establish the foundation for developing a smart laboratory and injecting new vitality into materials science advancements.
- Research Article
2
- 10.6083/m4tt4p9c
- Jan 1, 1993
- OHSU Digital Commons
Scientific data management has recently become a critical research issue due to computerization and automation of scientific research procedures and the resulting explosion of electronic data. For proper management of scientific data, adequate data types must be provided for representing the semantics of scientific data directly. Issues such as large data volume and data evolution are common among scientific applications, so traditional database support, such as storage management, are also beneficial to them. Current approaches such as standardized data formats and adoption of traditional databases do not accommodate the data support requirements of existing scientific applications well. Object-oriented databases, on the other hand, seem to hold promise, by combining flexible data models with traditional database support. In this dissertation, we constructed an experimental platform for exploring the design space for a scientific data management system. With this experimental approach, we could dynamically examine possible architectures for data support against actual instances of scientific data and typical operations executed on them. As many of the dynamic features of existing scientific applications, such as data access patterns, are yet to be discovered, the platform also provides an opportunity to explore such features. Based on the observed potential of object-oriented databases for scientific data management, the GemStone object-oriented database management system was chosen for the baseline data management architecture of the platform. Among scientific applications where better data support is desired, a scientific persistent language and data analysis environment called NewS and applications using it were selected as a target for the study. Connecting NewS with GemStone provided a cost-effective experimental platform where we could investigate a variety of scientific applications with single implementation. We incorporated GemStone into the NewS environment in such a manner that we could use existing NewS applications without modifications for experiments. We describe the design and implementation of our platform, GemStone-based NewS, and experiments performed on the platform. At the end, we describe primary contributions of the work, and assess whether or not our approach was a productive initial step toward improved scientific data management.
- Research Article
2
- 10.1016/0740-624x(93)90050-a
- Jan 1, 1993
- Government Information Quarterly
Data management and global change research: Technology and infrastructure
- Research Article
1
- 10.31891/2219-9365-2023-76-22
- Nov 30, 2023
- MEASURING AND COMPUTING DEVICES IN TECHNOLOGICAL PROCESSES
The paper introduces a sophisticated system designed for the meticulous selection and analysis of data in the realm of scientific research. This system empowers researchers to perform various data operations, including sorting, duplicate removal, data clustering, parameter-based filtration, elimination of empty records, and more. Data Selection System, Data Analysis in Scientific Research, Data Selection Methods for Scientific Research, Systems for Scientific Research, Data Processing and Interpretation. In this work, a robust system is presented to facilitate the systematic selection and analysis of data, catering specifically to the intricacies of scientific research. The system offers a comprehensive suite of operations, allowing researchers to perform essential tasks such as sorting data for improved organization, eliminating duplicate entries to enhance data integrity, clustering data to uncover patterns, filtering based on specific parameters to focus on relevant subsets, and clearing empty records for a refined dataset. Researchers can leverage advanced sorting functionalities to organize data based on specified parameters, enhancing data readability and facilitating a structured approach to analysis. The system incorporates mechanisms for identifying and removing duplicate entries, ensuring data accuracy and reliability in scientific investigations. Advanced clustering algorithms empower researchers to discern patterns within datasets, providing valuable insights crucial for scientific exploration. A key feature enables researchers to apply targeted filters based on specific parameters, refining datasets to focus on subsets of information relevant to their research objectives. Recognizing the importance of data completeness, the system adeptly manages and clears empty records, ensuring the integrity of analyses and facilitating more accurate research outcomes. A comprehensive tool designed for the purposeful extraction and refinement of data in scientific research workflows. The systematic examination and interpretation of data to derive meaningful insights within the context of scientific investigations. Varied approaches and techniques employed to selectively curate and prepare data for rigorous scientific analysis. Technological frameworks developed to enhance efficiency and effectiveness in scientific inquiry. The systematic manipulation and understanding of data to extract valuable information and insights. This framework represents a pivotal advancement in the realm of scientific data management, providing researchers with a versatile toolset to elevate the precision and efficiency of their data-driven investigations.
- Research Article
11
- 10.3389/ftox.2022.893924
- Jun 22, 2022
- Frontiers in Toxicology
Research in environmental health is becoming increasingly reliant upon data science and computational methods that can more efficiently extract information from complex datasets. Data science and computational methods can be leveraged to better identify relationships between exposures to stressors in the environment and human disease outcomes, representing critical information needed to protect and improve global public health. Still, there remains a critical gap surrounding the training of researchers on these in silico methods. We aimed to address this gap by developing the inTelligence And Machine lEarning (TAME) Toolkit, promoting trainee-driven data generation, management, and analysis methods to “TAME” data in environmental health studies. Training modules were developed to provide applications-driven examples of data organization and analysis methods that can be used to address environmental health questions. Target audiences for these modules include students, post-baccalaureate and post-doctorate trainees, and professionals that are interested in expanding their skillset to include recent advances in data analysis methods relevant to environmental health, toxicology, exposure science, epidemiology, and bioinformatics/cheminformatics. Modules were developed by study coauthors using annotated script and were organized into three chapters within a GitHub Bookdown site. The first chapter of modules focuses on introductory data science, which includes the following topics: setting up R/RStudio and coding in the R environment; data organization basics; finding and visualizing data trends; high-dimensional data visualizations; and Findability, Accessibility, Interoperability, and Reusability (FAIR) data management practices. The second chapter of modules incorporates chemical-biological analyses and predictive modeling, spanning the following methods: dose-response modeling; machine learning and predictive modeling; mixtures analyses; -omics analyses; toxicokinetic modeling; and read-across toxicity predictions. The last chapter of modules was organized to provide examples on environmental health database mining and integration, including chemical exposure, health outcome, and environmental justice indicators. Training modules and associated data are publicly available online (https://uncsrp.github.io/Data-Analysis-Training-Modules/). Together, this resource provides unique opportunities to obtain introductory-level training on current data analysis methods applicable to 21st century science and environmental health.
- Book Chapter
- 10.1017/9781316888773.006
- Jul 12, 2018
Chapter Objectives In this chapter, you will learn to: • identify the basic concepts of data management; • understand the role and importance of catalogs, metadata, data quality, and data governance; • identify key roles in database modeling and management; • understand the differences between an information architect, database designer, data owner, data steward, database administrator, and data scientist. Opening Scenario Sober realizes that the success of its entire business model depends on data, and it wants to make sure the data are managed in an optimal way. The company is looking at how to organize proper data management and wondering about the corresponding job profiles to hire. The challenge is twofold. On the one hand, Sober wants to have the right data management team to ensure optimal data quality. On the other hand, Sober only has a limited budget to build that team. In this chapter, we zoom into the organizational aspects of data management. We begin by elaborating on data management and review the essential role of catalogs and metadata, both of which were introduced in Chapter 1. We then discuss metadata modeling, which essentially follows a similar database design process to that which we described in Chapter 3. Data quality is also extensively covered in terms of both its importance and its underlying dimensions. Next, we introduce data governance as a corporate culture to safeguard data quality. We conclude by reviewing various roles in data modeling and management, such as information architect, database designer, data owner, data steward, database administrator, and data scientist. Data Management Data management entails the proper management of data and the corresponding data definitions or metadata. It aims at ensuring that (meta-)data are of good quality and thus a key resource for effective and efficient managerial decision-making. In the following subsections we first review catalogs, the role of metadata, and the modeling thereof. This is followed by a discussion on data quality and data governance. Catalogs and the Role of Metadata The importance of good metadata management cannot be understated. In the past this was often neglected, resulting in significant problems when applications needed to be updated or maintained. In the file-based approach to data management, the metadata were stored in each application separately, creating the issues discussed in Chapter 1.
- Research Article
3
- 10.3969/j.issn.1674-764x.2011.03.010
- Jan 31, 2012
- Journal of resources and ecology
Abstract: Northeast Asia is a key area for Earth system studies, global change frontier science research and regional sustainable development research. It has a complex ecological environment, a variety of climatic zones and typical human-Earth relationships. This paper outlines a data resources integration system fulfilling the data accumulation and management requirements of the Northeast Asia Resources and Environment Scientific Expedition. The data resources integration system has three subsystems: (i) data resources collection and management standards and specifications system, (ii) data classification system and (iii) a data management and publication software platform. The data resources collection and management standard and specification system has 23 specifications, divided into three types. They are: (i) data collection and processing specification type, (ii) data analysis and archiving specification and (iii) data management and sharing specification. The data resources classification system has four classes, 25 sub classes and 128 data elements. The data management and publication software platform has five function models: (i) data catalogue search model, (ii) metadata management model, (iii) data publication and virtualization model, (iv) data view model and (v) data download model. Based on the designed data integration system a prototype system has been developed and is supported by computer and Web GIS technologies. So far 144 datasets have been integrated into this data system. As more data are accumulated and integrated, it will play an important role in future scientific expedition data application and analysis.
- Research Article
17
- 10.1109/access.2021.3117780
- Jan 1, 2021
- IEEE Access
The current pandemic has significantly impacted educational practices, modifying many aspects of how and when we learn. In particular, remote learning and the use of digital platforms have greatly increased in importance. Online teaching and e-learning provide many benefits for information retention and schedule flexibility in our on-demand world while breaking down barriers caused by geographic location, physical facilities, transportation issues, or physical impediments. However, educators and researchers have noticed that students face a learning and performance decline as a result of this sudden shift to online teaching and e-learning from classrooms around the world. In this paper, we focus on reviewing eye-tracking techniques and systems, data collection and management methods, datasets, and multi-modal learning data analytics for promoting pervasive and proactive learning in educational environments. We then describe and discuss the crucial challenges and open issues of current learning environments and data learning methods. The review and discussion show the potential of transforming traditional ways of teaching and learning in the classroom, and the feasibility of adaptively driving learning processes using eye-tracking, data science, multimodal learning analytics, and artificial intelligence. These findings call for further attention and research on collaborative and intelligent learning systems, plug-and-play devices and software modules, data science, and learning analytics methods for promoting the evolution of face-to-face learning and e-learning environments and enhancing student collaboration, engagement, and success.
- Research Article
13
- 10.3844/jcssp.2014
- Dec 10, 2014
- Journal of Computer Science
“Data mining” for “knowledge discovery in databases” and associated computational operations first introduced in the mid-1990 s can no longer cope wit h the analytical issues relating to the so-called “ big data”. The recent buzzword big data refers to large volumes of diverse, dynami c, complex, longitudinal and/or distributed data generated from instruments, sensors, Internet transactions, email, video, clic k streams, noisy, structured/unstructured and/or all other digital sources available today and in the fu ture at speeds and on scales never seen before in human his tory. The big data also being described using 3 Vs, volume, variety and velocity (with an additional 4t h V for “veracity” and more recently with a 5th V f or “value”), requires a set of new technologies, such as high performance computing i.e., exascale, architectures (distributed or grid), algorithms (fo r data clustering and generating association rules) , programming languages, automated and scalable software tools, to uncover hidden patterns, unknown correlations and other useful information lately re ferred to as “actionable knowledge” or “data produc ts” from the massive volumes of complex raw data. In view of the above facts, the paper gives an introduct ion to the synergistic challenges in “data-intensive” s cience and “exascale” computing for resolving “big data analytics” and “data science” issues in four main d isciplines namely, computer science, computational science, statistics and mathematics. For the realis ation of vital identified foundational aspects of a n effective cyber infrastructure, basic problems need to be add ressed adequately in the respective disciplines and are outlined. Finally, the paper looks at five scientif ic research projects that are urgently in need of h igh performance computing; this is in contrast to the e arlier situations where private business enterprise s were the drivers of better modern and faster technologie s.
- Research Article
18
- 10.1007/s10586-018-1800-4
- Jan 29, 2018
- Cluster Computing
Data clustering partitions the information into helpful classes or groups with no earlier learning. This is a fundamental method in the field of computer data mining and it has turned into an essential element in many other engineering areas including cloud computing. This paper purports a novel clustering technique based on the application of krill herd Efficient Stud Krill Herd—Clustering (ESKH-C) technique. It is an optimisation approach for data clustering problem in which a swarm of krill (candidate solutions) moves to converge to specific positions as final cluster centres by minimizing the fitness function. The accuracy of the purposed methodology is blazed on different well familiar bench mark data sets. Analysed with the common clustering methods such as k-means clustering algorithm, data clustering using particle swarm optimization algorithm, ant colony optimization based data clustering, and clustering method using bacterial foraging algorithm, MATLAB simulation results evidence that the proposed technique is an effectual data clustering method. The proposed data clustering method can be employed to manipulate vast data sets with different cluster sizes, multi dimensional and densities.
- Preprint Article
- 10.5194/egusphere-egu2020-10057
- Mar 23, 2020
&lt;p&gt;The Istituto Nazionale di Geofisica e Vulcanologia (INGV) has a long tradition of sharing scientific data, well before the Open Science paradigm was conceived. In the last thirty years, a great deal of geophysical data generated by research projects and monitoring activities were published on the Internet, though encoded in multiple formats and made accessible using various technologies.&lt;/p&gt;&lt;p&gt;To organise such a complex scenario, a working group (PoliDat) for implementing an institutional data policy operated from 2015 to 2018. PoliDat published three documents: in 2016, the data policy principles; in 2017, the rules for scientific publications; in 2018, the rules for scientific data management. These documents are available online in Italian, and English (https://data.ingv.it/docs/).&lt;/p&gt;&lt;p&gt;According to a preliminary data survey performed between 2016 and 2017, nearly 300 different types of INGV-owned data were identified. In the survey, the compilers were asked to declare all the available scientific data differentiating by the level of intellectual contribution: level 0 identifies raw data generated by fully automated procedures, level 1 identifies data products generated by semi-automated procedures, level 2 is related to data resulting from scientific investigations, and level 3 is associated to integrated data resulting from complex analysis.&lt;/p&gt;&lt;p&gt;A Data Management Office (DMO) was established in November 2018 to put the data policy into practice. DMO first goal was to design and establish a Data Registry aimed to satisfy the extremely differentiated requirements of both internal and external users, either at scientific or managerial levels. The Data Registry is defined as a metadata catalogue, i.e., a container of data descriptions, not the data themselves. In addition, the DMO supports other activities dealing with scientific data, such as checking contracts, providing advice to the legal office in case of litigations, interacting with the INGV Data Transparency Office, and in more general terms, supporting the adoption of the Open Science principles.&lt;/p&gt;&lt;p&gt;An extensive set of metadata has been identified to accommodate multiple metadata standards. At first, a preliminary set of metadata describing each dataset is compiled by the authors using a web-based interface, then the metadata are validated by the DMO, and finally, a DataCite DOI is minted for each dataset, if not already present. The Data Registry is publicly accessible via a dedicated web portal (https://data.ingv.it). A pilot phase aimed to test the Data Registry was carried out in 2019 and involved a limited number of contributors. To this aim, a top-priority data subset was identified according to the relevance of the data within the mission of INGV and the completeness of already available information. The Directors of the Departments of Earthquakes, Volcanoes, and Environment supervised the selection of the data subset.&lt;/p&gt;&lt;p&gt;The pilot phase helped to test and to adjust decisions made and procedures adopted during the planning phase, and allowed us to fine-tune the tools for the data management. During the next year, the Data Registry will enter its production phase and will be open to contributions from all INGV employees.&lt;/p&gt;
- Research Article
6
- 10.3233/ds-190017
- Apr 24, 2019
- Data Science
The sharing of scientific and scholarly data has been increasingly promoted over the last decade, leading to open repositories in many different scientific domains. However, data sharing and open data are not final goals in themselves, the real benefit is in data reuse, which allows leveraging investments in research and enables large-scale data-driven research progress. Focusing on reuse, this paper discusses the design of an integrated framework to automatically take advantage of large amounts of scientific data extracted from the literature to support research, and in particular scientific model development. Scientific models reproduce and predict complex phenomena and their development is a rather challenging task, within which scientific experiments have a key role in their continuous validation. Starting from the combustion kinetics domain, this paper discusses a set of use cases and a first prototype for such a framework which leads to a set of new requirements and an architecture that can be generalized to other domains. The paper analyzes the needs, the challenges and the research directions for such a framework, in particular those related to data management, automatic scientific model validation, data aggregation and data analysis, to leverage large amounts of published scientific data for new knowledge extraction.