Critical analysis of Big Data challenges and analytical methods
Critical analysis of Big Data challenges and analytical methods
- Research Article
16
- 10.1002/poi3.258
- Jun 14, 2021
- Policy & Internet
The use of big data and data analytics are slowly emerging in public policy-making, and there are calls for systematic reviews and research agendas focusing on the impacts that big data and analytics have on policy processes. This paper examines the nascent field of big data and data analytics in public policy by reviewing the literature with bibliometric and qualitative analyses. The study encompassed scientific publications gathered from SCOPUS (N = 538). Nine bibliographically coupled clusters were identified, with the three largest clusters being big data's impact on the policy cycle, data-based decision-making, and productivity. Through the qualitative coding of the literature, our study highlights the core of the discussions and proposes a research agenda for further studies. 大数据和数据分析的使用已逐渐出现在公共决策中,对此需要展开系统性综述和研究议程,聚焦大数据和数据分析对政策过程产生的影响。本文通过文献计量分析和定性分析,分析了公共政策中大数据和数据分析这一新兴领域。本研究包括了从SCOPUS中选取的科学刊物(N=538)。识别了9个文献耦合簇,其中最大的三个簇分别为:大数据对政策周期的影响、基于数据的决策、生产率。通过对文献进行定性编码,我们的研究强调了探讨的核心,并提出了进一步的研究议程。 El uso de macrodatos y análisis de datos está emergiendo lentamente en la formulación de políticas públicas, y hay pedidos de revisiones sistemáticas y agendas de investigación que se centren en los impactos que los macrodatos y el análisis tienen en los procesos de políticas. Este artículo examina el campo naciente del big data y el análisis de datos en las políticas públicas mediante la revisión de la literatura con análisis bibliométricos y cualitativos. El estudio abarcó publicaciones científicas recopiladas de SCOPUS (N = 538). Se identificaron nueve grupos acoplados bibliográficamente, siendo los tres grupos más grandes el impacto de los macrodatos en el ciclo de políticas, la toma de decisiones basada en datos y la productividad. A través de la codificación cualitativa de la literatura, nuestro estudio destaca el núcleo de las discusiones y propone una agenda de investigación para estudios posteriores. Big data and data analytics have been seen as augmenting knowledge, ultimately leading to better decision-making. Arguments such as that the broad-based use of big data and data analytics will lead to the end-of-theory speak volumes about our expectations of big data and data analytics technologies' transformative power. While industry has been leading the way to test big data and analytics, public actors have been slower to engage (Poel et al., 2018), despite an equal opportunity for big data and data analytics to augment the public policy process. Utilizing big data and data analytics has become a near necessity due to our increasing capability for creating and collecting data at an extraordinary rate. The terms "big data" and "data analytics" have been among the buzzwords of recent years, leading to an upsurge in research, industry, and government applications (Zhou et al., 2014). The increased interest in big data in public policy can be seen in the scientific literature Figure 1, highlighting the increase in big data and analytics related literature. We also see public organizations increasingly engaging with big data analytics to solve challenges like the sustainability crisis and pandemics.1 Scholarly discourse has highlighted case studies and narratives on implementing big data and data analytics in the policy process. However, the literature lacks a systematic view of the current state of big data and data analytics in public policy, and there are identifiable research gaps (Desouza & Jacob, 2017). RQ1. What are the thematic communities of big data and data analytics literature concerning public policy-making? RQ2. What are the research questions emerging under each of the thematic research communities? Our study adopted a mixed-method systematic literature review approach based on a robust empirical bibliometric analysis followed by a qualitative analysis of the core documents to answer these questions. Using a well-established bibliometric method, bibliographic coupling, we identified thematic differences within the literature, and here, we highlight points of departure from the extant literature. The bibliometric analysis was, in turn, used as a basis for the qualitative analysis of the core literature, which is used to propose a research agenda. We find nine contemporary research communities addressing different aspects of big data and data analytics in public policy. While these communities have significant overlap, our analysis identifies them drawing from different theoretical foundations. Moreover, we demonstrate three larger research strands taking different vantage points, namely building strategic capability, data-based decision-making and productivity increases. Finally, our work proposes a research agenda focusing on the role of strategic capability, data-based decision-making, how to address expectations for better services while simultaneously increasing productivity and how to leverage policy analytics and empiricism. Our results offer scholars in public policy a vantage point to the theoretical foundations of research in big data and data analytics in public policy-making. We also draw from the identified communities to highlight emerging research themes that can guide research forward. For policymakers, our results highlight the on-going scholarly debate that focuses on addressing critical issues in the adoption of big data in public policy-making, namely capability building and the extent of data-based decision-making. This article will proceed as follows: next, we review the central elements of big data in policy-making. This is followed by a description of the data and our mixed-method approach. Finally, the empirical results are described and followed by a discussion to make sense of the research themes emerging from the analysis. "Big data" is a general term used for the process of gathering massive amounts of data from different sources. Sources can include human-input data but also includes data from sensors or different types of monitoring systems that create process data while running. It is clear that we are accumulating data at a never before seen rate. Already, in 2014, the pace was staggering, with 90% of the world's data being collected during the prior 2 years and 2.5 quintillion bytes of data added each day (Kim et al., 2014). Having access to massive amounts of data has enabled significant innovation in both the public and private domains. Looking at companies like Google and Amazon, with their innovation of new services for consumers, or at the recent ability for doctors to detect cancer cells more precisely thanks to massive training data about what a cancerous cell is, we can see that we are very much on the cusp of creating a broad utility of big data and analytics. This has been seen as a shift in the Industrial Revolution's magnitude (Richards & King, 2014) and has been widely hyped in business (Margetts & Sutcliffe, 2013). That said, public policy is not at the forefront of the use of big data and data analytics in decision making (Kaski et al., 2019; Poel et al., 2018). This nonadoption is due to multiple factors limiting these technologies' utility (Malomo & Sena, 2017). The ever-increasing amount of data offers possibilities for discovering new relationships and inferencing a multitude of problems. However, this comes with new challenges involving reproducibility, complexity, security, and risks to privacy and a need for new technology and human skills. This is very much the case in public policy, where we need to clearly identify where big data can add value in an ethical and trustworthy manner. In a review, Giest (2017) highlighted three underlying factors to consider. First, institutional capacities have a significant role in the use of big data in public policy, producing solutions that can enable users to easily interact with data while also taking into account the siloed data structures in the public domain. However, we know from previous research that siloed structures are an important limiting factor for public policy utilization of big data (Malomo & Sena, 2017). Second, hand-in-hand with big data comes the broader digitalization of public services. Digitalization allows for mediums to interact with big data but also enables the creation of new data. There is, however, evidence that digitalization changes the interactions between citizens and public officials and requires new skills from both parties. Third, big data information will have an impact on the policy cycle. Studies have found that there has been limited progress in taking advantage of big data and analytics (Poel et al., 2018) because it requires a significant change in the policy cycle (Höchtl et al., 2016). Giest (2017) highlights two issues, the substantive role and the procedural role of big data in policy instruments. Procedural activities focus on regulatory activities, such as enabling open data, while substantive actions relate to collecting data for enhancing, for example, evidence-based policy making. Capacities, digitalization, and the role of big data in the (substantive and procedural) policy cycle are core to digital-era governance and evidence-based policy making. In this, it is important to note that policy-makers are not a homogeneous group, and policy cycles vary. Thus, the objectives of analytics throughout the policy cycle vary significantly (Daniell et al., 2016) whether or not we approach the policy cycle as separate discrete stages (Jann & Wegrich, 2007), and it has been shown that big data analytics, when used more in some policy stages than in others, notably improved government transparency, policy evaluation, foresight, and agenda setting (Poel et al., 2018). This should be reflected against findings that data analytics have been politically significant in all policy cycle stages (Van der Voort et al., 2019). To overcome the challenges, Poel et al. (2015) highlighted multiple topics that must be addressed to enable capacity building, digitalization, and data integration into the policy cycle. These are (1) a skills gap, (2) reduced transparency due to data analytics, (3) sources and tools, (4) standardization of methods and tools, (5) linking of policy experiments with impact assessments, and (6) enabling policy-makers to be informed about the tools that are developed and piloted. The highlighted themes give context to the issue of big data in policy. While we see the significant impacts being created by the use of big data in policy making, along with the subsequent adaptation of data analytics, we need to better explain and make transparent the utility and complementarity of big data-driven analyses for the policy cycle (Vydra & Klievink, 2019). The challenge highlighted by Poel et al. (2015) and Giest et al. (2017), is also reflected in Pencheva et al. (2020) and Ingram (2019). Both note that big data in public policy has focused more on the "techno-rational factors," dismissing the importance of interaction with the policy process. We know that technology adoption is dependent on the perceived usefulness by the user (Venkatesh & Davis, 2000) and that there is scepticism towards the use of big data and data analytics is public policy-making (Guenduez et al., 2020). This can be the result of a mismatch in practise and expectations. Durrant et al. (2018) show how there is a aspirational motivation to the use of big data and data analytics, not reflected by its everyday utility. While we know that by employing data drive approaches, there is significant potential for anticipatory governance, interaction among stakeholders is the key to draw value from big data and analytics (Maffei et al., 2020; Starke & Lünich, 2020). While the increased stakeholder involvement does not protect from big data and data analytics policy-making creating hard to detect inequalities (Giest & Samuels, 2020; van Veenstra et al., in press). However, engaging with a large pool of stakeholders in the public policy-making process will increase the complexities of adopting big data and data analytics (Janssen et al., 2017). In addition to stakeholder interaction, the ability to build capacity and also evaluate deficiencies is important (Okuyucu & Yavuz, 2020). Building capacity should not be merely seen as the technical capacity, while that is also important (Poel et al., 2018), but a more holistic capability to integrate big data and analytics into the policy cycle (Höchtl et al., 2016). Policy-making organizations, while being exceptionally technically capable, can be in a situation where the benefits from big data and data analytics remain small due to "applications" not fitting "their organizations and main statutory tasks" (Klievink et al., 2017). This to say that when we talk about big data and data analytics capabilities in public policy-making, the literature focuses on technical issues but also on the ability of big data and data analytics to produce policy-relevant applications. While we see an increasing and diverse set of research addressing different challenges of big data and data analytics, the current body of literature lacks holistic research agendas (Desouza & Jacob, 2017) addressing the issues highlighted from practice by Giest (2017) and Poel et al. (2015). While we can note emerging fields such as policy analytics (De Marchi et al., 2016; Tsoukias et al., 2013), there is a need to better understand the theoretical grounding and research gap of big data and data analytics in public policy-making. This study's methodological approach was based on a mixed method of quantitative and qualitative analyses of the bibliometric data and the publication's content. The selected four-step mixed-methods approach, described below, enables a holistic approach to comprehending the current state-of-the-art and allows us to propose an agenda for going forward. The first phase focused on retrieving the sample of relevant articles and their bibliometric data for analysis. The second phase involved the bibliometric analysis of the retrieved data, which was performed by analysing descriptive statistics, bibliographical coupling, network analytics, and community detection. By gaining a comprehensive view of the more extensive body of literature, we could implement more filtering process based on eigenvector centrality to get a shortlist of papers for the next phase. The third phase, qualitative analysis, continued the process with an in-depth review and coding of the articles' full text. Finally, in the fourth phase, Synthesis, we draw insights from the MAXQDA coding analysis and reporting. This four-phase process is shown in Figure 2. The data used in this study were retrieved from the Scopus database. Scopus is Elsevier's abstract and citation database that has over 1.7 billion cited references dating back to 1970. A central aspect of the quality of the results is that the query used to search for relevant articles was correctly designed. The study focuses on public policy-making and big data and data analytics. The study's scope is relatively narrow, focusing solely on publications that address policy-making and the policy process. The decision of the scope excludes articled that focus on, for example, big data or data analytics, but lack the specific aspect of policy-making. This was the key inclusion criteria of articles into the data set. To focus on this specific scope, we used an iterative approach where multiple search strings were tested, and after each search, the abstracts of the 10 most cited articles and the 10 most recent articles were reviewed to understand if the query results reflected the objectives of the study. In practice, the process started with a seed query of "big data" or "data analytics" and "public policy." The query results were reviewed to estimate which articles focused on big data or data analytics and policy-making. These articles were reviewed to see if new terms emerged through the titles, abstracts, and keywords that needed to be included in the analysis. The process adjusted based on a subjective evaluation of the number of false-positives in the 10 most recent and 10 most cited publications and the number of articles retrieved. This method of short-listing the important literature is known as the snowball method, and the process includes consulting the bibliographies of the key documents to find other relevant titles in the subject (Jalali & Wohlin, 2012). After multiple tests of a comprehensive query that also limited the number of false-positives, we downloaded the metadata for 538 documents. These documents were retrieved using the query "public policy," "policy analysis," "policy making," or "public administration," with the terms "big data," "data analytics," or "automated decision-making" in the title, abstract, or keywords of the document. To analyze the literature, we used the well-established bibliometric method of bibliographical coupling. Bibliographical coupling allows for analysis of the publications' shared intellectual background (Kessler, 1963), highlighting contemporary research (Youtie et al., 2013). It is an approach to analyzing the shared theoretical background of scientific publications where the link between documents is calculated by the number of references the two documents share. Kessler (1963) elaborates, "A single item of reference shared by two documents is defined as a unit of coupling between them," and if multiple references are shared, the weight of the coupling increases. Bibliographical coupling is able to highlight hot topics (Glanzel & Czerwon, 1996) and links documents with a similar research focus (Jarneving, 2007), ultimately creating a "contemporaneous representation of knowledge" (Youtie et al., 2013). This approach has been used in several research papers to form the basis for research agenda building (Suominen et al., 2019; Yuan et al., 2015). Using the retrieved publication metadata, the VOSviewer tool (van Eck & Waltman, 2009) was selected to calculate bibliographical coupling weights for all the documents in our data set. VOSviewer is a free tool used for bibliometrics and was selected due to the exports available in the software allowing for deeper network analysis in Gephi. The SCOPUS data export was used as an input to the VOSviewer. During the analysis process, we selected documents as the level of analysis, minimum number of citations for a document was set to zero and the full set was selected for the analysis. The full counting method, which assigns each researcher with full credit of one publication rather than a fractional share per the number of authors, was used for the calculation method. Finally, we accepted VOSviewer default to keep the most extensive set of related items, which limited the analysis to the largest, by created by the bibliographical coupling analysis. This limited the analysis to documents. Bibliographical coupling analysis of a data set a by a set of and a set of we calculated the link weight between each publication we created a This data, with the VOSviewer was to because it allows for more network and community detection. network analysis, network descriptive for example, for were calculated in Gephi. were identified using et al., The is one of the most methods to find of in The methods by each to a separate community there after the in by in a This process is continued through the ultimately creating a new network with to communities et al., for a In the the can be by a that the number of communities the This was to the number of We increased the value the community has one share of the documents. We also calculated the eigenvector centrality for each document in the data set. In eigenvector centrality the a has in A value is calculated to all based on an that to other important are more important than equal This centrality value a publication's to the network created by the bibliographical coupling analysis. for we identified the most central publications from each community to be selected for systematic the most central publications was to keep the sample of documents in the coding phase and the most eigenvector central publications was a method to the most important publications from each community for a deeper analysis. of study of study or empirical study findings The offer the as a to be to the study's at For the current we adopted the of elements in the creating a approach (1) the context of the (2) and (3) the scope defined as a research We the or empirical study and to be the (4) for the theoretical or methodological We also included a separate for (5) the method used to understand the specific approach used in the study. For the key we also created The first focused on (6) results and the second focused on discussion and This was to highlight other study results and to the analysis in the scholarly debate the research To add to the coding process, the qualitative analysis was using MAXQDA The software allows for coding the documents with the and drawing from the created The use of the software transparency and the of the analysis et al., & 2012). After the coding was the coding process was by one the eigenvector central documents from each and the documents based on the The second researcher the role of the The interaction between the was to that all of the in the coding were The second researcher through the papers and by the first researcher to make that each if was However, as publications can the our approach was to we not the information multiple per This of an approach for example, The MAXQDA analysis software used in the coding created of the documents and created a document about the of text. These were used in the phase. The of by MAXQDA was used to the The two researcher with the MAXQDA of the communities After review of the the on their There after the to draw insights from the coding of the core of The retrieved publications are with the first publication in the data set in as seen in Figure The data were retrieved in 2020; publications for the first and one should a in publication While Figure it is important to this into Figure the of the topics of big data and data analytics in the public policy related literature concerning the public policy literature for the The publication volumes are on 10 to be able to them in one It is clear that the body of literature focused on big data and data analytics in public policy is a much increase in interest in to the public policy body of literature. of the publications are from or seen in 1, the three have over highlights all with over publications in the data set. the are highlighting the use of big data and data analytics for example, the topics of and issues and The descriptive analyses of the data also highlight the different that are on the shown in the publication sources for the articles included in the data set are and information we should note our The study focused solely on big data or data analytics and its use in public policy-making. In the the number of articles focusing on single for example, big data is much We should also note that the search in SCOPUS at the title, abstract and making articles big data or data analytics and policy-making in by the It is that the publication sources with at publications all or of all This that publications are over different publication and an on-going debate on the subject is hard to to a specific The of the publications by with that of scientific publication with is significantly the with The two largest are followed by the and and with Looking at the of the between different seen in of has a significant of publications in the data and with the publication there were publications per To understand more in-depth the we the and keywords of the publications in the We the and fields by and in the Figure the terms are as focusing on big data and public policy. thematic terms emerging to the most terms relate to and into the articles' the descriptive analyses highlight that the data set recent among multiple sources and with relatively volumes by The terms in the publications with the search The 538 publications' bibliometric data were for bibliographic coupling using VOSviewer During the calculation process, the software first if of the publications were and from the network emerging from the data. VOSviewer the largest in the full and offers an to used the largest set. to use the largest set allows focusing on the core the In our data the largest set of creating a network was documents. The documents were from the as these articles not be included in the bibliometric coupling based clusters but remain throughout the analysis. with the sample of the VOSviewer created network was to software for further analysis. were and the bibliographic coupling network a with the of the network being This to say that there is significant between in the communities as is but the full network is not as the is the communities were using the et al., The was increased from its default value of the community an share of the documents. a of the in nine with the the community has of the The largest community of the A representation of the created can be seen in Figure In this the the clusters created using the The of a the number of citations of an article in the the the network was created by two large communities
- Research Article
- 10.21303/2585-6847.2021.002228
- Dec 22, 2021
- Technology transfer: fundamental principles and innovative technical solutions
Big Data has been created from virtually everything around us at all times. Every digital media interaction generates data, from computer browsing and online retail to iTunes shopping and Facebook likes. This data is captured from multiple sources, with terrifying speed, volume and variety. But in order to extract substantial value from them, one must possess the optimal processing power, the appropriate analysis tools and, of course, the corresponding skills. The range of data collected by businesses today is almost unreal. According to IBM, more than 2.5 times four million data bytes generated per year, while the amount of data generated increases at such an astonishing rate that 90 % of it has been generated in just the last two years. Big Data have recently attracted substantial interest from both academics and practitioners. Big Data Analytics (BDA) is increasingly becoming a trending practice that many organizations are adopting with the purpose of constructing valuable information from BD. The analytics process, including the deployment and use of BDA tools, is seen by organizations as a tool to improve operational efficiency though it has strategic potential, drive new revenue streams and gain competitive advantages over business rivals. However, there are different types of analytic applications to consider. This paper presents a view of the BD challenges and methods to help to understand the significance of using the Big Data Technologies. This article based on a bibliographic review, on texts published in scientific journals, on relevant research dealing with the big data that have exploded in recent years, as they are increasingly linked to technology
- Single Book
33
- 10.1201/9781315373966
- Oct 3, 2016
Proven Methods for Big Data Analysis As big data has become standard in many application areas, challenges have arisen related to methodology and software development, including how to discover meaningful patterns in the vast amounts of data. Addressing these problems, Applied Biclustering Methods for Big and High-Dimensional Data Using R shows how to apply biclustering methods to find local patterns in a big data matrix. The book presents an overview of data analysis using biclustering methods from a practical point of view. Real case studies in drug discovery, genetics, marketing research, biology, toxicity, and sports illustrate the use of several biclustering methods. References to technical details of the methods are provided for readers who wish to investigate the full theoretical background. All the methods are accompanied with R examples that show how to conduct the analyses. The examples, software, and other materials are available on a supplementary website.
- Research Article
167
- 10.3390/bdcc4020004
- Mar 26, 2020
- Big Data and Cognitive Computing
Big data is the concept of enormous amounts of data being generated daily in different fields due to the increased use of technology and internet sources. Despite the various advancements and the hopes of better understanding, big data management and analysis remain a challenge, calling for more rigorous and detailed research, as well as the identifications of methods and ways in which big data could be tackled and put to good use. The existing research lacks in discussing and evaluating the pertinent tools and technologies to analyze big data in an efficient manner which calls for a comprehensive and holistic analysis of the published articles to summarize the concept of big data and see field-specific applications. To address this gap and keep a recent focus, research articles published in last decade, belonging to top-tier and high-impact journals, were retrieved using the search engines of Google Scholar, Scopus, and Web of Science that were narrowed down to a set of 139 relevant research articles. Different analyses were conducted on the retrieved papers including bibliometric analysis, keywords analysis, big data search trends, and authors’ names, countries, and affiliated institutes contributing the most to the field of big data. The comparative analyses show that, conceptually, big data lies at the intersection of the storage, statistics, technology, and research fields and emerged as an amalgam of these four fields with interlinked aspects such as data hosting and computing, data management, data refining, data patterns, and machine learning. The results further show that major characteristics of big data can be summarized using the seven Vs, which include variety, volume, variability, value, visualization, veracity, and velocity. Furthermore, the existing methods for big data analysis, their shortcomings, and the possible directions were also explored that could be taken for harnessing technology to ensure data analysis tools could be upgraded to be fast and efficient. The major challenges in handling big data include efficient storage, retrieval, analysis, and visualization of the large heterogeneous data, which can be tackled through authentication such as Kerberos and encrypted files, logging of attacks, secure communication through Secure Sockets Layer (SSL) and Transport Layer Security (TLS), data imputation, building learning models, dividing computations into sub-tasks, checkpoint applications for recursive tasks, and using Solid State Drives (SDD) and Phase Change Material (PCM) for storage. In terms of frameworks for big data management, two frameworks exist including Hadoop and Apache Spark, which must be used simultaneously to capture the holistic essence of the data and make the analyses meaningful, swift, and speedy. Further field-specific applications of big data in two promising and integrated fields, i.e., smart real estate and disaster management, were investigated, and a framework for field-specific applications, as well as a merger of the two areas through big data, was highlighted. The proposed frameworks show that big data can tackle the ever-present issues of customer regrets related to poor quality of information or lack of information in smart real estate to increase the customer satisfaction using an intermediate organization that can process and keep a check on the data being provided to the customers by the sellers and real estate managers. Similarly, for disaster and its risk management, data from social media, drones, multimedia, and search engines can be used to tackle natural disasters such as floods, bushfires, and earthquakes, as well as plan emergency responses. In addition, a merger framework for smart real estate and disaster risk management show that big data generated from the smart real estate in the form of occupant data, facilities management, and building integration and maintenance can be shared with the disaster risk management and emergency response teams to help prevent, prepare, respond to, or recover from the disasters.
- Conference Article
- 10.2991/icismme-15.2015.375
- Jan 1, 2015
Complex network system has characteristics of numerous units, complex structure, and dynamic changes, which brings many problems for the study of network reliability. However, the big data of network system reliability and big data analysis technology provide more data and thought for supporting traditional reliability research. This paper explores the source, collecting process, preprocessing of the big data of network system reliability and research idea of big data analysis, meanwhile introduces research field to which big data of network system reliability relates at present. Introduction With the innovation of science and technology and the development of society, the scale and structure of the modern network system and the platform is becoming larger, whose units and connection mode is more complex, then the research problem of reliability is becoming more prominent. Network system is complex systems with particular structure and function, composed of many nodes (equipment, personnel, etc.), which is connected by edges (material, information, energy). Modern network system, such as communication network, traffic network, electric power network, has certain forming principle and task requires. It has characteristics of numerous units, complex structure, dynamic changes, etc. Traditional research of network system reliability usually have two kinds of thought: The one kind is that describing, simplifying and calculating network with series-parallel way. The other kind is that doing multiple simulations for network system if people have suitable condition. In recent years, information technology and computer network technology develops quickly and technology and services such as cloud computing, Internet, mobile phone constantly improve. Meanwhile, data scale of many kinds of systems has explosive growth. Big Data has become a hot issue in today's social studies[1]. Huge amounts of data acquisition in network system reliability and utilization become more convenient with the help of big data technology and analysis method. The big data hidden behind nodes, edges and the connection relations in the network system has a broad application space. Laws and properties revealed in big data play an important role in aspects such as exploring network system reliability problems in key stage, data correlation in models, prediction of network system reliability, which provide more data and thought for supporting traditional reliability research. This paper discusses the related contents and process of the study of network reliability based on the big data. Specific structure is as follows: the second part expounds the basic content of network reliability and some big data examples of network system reliability. In the third part the paper analyzes collection of the network reliability data. In the fourth part, starting from the concept of big data analysis, this paper introduces the general methods of big data analysis, and put forward some corresponding research idea combing with the network reliability data. The fifth part lists the field of study that the big data of network reliability is involved and applied preliminarily. The paper ends with a summary in the sixth part. International Conference on Information Sciences, Machinery, Materials and Energy (ICISMME 2015) © 2015. The authors Published by Atlantis Press 1828 Description of Network Reliability and Its Reliability Big Data Network Reliability Definition. In the process of the development of the network system reliability research, according to different network characteristics and research target, the standard of the definition of network reliability is not uniform. For example, in the communication network, according to the connotation of basic reliability and mission reliability, the network reliability is defined as the ability that complex network keeps entire network connectivity or realizes communication function under prescribed conditions and within the stipulated time. For other network systems, refers to the ability that transmits the material, information and energy completely and correctly in the network for certain time range of users’ expectations under prescribed conditions and within the stipulated time. Network Reliability Measure Parameters. The research of reliability problems is based on the analysis and evaluation of reliability measure parameters. At present, compared with the relatively perfect reliability measure parameter system, network reliability measure parameters have not yet formed a consistent view. But it always stresses that an important standard of reliability evaluation is to meet the needs of users in most of the definition. Therefore, by analyzing the inherent requirement of reliability and some existing research results, this paper briefly summaries some relevant parameters of the network system reliability as the following three aspects. Network System Reliability Measure Parameters Basic reliability Network Performance Random probability indicator
- Research Article
970
- 10.1109/access.2017.2689040
- Jan 1, 2017
- IEEE Access
Voluminous amounts of data have been produced, since the past decade as the miniaturization of Internet of things (IoT) devices increases. However, such data are not useful without analytic power. Numerous big data, IoT, and analytics solutions have enabled people to obtain valuable insight into large data generated by IoT devices. However, these solutions are still in their infancy, and the domain lacks a comprehensive survey. This paper investigates the state-of-the-art research efforts directed toward big IoT data analytics. The relationship between big data analytics and IoT is explained. Moreover, this paper adds value by proposing a new architecture for big IoT data analytics. Furthermore, big IoT data analytic types, methods, and technologies for big data mining are discussed. Numerous notable use cases are also presented. Several opportunities brought by data analytics in IoT paradigm are then discussed. Finally, open research challenges, such as privacy, big data mining, visualization, and integration, are presented as future research directions.
- Research Article
2
- 10.18768/ijaedu.819010
- Dec 31, 2020
- IJAEDU- International E-Journal of Advances in Education
Ability to solve realistic problems in chemical big data, the CDIO model inherits and develops the concept of the significant reform of engineering education in Europe and the United States for more than 20 years, and nearly 40 universities in China have become the demonstration university of China's CDIO engineering education model. China CDIO Engineering Education Model has officially accredited Shenyang University of Chemical Technology's Computer Science Technology College in 2020. This article is a review of Shenyang University of Chemical Technology's Computer Science Technology College in the process of China's CDIO Engineering Education Model Certification. Universal experience with CDIO engineering education has shown that CDIO standards not only contribute to enhancing the quality of education but also provide an opportunity to improve the quality of engineering education. A good foundation is laid for the systematic development of education. To this end, we introduced the CDIO engineering education model into the Shenyang University of Chemical Technology's big data analytics and applications, and in the design of pragmatic course teaching. We apply the CDIO concept to the human workforce training program, course system construction for chemistry majors, and faculty of Computer Science Technology majors. A series of investigations and practices have been carried out. To address the problem of offering big data analytics and applications courses for chemistry students, we introduce the CDIO engineering education model into the content, teaching methods, and teacher preparation are all aspects of this project. We have given a "problem-oriented" for both theoretical teaching and practical teaching, considering the learning situation of chemistry students. Based learning with an inspiring interactive theoretical teaching method, "result-oriented" practical teaching mode, and project design-oriented, students learn professional knowledge and engineering thinking through practical engineering projects and improve their engineering awareness and practical skills. Finally, through the combination of theory and practice, chemistry majors can understand the theories and methods of big data analysis and applications, as well as the principles and methods of big data analysis and processing. Keywords: CDIO, engineering education, big data analytics and applications, chemistry majors, talent training mode
- Research Article
14
- 10.1186/s40537-022-00560-z
- Jan 1, 2022
- Journal of Big Data
The study aimed to present an integrated model for evaluation of big data (BD) challenges and analytical methods in recommender systems (RSs). The proposed model used fuzzy multi-criteria decision making (MCDM) which is a human judgment-based method for weighting of RSs’ properties. Human judgment is associated with uncertainty and gray information. We used fuzzy techniques to integrate, summarize, and calculate quality value judgment distances. Then, two fuzzy inference systems (FIS) are implemented for scoring BD challenges and data analytical methods in different RSs. In experimental testing of the proposed model, A correlation coefficient (CC) analysis is conducted to test the relationship between a BD challenge evaluation for a collaborative filtering-based RS and the results of fuzzy inference systems. The result shows the ability of the proposed model to evaluate the BD properties in RSs. Future studies may improve FIS by providing rules for evaluating BD tools.
- Conference Article
- 10.2991/meita-15.2015.70
- Jan 1, 2015
Data type and amount in human society is growing in amazing speed which is caused by emerging new services such as cloud computing, internet of things (IoT) and social network, the era of big data has come. Data is a fundamental resource from simple dealing object. In order to fully understand the connotation of big data, this paper expounds the conception of big data, combines with the application requirements of manufacturing big data, the five layer stack type big data processing frameworkis proposed, and the key technologies of manufacturing big data applications are discussed and analyzed.
- Research Article
2
- 10.24025/2306-4420.63.2021.248464
- Dec 21, 2021
- Proceedings of Scientific Works of Cherkasy State Technological University Series Economic Sciences
The article is devoted to the study of the impact of Big Data technology on the effectiveness of digital marketing strategies It is emphasized that at present it is almost impossible to implement a marketing strategy that is not based on research, search and analysis of information obtained. Big Data technology is actively used to understand the target audience, consumer behavior, market conditions, etc. It is noted that working with «big data» is the key to creating a quality product or service and successfully promoting them in the market. The current state and features of the development of this technology in recent years are analyzed. The article describes the main characteristics and forms of big data, considers the types of data that are of particular importance to professionals in the field of digital marketing. There are a number of tasks that Big Data technology helps to solve for marketers: audience segmentation, mood analysis, targeted marketing, forecasting and policy analysis, measuring the results of campaigns, etc. The main methods of Big Data analysis, as well as tools and technologies for their processing are presented. A number of problems and consequences related to the use of «big data» in digital marketing are described: unauthorized use of personal information, scaling problem, integration of previously collected data, fragmentation of information collection and processing systems, lack of highly qualified specialists, problem of universal language of communication and work with data within the company, the threat of hacking, errors and malfunctions, the potential for damage to the company's reputation, etc. It has been established that data-driven digital marketing allows to effectively segment consumers, which further facilitates the ability to make optimal personal offers to customers, and can significantly reduce sales funnels and increase profitability of companies. In the course of the research the necessity of using «big data» in the implementation of digital marketing strategies by modern companies was identified and substantiated, the most susceptible areas for the use of Big Data were identified.
- Research Article
1080
- 10.1016/j.compag.2017.09.037
- Oct 10, 2017
- Computers and Electronics in Agriculture
A review on the practice of big data analysis in agriculture
- Research Article
111
- 10.5204/mcj.561
- Oct 11, 2012
- M/C Journal
Lists and Social MediaLists have long been an ordering mechanism for computer-mediated social interaction. While far from being the first such mechanism, blogrolls offered an opportunity for bloggers to provide a list of their peers; the present generation of social media environments similarly provide lists of friends and followers. Where blogrolls and other earlier lists may have been user-generated, the social media lists of today are more likely to have been produced by the platforms themselves, and are of intrinsic value to the platform providers at least as much as to the users themselves; both Facebook and Twitter have highlighted the importance of their respective “social graphs” (their databases of user connections) as fundamental elements of their fledgling business models. This represents what Mejias describes as “nodocentrism,” which “renders all human interaction in terms of network dynamics (not just any network, but a digital network with a profit-driven infrastructure).”The communicative content of social media spaces is also frequently rendered in the form of lists. Famously, blogs are defined in the first place by their reverse-chronological listing of posts (Walker Rettberg), but the same is true for current social media platforms: Twitter, Facebook, and other social media platforms are inherently centred around an infinite, constantly updated and extended list of posts made by individual users and their connections.The concept of the list implies a certain degree of order, and the orderliness of content lists as provided through the latest generation of centralised social media platforms has also led to the development of more comprehensive and powerful, commercial as well as scholarly, research approaches to the study of social media. Using the example of Twitter, this article discusses the challenges of such “big data” research as it draws on the content lists provided by proprietary social media platforms.Twitter Archives for ResearchTwitter is a particularly useful source of social media data: using the Twitter API (the Application Programming Interface, which provides structured access to communication data in standardised formats) it is possible, with a little effort and sufficient technical resources, for researchers to gather very large archives of public tweets concerned with a particular topic, theme or event. Essentially, the API delivers very long lists of hundreds, thousands, or millions of tweets, and metadata about those tweets; such data can then be sliced, diced and visualised in a wide range of ways, in order to understand the dynamics of social media communication. Such research is frequently oriented around pre-existing research questions, but is typically conducted at unprecedented scale. The projects of media and communication researchers such as Papacharissi and de Fatima Oliveira, Wood and Baughman, or Lotan, et al.—to name just a handful of recent examples—rely fundamentally on Twitter datasets which now routinely comprise millions of tweets and associated metadata, collected according to a wide range of criteria. What is common to all such cases, however, is the need to make new methodological choices in the processing and analysis of such large datasets on mediated social interaction.Our own work is broadly concerned with understanding the role of social media in the contemporary media ecology, with a focus on the formation and dynamics of interest- and issues-based publics. We have mined and analysed large archives of Twitter data to understand contemporary crisis communication (Bruns et al), the role of social media in elections (Burgess and Bruns), and the nature of contemporary audience engagement with television entertainment and news media (Harrington, Highfield, and Bruns). Using a custom installation of the open source Twitter archiving tool yourTwapperkeeper, we capture and archive all the available tweets (and their associated metadata) containing a specified keyword (like “Olympics” or “dubstep”), name (Gillard, Bieber, Obama) or hashtag (#ausvotes, #royalwedding, #qldfloods). In their simplest form, such Twitter archives are commonly stored as delimited (e.g. comma- or tab-separated) text files, with each of the following values in a separate column: text: contents of the tweet itself, in 140 characters or less to_user_id: numerical ID of the tweet recipient (for @replies) from_user: screen name of the tweet sender id: numerical ID of the tweet itself from_user_id: numerical ID of the tweet sender iso_language_code: code (e.g. en, de, fr, ...) of the sender’s default language source: client software used to tweet (e.g. Web, Tweetdeck, ...) profile_image_url: URL of the tweet sender’s profile picture geo_type: format of the sender’s geographical coordinates geo_coordinates_0: first element of the geographical coordinates geo_coordinates_1: second element of the geographical coordinates created_at: tweet timestamp in human-readable format time: tweet timestamp as a numerical Unix timestampIn order to process the data, we typically run a number of our own scripts (written in the programming language Gawk) which manipulate or filter the records in various ways, and apply a series of temporal, qualitative and categorical metrics to the data, enabling us to discern patterns of activity over time, as well as to identify topics and themes, key actors, and the relations among them; in some circumstances we may also undertake further processes of filtering and close textual analysis of the content of the tweets. Network analysis (of the relationships among actors in a discussion; or among key themes) is undertaken using the open source application Gephi. While a detailed methodological discussion is beyond the scope of this article, further details and examples of our methods and tools for data analysis and visualisation, including copies of our Gawk scripts, are available on our comprehensive project website, Mapping Online Publics.In this article, we reflect on the technical, epistemological and political challenges of such uses of large-scale Twitter archives within media and communication studies research, positioning this work in the context of the phenomenon that Lev Manovich has called “big social data.” In doing so, we recognise that our empirical work on Twitter is concerned with a complex research site that is itself shaped by a complex range of human and non-human actors, within a dynamic, indeed volatile media ecology (Fuller), and using data collection and analysis methods that are in themselves deeply embedded in this ecology. “Big Social Data”As Manovich’s term implies, the Big Data paradigm has recently arrived in media, communication and cultural studies—significantly later than it did in the hard sciences, in more traditionally computational branches of social science, and perhaps even in the first wave of digital humanities research (which largely applied computational methods to pre-existing, historical “big data” corpora)—and this shift has been provoked in large part by the dramatic quantitative growth and apparently increased cultural importance of social media—hence, “big social data.” As Manovich puts it: For the first time, we can follow [the] imaginations, opinions, ideas, and feelings of hundreds of millions of people. We can see the images and the videos they create and comment on, monitor the conversations they are engaged in, read their blog posts and tweets, navigate their maps, listen to their track lists, and follow their trajectories in physical space. (Manovich 461) This moment has arrived in media, communication and cultural studies because of the increased scale of social media participation and the textual traces that this participation leaves behind—allowing researchers, equipped with digital tools and methods, to “study social and cultural processes and dynamics in new ways” (Manovich 461). However, and crucially for our purposes in this article, many of these scholarly possibilities would remain latent if it were not for the widespread availability of Open APIs for social software (including social media) platforms. APIs are technical specifications of how one software application should access another, thereby allowing the embedding or cross-publishing of social content across Websites (so that your tweets can appear in your Facebook timeline, for example), or allowing third-party developers to build additional applications on social media platforms (like the Twitter user ranking service Klout), while also allowing platform owners to impose de facto regulation on such third-party uses via the same code. While platform providers do not necessarily have scholarship in mind, the data access affordances of APIs are also available for research purposes. As Manovich notes, until very recently almost all truly “big data” approaches to social media research had been undertaken by computer scientists (464). But as part of a broader “computational turn” in the digital humanities (Berry), and because of the increased availability to non-specialists of data access and analysis tools, media, communication and cultural studies scholars are beginning to catch up. Many of the new, large-scale research projects examining the societal uses and impacts of social media—including our own—which have been initiated by various media, communication, and cultural studies research leaders around the world have begun their work by taking stock of, and often substantially extending through new development, the range of available tools and methods for data analysis. The research infrastructure developed by such projects, therefore, now reflects their own disciplinary backgrounds at least as much as it does the fundamental principles of computer science. In turn, such new and often experimental tools and methods necessarily also provoke new epistemological and methodological challenges. The Twitter API and Twitter ArchivesThe Open
- Research Article
61
- 10.1016/j.outlook.2016.11.021
- Dec 8, 2016
- Nursing Outlook
Big data science: A literature review of nursing research exemplars
- Book Chapter
- 10.1007/978-3-030-58669-0_35
- Sep 20, 2020
Nowadays, with the development and maturity of big data technology, big data analysis technology is more and more widely used in practice, and more and more data are gradually applied to the smart grid of our country. Since the data in the smart grid meets the 4 V characteristics of big data (large quantity, fast speed, many types, low value density), the use of big data technology can provide more accurate and cheaper data for power information generation Economic value and significance. Through the analysis of the development process of China’s power industry, the development of distribution network in China obviously lags behind the development of power generation and transmission network. At present, more than 95% of the blackouts are caused by the distribution network, and half of the power loss occurs in the distribution network, so the automation of the distribution network system urgently needs the support of new technologies. This paper first enumerates several key points of big data technology, including big data collection, storage and analysis, and then expounds several methods of big data analysis. On this basis, big data technology is applied to the field of intelligent distribution network. Especially in the application of distribution forecasting, it can provide more powerful technical support for the operation of smart distribution network, continuously improve the technical level of China’s smart distribution network, and promote the optimization and upgrading of smart grid system. Finally, an optimized prediction model is proposed, and the application of the new technology (5G technology) developed at the present stage is prospected, and its contribution to the data acquisition and application of big data technology is analyzed.
- Research Article
9
- 10.5391/jkiis.2015.25.5.470
- Oct 25, 2015
- Journal of Korean Institute of Intelligent Systems
Big data has been used in diverse areas. For example, in computer science and sociology, there is a differ-ence in their issues to approach big data, but they have same usage to analyze big data and imply the anal-ysis result. So the meaningful analysis and implication of big d ata are needed in most areas. Statistics and machine learning provide various methods for big data analysis. In this paper, we study a process for big data analysis, and propose an efficient methodology of entire p rocess from collecting big data to implying the result of big data analysis. In addition, patent documents have the characteristics of big data, we pro-pose an approach to apply big da ta analysis to patent data, and imply the result of patent big data to build R&D strategy. To illustrate how to use our proposed methodology for real problem, we perform a case study using applied and registered patent documents retrieved f rom the patent databases in the world. Key Words : Big Data Analysis, Statistics, Natural Language Processing, Text Mining, Patent Analysis, Linear Model.Received: Aug. 28, 2015Revised : Sep. 17, 2015Accepted: Sep. 19, 2015