The digital transformation of tax audits: how AI, big data, blockchain, and advanced analytics are reshaping tax evasion detection
ABSTRACT Purpose This study examines how emerging digital technologies—particularly artificial intelligence (AI), machine learning, big data analytics, and blockchain—are transforming tax audit practices. The research aims to synthesize existing evidence on the role of these technologies in improving audit efficiency, detection accuracy, and enforcement effectiveness across different jurisdictions. Design/Methodology/Approach A systematic literature review was conducted using a PRISMA-guided methodology. Peer-reviewed publications published between 2017 and 2024 were collected and screened, resulting in a final sample of 23 studies. The selected studies were analyzed and synthesized to identify dominant technological applications, institutional contexts, and emerging research patterns in technology-enabled tax auditing. Findings The review identifies four major thematic clusters: AI-enabled automation in audit processes, blockchain-based audit integrity and traceability, data-analytics-driven risk detection, and regulatory and institutional readiness. The findings indicate that technologically advanced jurisdictions achieve significant improvements in audit precision, real-time anomaly detection, and transactional transparency. In contrast, developing economies face persistent challenges related to digital infrastructure limitations, technical capacity constraints, and fragmented regulatory frameworks. Originality/Value This study contributes to the tax administration literature by providing a comprehensive taxonomy of emerging tax audit technologies and offering a cross-jurisdictional synthesis of their effectiveness. It also proposes a conceptual framework integrating technological drivers, institutional moderators, and audit performance outcomes, thereby offering a structured perspective for analyzing digital audit transformation. Practical Implications The findings highlight important policy considerations for tax authorities and policymakers, particularly the need to strengthen data governance frameworks, invest in auditor digital competencies, and adapt regulatory systems to support the implementation of technology-enabled tax audit mechanisms.
- Research Article
- 10.59141/jrssem.v5i3.1109
- Oct 3, 2025
- Journal Research of Social Science, Economics, and Management
The rapid advancement of digital technologies has driven a fundamental transformation in auditing practices, shifting from manual, sampling-based methods toward automation and data analytics. Although innovations such as Artificial Intelligence, Big Data Analytics, Blockchain, Cloud Computing, and Robotic Process Automation have been introduced, a comprehensive understanding of their applications, challenges, and impacts on audit quality remains limited. This study aims to (1) identify digital innovations applied in the audit process, (2) analyze key challenges in their implementation, and (3) evaluate the impact of digital transformation on audit quality in terms of effectiveness, efficiency, transparency, accuracy, and timeliness. The research employs a systematic literature review (SLR) using the PRISMA protocol and the PICOC framework, covering 84 academic articles published between 2015 and 2025. Data were analyzed through thematic synthesis and narrative synthesis. Findings reveal that AI and data analytics dominate digital audit innovations, while Blockchain enhances transparency and RPA accelerates routine procedures. Key challenges include limited IT infrastructure, organizational resistance, auditors’ skill gaps, and a lack of specific regulations for digital auditing. Nevertheless, digital transformation significantly improves accuracy, efficiency, and real-time anomaly detection, although issues of data integrity and professional ethics persist. Digital transformation enhances audit quality and repositions auditors as strategic partners within organizations. However, its success largely depends on technological readiness, human capital competence, and adaptive regulatory frameworks.
- Research Article
3
- 10.52214/vib.v7i.8403
- Jun 2, 2021
- Voices in Bioethics
Photo by Josh Riemer on Unsplash
 Introduction
 With the rapid advancements in neurotechnological machinery and improved analytical insights from machine learning in neuroscience, the availability of big brain data has increased tremendously. Neurological health research is done using digitized brain data.[1] There must be adequate data governance to secure the privacy of subjects participating in brain research and treatments. If not properly regulated, the research methods could lead to significant breaches of the subject’s autonomy and privacy. This paper will address the necessity for neuroprotection laws, which effectively govern the use of big brain data to ensure respect for patient privacy and autonomy.
 Background
 Artificial intelligence and machine learning can be integrated with neuroscience big brain data to drive research studies. This integrative technology allows patterns of electrical activity in neurons to be studied in detail.[2]Specifically, it uses a robotic system which can reason, plan, and exhibit biologically intelligent behavior. Machine learning is a method of computer programming where the code can adapt its behavior based on big brain data.[3] The big brain data is the collection of large amounts of information for the purpose of deciphering patterns through computer analysis using machine learning.[4] The information that these technologies provide is extensive enough to allow a researcher to read a patient’s mind. AI and machine learning technologies work by finding the underlying structure of brain data, which is then described by patterns known as latent factors, eventually resulting in an understanding of the brain’s temporal dynamics.[5]
 Through these technologies, researchers are able to decipher how the human brain computes its performances and thoughts. However, due to the extensive and complex nature of the data processed through AI and machine learning, researchers may gain access to personal information a patient may not wish to reveal. From a bioethical lens, tensions arise in the realm of patient autonomy. Patients are not able to control the transmission of data from their brains that is analyzed by researchers. Governing brain data through laws may enhance the extent of patient privacy in the case where brain data is being used through AI technologies.[6] A responsible approach to governing brain data would require a sophisticated legal structure.
 Analysis
 Impact on Patient Autonomy and Privacy 
 In research pertaining to big brain data, the consent forms do not fully cover the vast amounts of information that is collected. According to research, personal data has become the most sought out commodity to provide content to corporations and the web-based service industry. Unfortunately, data leaks that release private information frequently occur.[7] The storage of an individual’s data on technologies accessible on the internet during research studies makes it vulnerable to leaks, jeopardizing an individual’s privacy. These data leaks may cause the patient to be identified easily, as the degree of information provided by AI technologies are personalized and may be decoded through brain fingerprinting methods.[8]
 There has been an extensive growth in the development and use of AI. It is efficient in providing information to radiologists who diagnose various diseases including brain cancer and psychiatric disease, and AI assists in the delivery of telemedicine.[9] However, the ethical pitfall of reduced patient autonomy must be addressed by analyzing current AI technologies and creating more options for patient preference in how the data may be used. For instance, facial recognition technology[10] commonly used in health care produces more information than listed in common consent forms, threatening to undermine informed consent. Facial recognition software collects extensive data and may disclose more information than a person would prefer to provide despite being a useful tool for diagnosing medical and genetic conditions.[11] In addition, people may not be aware that their images are being used to generate more clinical data for other purposes. It is difficult to guarantee the data is anonymized. Consent requirements must include informing people about the complexity of the potential uses of the data; software developers should maximize patient privacy.[12] Furthermore, there is a “human element” in the use of AI technologies as medical providers control the use and the extent to which data is captured or accessed through the AI technologies.[13] People must understand the scope of the technology and have clear communication with the physician or health care provider about how the medical information will be used. 
 Existing Laws for Brain Data Governance 
 A strict system of defined legal responsibilities of medical providers will ensure a higher degree of patient privacy and autonomy when AI technologies and data from machine learning are used. Governing specific algorithmic data is crucial in safeguarding a patient’s privacy and developing a gold standard treatment protocol following the procurement of the information.[14] Certain AI technologies provide more data than others, and legal boundaries should be established to ensure strong performance, quality control, and scope for patient privacy and autonomy. For instance, currently AI technologies are being used in the realm of intensive neurological care. However, there is a significant level of patient uncertainty about how much control patients have over the data’s uses.[15] Calibrated legal and ethical standards will allow important brain data to be securely governed and monitored.
 Once brain signals are recorded and processed from one individual, the data may be merged with other data in Brain Computer Interface Technology (BCI).[16] To ensure a right and ability to retrieve personal data or pull it from the collection, specific regulations for varying types of data are needed.[17] The importance of consent and patient privacy must be considered through giving patients a transparent view of how brain data is governed.[18] The legal system must address discriminatory issues and risks to patients whose data is used in studies. Laws like the General Data Protection Regulation (GDPR) and the California Consumer Privacy Protection Act (CCPA) can serve as effective models to protect aggregated data. These laws govern consumer information and ensure the compliance when personal data is collected.[19] California voters recently approved expansion of the CCPA to health data. The Washington Privacy Act, which would have provided rights to access, change, and withdraw personal data, failed to pass. Other states should improve privacy as well,[20] although a federal bill would be preferable. Scientists at the Heidelberg Academy of Sciences argue for data security to be governed in a manner that balances patient privacy and autonomy with the commercial interests of researchers.[21] The balance could be achieved through privacy protections like those in the Washington Privacy Act. Although the Health Insurance Portability and Accountability Act (HIPAA) provides an overall framework to deter the likelihood of dangers to patient protection and privacy, more thorough laws are warranted to combat pervasive data transfer and analysis that technology has brought to the health care industry.[22] Breaches of patient privacy under current HIPAA regulations include releasing patient information to a reporter without their consent and sending HIV data to a patient’s employer without consent.[23] HIPAA does not cover information being shared with outside contractors who do not have an agreement with technology companies to keep patient data confidential. HIPAA regulations also do not always address blatant breaches on patient data confidentiality.[24] Patients must be provided with methods to monitor the data being analyzed to be able to view the extent of private information being generated via AI technologies. In health research, the medical purposes of better diagnosis, earlier detection of diseases, or prevention are ethical justifications for the use of the data if it was collected with permission, the person understood and approved the uses of the data, and the data was deidentified.
 A standard governance framework is required in providing the fairest system of care to patients who allow their brain data to be examined. Informed consent in the neuroscience field could reaffirm the privacy and autonomy of patients by ensuring that they understand the type of information collected. Laws also could protect data after a patient’s death. Malpractice in the scope of brain data could give people a cause of action critical in safeguarding patient’s rights. Data breach lawsuits will become common but generally do not cover deidentified data that becomes part of big data collection. A more synchronized approach to the collection and consent process will encourage an understanding of how big data is used to diagnose and treat patients. Some altruistic people may even be more likely to consent if they know the largescale data collection is helpful to treat and diagnose people. Others should have the ability to opt out of sharing neurological data, especially when there is not certainty surrounding deidentification.[25]
 Conclusion
 Artificial intelligence and machine learning technologies have the potential to aid in the diagnosis and treatment of people globally by extracting and aggregating brain data specific to individuals. However, the secure use of the data is necessary to build trust between care providers and patients, as well as in balancing the bioethical principles of beneficence and patient autonomy. We must ensure the highest quality of care to patients, while protecting their privacy, informed consent, and clinical trust. More sophis
- Research Article
- 10.62951/icistech.v5i1.280
- Jun 30, 2025
- Proceeding of The International Conference of Inovation, Science, Technology, Education, Children, and Health
Indonesia's healthcare system continues to face significant challenges in delivering equitable services across diverse and remote regions. The digital transformation of healthcare—through the integration of Artificial Intelligence (AI), big data, and telemedicine—offers promising solutions to overcome disparities in access, infrastructure, and service delivery. This study aims to comprehensively analyze global and national research trends related to the digital transformation of healthcare using a Systematic Literature Review (SLR) approach, supported by bibliometric analysis through the VOSviewer software. Following the PRISMA protocol, a total of 30 relevant articles published between 2020 and 2025 were identified and analyzed. The network and overlay visualizations generated reveal four major thematic clusters: digital transformation and service quality, big data and pandemic response, AI and data privacy, and community engagement in digital health services. Overlay visualization also shows a clear shift in research focus—from early pandemic responses toward system optimization, ethical governance, and technological inclusivity in recent years. The findings highlight that digital healthcare transformation has increasingly evolved from emergency responses to COVID-19 into a strategic framework for long-term system improvement. Moreover, AI and big data have played pivotal roles in enhancing diagnostics, predicting outbreaks, and improving resource allocation. However, concerns related to privacy, digital literacy, and unequal technological access remain prominent. The study concludes that the integration of AI, big data, and telemedicine not only enhances healthcare efficiency but also requires strong regulatory frameworks, infrastructure readiness, and public engagement. Future research should incorporate co-citation and cross-country comparative analyses to enrich the understanding of digital health transformation in a global context, especially in low- and middle-income countries like Indonesia.
- Research Article
50
- 10.1016/j.fertnstert.2020.10.040
- Nov 1, 2020
- Fertility and Sterility
Predictive modeling in reproductive medicine: Where will the future of artificial intelligence research take us?
- Research Article
- 10.29038/2786-4618-2025-02-31-41
- Jul 21, 2025
- Economic journal of Lesya Ukrainka Volyn National University
Introduction. The rapid digitalization of the economy is fundamentally transforming managerial functions, particularly accounting and business analytics. The widespread adoption of cloud technologies, robotic process automation (RPA), artificial intelligence (AI), blockchain, and big data analytics is redefining the roles of accountants and business analysts, shifting their focus from retrospective data recording to strategic insight generation. The purpose of the article. This article aims to explore the impact of digital transformation on the accounting and analytical functions of enterprises, identify the key technological trends reshaping these professions, and determine the competencies required for professionals to remain relevant and effective in the digital economy. Methods. The research methodology includes a systematic review of recent academic literature and consulting reports (Deloitte, PwC, Gartner, ACCA, Forbes), comparative analysis of traditional versus digital accounting systems, and synthesis of best practices in implementing innovative technologies. Empirical data, expert forecasts, and analytical models were used to assess the potential and challenges of digital integration in financial processes. Results. The study finds that RPA significantly enhances the efficiency and accuracy of routine accounting tasks, enabling professionals to focus on analytical and advisory roles. AI and machine learning support predictive modeling, fraud detection, and scenario planning. Blockchain increases transparency, data immutability, and security in financial transactions, while cloud platforms improve accessibility, scalability, and data integrity. Additionally, modern business analytics evolves into a proactive and prescriptive tool that integrates data from multiple internal and external sources, enabling real-time decision-making and strategic foresight. The roles of accountants and business analysts are converging towards that of data-driven strategic partners, requiring hybrid competencies that merge financial literacy with IT skills, data visualization, and cybersecurity awareness. Conclusions. Digital transformation in accounting and analytics is irreversible and necessitates the rethinking of traditional roles, processes, and education. To remain competitive, enterprises must invest in lifelong learning, IT infrastructure, and corporate cultural adaptation. Professionals must develop hybrid expertise in finance, data science, and business communication. The digital shift is not merely a challenge but a major opportunity for sustainable growth and value creation. Future research should focus on adaptive models of managing accounting systems, integration of diverse data sources, and scenario-based planning supported by predictive analytics tailored to national and international standards.
- Research Article
1
- 10.63278/1533
- Apr 16, 2025
- Metallurgical and Materials Engineering
In the context of the rapid physical pace of digitalization and the emergence of an environmental agenda, the embrace of Industry 4.0 technologies within sustainability frameworks around the globe is in an accelerated state. This is especially true in developing economies, which are looking into how these advances in technology (artificial intelligence (AI), Internet of Things (IoT), Blockchain, and Big Data) propel green transitions and address environmental challenges. The present research involves a systematic literature review of the existing literature that reviews the intersections between digital transformation and environmental sustainability with meaning in developing contexts. Through the Systematic Literature Review (SLR) approach, we reviewed the peer reviewed articles obtained from Scopus, Web of Science, and ScienceDirect databases, based on the PRISMA framework for article selection The review frames the literature, limited to 2010-2023, and synthesizes to structure around the themes of adoption, barriers to implementation, sectoral considerations, and policy considerations. The review indicates that Industry 4.0 technologies present a significant opportunity for improved energy efficiency, waste reduction, and green innovation, but their uptake is limited by infrastructure, availability of digital skills, and weak policy support in less developed contexts. In contrast, enablers include cross-sectoral collaboration, institutional support, and ESG frameworks. Numerous opportunities for application exist including, but not limited to, manufacturing, energy, agriculture, and logistics. The review provides a contribution to knowledge for academia by bringing together otherwise fragmented knowledge and offering a conceptual platform to guide future research. Policymakers and practitioners will find this piece useful in adding structure and direction to emerging digital sustainability strategies. Overall, the review reinforces the necessity for local and inclusive approaches to realise the transformative opportunities Industry 4.0 can present for sustainable development.
- Research Article
- 10.62951/icistech.v5i1.270
- Jun 30, 2025
- Proceeding of The International Conference of Inovation, Science, Technology, Education, Children, and Health
Indonesia's healthcare system continues to face significant challenges in delivering equitable services across diverse and remote regions. The digital transformation of healthcare—through the integration of Artificial Intelligence (AI), big data, and telemedicine—offers promising solutions to overcome disparities in access, infrastructure, and service delivery. This study aims to comprehensively analyze global and national research trends related to the digital transformation of healthcare using a Systematic Literature Review (SLR) approach, supported by bibliometric analysis through the VOSviewer software. Following the PRISMA protocol, a total of 30 relevant articles published between 2020 and 2025 were identified and analyzed. The network and overlay visualizations generated reveal four major thematic clusters: digital transformation and service quality, big data and pandemic response, AI and data privacy, and community engagement in digital health services. Overlay visualization also shows a clear shift in research focus—from early pandemic responses toward system optimization, ethical governance, and technological inclusivity in recent years. The study concludes that the integration of AI, big data, and telemedicine not only enhances healthcare efficiency but also requires strong regulatory frameworks, infrastructure readiness, and public engagement. Future research should incorporate co-citation and cross-country comparative analyses to enrich the understanding of digital health transformation in a global context.
- Research Article
- 10.1002/fsat.3503_11.x
- Sep 1, 2021
- Food Science and Technology
Intrinsic value of food chain data
- Research Article
- 10.36887/2524-0455-2025-3-21
- May 28, 2025
- Actual problems of innovative economy and law
Digitalization penetrates all areas of economic activity, including the banking industry. Therefore, the purpose of the article is to analyze the transformation of banking services in the digital era based on an integrated approach. In the course of conducting scientific research, the authors used a set of scientific methods of cognition: the method of system analysis; the bibliographical method; the method of extrapolation and forecasting, which made it possible to identify patterns and trends in the development of digital transformation in banking. The digital transformation of banking leads to the introduction of innovative technologies to protect personal data, transaction data, the array of collected data, convenience and ease of use of mobile banking applications, speed of payments, etc. Such innovative technologies in the banking and financial sector are: mobile banking, AI, Blockchain, Machine Learning, Internet of Things, Big Data and open banking, etc. Each of the above areas of technological change contains the following characteristics: mobile banking – for digital interaction of banks with their clients; financial consultants based on artificial intelligence – to help in using the main functions of banking services, making informed management decisions; Blockchain – to reduce banking costs, the number of employees and increase the security of transactions; Machine Learning – to obtain a positive customer experience; Internet of Things – to integrate disparate types of equipment, devices for data exchange; Big Data – for collecting, processing, analyzing large data sets and making forecasts based on them and Open Banking – for the free exchange of financial data between different financial providers. The digital technology market is dominated by companies from the USA (Oracle, Microsoft, IBM, Amazon, etc.), as well as the UK and the EU. Forecasts of well-known consulting companies indicate the growth of digital technologies and the gradual transition to a virtual environment, where there is no need for the physical presence of the client in a bank. Keywords: digital transformation, banking structures, banking services, digital technologies, artificial intelligence, mobile banking, Blockchain technologies, Internet of Things, machine learning, big data, open banking.
- Research Article
7
- 10.62304/ijmisds.v1i3.164
- Jun 4, 2024
- Global Mainstream Journal
This systematic literature review examines the integration of Industry 4.0 and Lean technologies in manufacturing, a topic of growing importance as industries seek to enhance efficiency and competitiveness. By analyzing 156 peer-reviewed journal articles, conference papers, and industry reports published between 2010 and 2023, this review identifies vital themes, benefits, challenges, and gaps in the literature. Industry 4.0, characterized by IoT, big data analytics, artificial intelligence (AI), and machine learning (ML), offers significant potential for improving real-time data collection, process automation, and advanced analytics. When integrated with Lean manufacturing principles, which focus on waste reduction and continuous improvement, these technologies can lead to more efficient operations, better quality control, and faster response times. However, the review also highlights several challenges, including high initial costs, the need for a skilled workforce, and the complexity of integrating new technologies with existing systems. Despite these challenges, numerous case studies and best practices demonstrate the successful implementation of these integrated approaches, providing valuable insights for future research and practical applications. This review concludes with recommendations for addressing the identified gaps and leveraging the synergies between Industry 4.0 and Lean technologies to achieve operational excellence in manufacturing.
- Research Article
16
- 10.1108/jfra-09-2024-0639
- Mar 17, 2025
- Journal of Financial Reporting and Accounting
Purpose This paper aims to provide a comprehensive bibliometric approach to analyze the integration of automation and artificial intelligence (AI) in accounting. The study identifies key trends, influential works and future directions to help academics, practitioners and regulators maximize the potential of automation and AI in accounting. Design/methodology/approach This paper conducted a bibliometric analysis, using performance analysis and science mapping techniques to examine 343 articles from the Scopus database covering the period from 2001 to 2024. Preferred Reporting Items for Systematic Reviews and Meta-analysis (PRISMA) protocol was used to ensure a systematic and objective process for identifying, screening and including relevant studies. The analysis used Biblioshiny to generate bibliometric indicators, such as publication trends, thematic maps and insights into co-citation patterns, thematic evolution and the intellectual framework outlining the scope of automation and AI in accounting. Findings The results reveal that the research area is structured around four main conceptual clusters: automation and AI as tools for enhancing accounting practices, the shift toward digital management accounting processes, emerging technologies such as blockchain and Internet of Things for process automation in accounting and auditing and machine learning (ML) and advanced data analytics for fraud detection, real-time reporting and cost optimization. Also, the analysis of theme evolution demonstrates a clear shift from automation (2001–2010) to AI and ML (2011–2020) with digital transformation, big data and data analytics as dominant themes in 2023–2024. Originality/value To the best of the authors’ knowledge, this study is the first comprehensive bibliometric analysis of the available literature on automation and AI in accounting. This analysis fills a critical gap by providing insights into unexpected areas such as the intellectual and social structure of the research area. Using PRISMA and Biblioshiny, this study outlines key trends and gaps, providing guidance for further studies in the digital era.
- Discussion
8
- 10.1016/j.ejmp.2021.05.008
- Mar 1, 2021
- Physica Medica
Focus issue: Artificial intelligence in medical physics.
- Research Article
1
- 10.1016/j.ijme.2025.101287
- Mar 1, 2026
- The International Journal of Management Education
Exploring the impact of artificial intelligence on business talent development in higher education:A systematic literature review and research agenda
- Research Article
32
- 10.3390/bdcc7010003
- Dec 20, 2022
- Big Data and Cognitive Computing
Digital transformation (or digitalization) is the process of continuous further development of digital technologies (such as smart devices, cloud services, and Big Data) that have a lasting impact on our economy and society. In this manner, digitalization is a huge driver for permanent change, even in the field of Sustainable Urban Development. In the wake of digitalization, expectations are changing, placing pressure at the societal level on the design and development of smart environments for everything that means Sustainable Urban Development. In this sense, the solution is the integration of Artificial Intelligence into Sustainable Urban Development, because technology can simplify people’s lives. The aim of this paper is to ascertain which Sustainable Urban Development dimensions are taken into account when integrating Artificial Intelligence and what results can be achieved. These questions formed the basic framework for this research article. In order to make the current state of Artificial Intelligence in Sustainable Urban Development as a snapshot visible, a systematic review of the current literature between 2012 and 2022 was conducted. The data were collected and analyzed using PRISMA. Based on the studies identified, we found a significant growth in studies, starting in 2018, and that Artificial Intelligence applications refer to the Sustainable Urban Development dimensions of environmental protection, economic development, social justice and equity, culture, and governance. The used Artificial Intelligence techniques in Sustainable Urban Development cover a broad field of Artificial Intelligence, such as Artificial Intelligence in general, Machine Learning, Deep Learning, Artificial Neuronal Networks, Operations Research, Predictive Analytics, and Data Mining. However, with the integration of Artificial Intelligence in Sustainable Urban Development, challenges are marked out. These include responsible municipal policies, awareness of data quality, privacy and data security, the formation of partnerships among stakeholders (e.g., local citizens, civil society, industry, and various levels of government), and transparency and traceability in the implementation and rollout of Artificial Intelligence. A first step was taken towards providing an overview of the possible applications of Artificial Intelligence in Sustainable Urban Development. It was clearly shown that Artificial Intelligence is also gaining ground in this sector.
- Research Article
16
- 10.1002/poi3.258
- Jun 14, 2021
- Policy & Internet
The use of big data and data analytics are slowly emerging in public policy-making, and there are calls for systematic reviews and research agendas focusing on the impacts that big data and analytics have on policy processes. This paper examines the nascent field of big data and data analytics in public policy by reviewing the literature with bibliometric and qualitative analyses. The study encompassed scientific publications gathered from SCOPUS (N = 538). Nine bibliographically coupled clusters were identified, with the three largest clusters being big data's impact on the policy cycle, data-based decision-making, and productivity. Through the qualitative coding of the literature, our study highlights the core of the discussions and proposes a research agenda for further studies. 大数据和数据分析的使用已逐渐出现在公共决策中,对此需要展开系统性综述和研究议程,聚焦大数据和数据分析对政策过程产生的影响。本文通过文献计量分析和定性分析,分析了公共政策中大数据和数据分析这一新兴领域。本研究包括了从SCOPUS中选取的科学刊物(N=538)。识别了9个文献耦合簇,其中最大的三个簇分别为:大数据对政策周期的影响、基于数据的决策、生产率。通过对文献进行定性编码,我们的研究强调了探讨的核心,并提出了进一步的研究议程。 El uso de macrodatos y análisis de datos está emergiendo lentamente en la formulación de políticas públicas, y hay pedidos de revisiones sistemáticas y agendas de investigación que se centren en los impactos que los macrodatos y el análisis tienen en los procesos de políticas. Este artículo examina el campo naciente del big data y el análisis de datos en las políticas públicas mediante la revisión de la literatura con análisis bibliométricos y cualitativos. El estudio abarcó publicaciones científicas recopiladas de SCOPUS (N = 538). Se identificaron nueve grupos acoplados bibliográficamente, siendo los tres grupos más grandes el impacto de los macrodatos en el ciclo de políticas, la toma de decisiones basada en datos y la productividad. A través de la codificación cualitativa de la literatura, nuestro estudio destaca el núcleo de las discusiones y propone una agenda de investigación para estudios posteriores. Big data and data analytics have been seen as augmenting knowledge, ultimately leading to better decision-making. Arguments such as that the broad-based use of big data and data analytics will lead to the end-of-theory speak volumes about our expectations of big data and data analytics technologies' transformative power. While industry has been leading the way to test big data and analytics, public actors have been slower to engage (Poel et al., 2018), despite an equal opportunity for big data and data analytics to augment the public policy process. Utilizing big data and data analytics has become a near necessity due to our increasing capability for creating and collecting data at an extraordinary rate. The terms "big data" and "data analytics" have been among the buzzwords of recent years, leading to an upsurge in research, industry, and government applications (Zhou et al., 2014). The increased interest in big data in public policy can be seen in the scientific literature Figure 1, highlighting the increase in big data and analytics related literature. We also see public organizations increasingly engaging with big data analytics to solve challenges like the sustainability crisis and pandemics.1 Scholarly discourse has highlighted case studies and narratives on implementing big data and data analytics in the policy process. However, the literature lacks a systematic view of the current state of big data and data analytics in public policy, and there are identifiable research gaps (Desouza & Jacob, 2017). RQ1. What are the thematic communities of big data and data analytics literature concerning public policy-making? RQ2. What are the research questions emerging under each of the thematic research communities? Our study adopted a mixed-method systematic literature review approach based on a robust empirical bibliometric analysis followed by a qualitative analysis of the core documents to answer these questions. Using a well-established bibliometric method, bibliographic coupling, we identified thematic differences within the literature, and here, we highlight points of departure from the extant literature. The bibliometric analysis was, in turn, used as a basis for the qualitative analysis of the core literature, which is used to propose a research agenda. We find nine contemporary research communities addressing different aspects of big data and data analytics in public policy. While these communities have significant overlap, our analysis identifies them drawing from different theoretical foundations. Moreover, we demonstrate three larger research strands taking different vantage points, namely building strategic capability, data-based decision-making and productivity increases. Finally, our work proposes a research agenda focusing on the role of strategic capability, data-based decision-making, how to address expectations for better services while simultaneously increasing productivity and how to leverage policy analytics and empiricism. Our results offer scholars in public policy a vantage point to the theoretical foundations of research in big data and data analytics in public policy-making. We also draw from the identified communities to highlight emerging research themes that can guide research forward. For policymakers, our results highlight the on-going scholarly debate that focuses on addressing critical issues in the adoption of big data in public policy-making, namely capability building and the extent of data-based decision-making. This article will proceed as follows: next, we review the central elements of big data in policy-making. This is followed by a description of the data and our mixed-method approach. Finally, the empirical results are described and followed by a discussion to make sense of the research themes emerging from the analysis. "Big data" is a general term used for the process of gathering massive amounts of data from different sources. Sources can include human-input data but also includes data from sensors or different types of monitoring systems that create process data while running. It is clear that we are accumulating data at a never before seen rate. Already, in 2014, the pace was staggering, with 90% of the world's data being collected during the prior 2 years and 2.5 quintillion bytes of data added each day (Kim et al., 2014). Having access to massive amounts of data has enabled significant innovation in both the public and private domains. Looking at companies like Google and Amazon, with their innovation of new services for consumers, or at the recent ability for doctors to detect cancer cells more precisely thanks to massive training data about what a cancerous cell is, we can see that we are very much on the cusp of creating a broad utility of big data and analytics. This has been seen as a shift in the Industrial Revolution's magnitude (Richards & King, 2014) and has been widely hyped in business (Margetts & Sutcliffe, 2013). That said, public policy is not at the forefront of the use of big data and data analytics in decision making (Kaski et al., 2019; Poel et al., 2018). This nonadoption is due to multiple factors limiting these technologies' utility (Malomo & Sena, 2017). The ever-increasing amount of data offers possibilities for discovering new relationships and inferencing a multitude of problems. However, this comes with new challenges involving reproducibility, complexity, security, and risks to privacy and a need for new technology and human skills. This is very much the case in public policy, where we need to clearly identify where big data can add value in an ethical and trustworthy manner. In a review, Giest (2017) highlighted three underlying factors to consider. First, institutional capacities have a significant role in the use of big data in public policy, producing solutions that can enable users to easily interact with data while also taking into account the siloed data structures in the public domain. However, we know from previous research that siloed structures are an important limiting factor for public policy utilization of big data (Malomo & Sena, 2017). Second, hand-in-hand with big data comes the broader digitalization of public services. Digitalization allows for mediums to interact with big data but also enables the creation of new data. There is, however, evidence that digitalization changes the interactions between citizens and public officials and requires new skills from both parties. Third, big data information will have an impact on the policy cycle. Studies have found that there has been limited progress in taking advantage of big data and analytics (Poel et al., 2018) because it requires a significant change in the policy cycle (Höchtl et al., 2016). Giest (2017) highlights two issues, the substantive role and the procedural role of big data in policy instruments. Procedural activities focus on regulatory activities, such as enabling open data, while substantive actions relate to collecting data for enhancing, for example, evidence-based policy making. Capacities, digitalization, and the role of big data in the (substantive and procedural) policy cycle are core to digital-era governance and evidence-based policy making. In this, it is important to note that policy-makers are not a homogeneous group, and policy cycles vary. Thus, the objectives of analytics throughout the policy cycle vary significantly (Daniell et al., 2016) whether or not we approach the policy cycle as separate discrete stages (Jann & Wegrich, 2007), and it has been shown that big data analytics, when used more in some policy stages than in others, notably improved government transparency, policy evaluation, foresight, and agenda setting (Poel et al., 2018). This should be reflected against findings that data analytics have been politically significant in all policy cycle stages (Van der Voort et al., 2019). To overcome the challenges, Poel et al. (2015) highlighted multiple topics that must be addressed to enable capacity building, digitalization, and data integration into the policy cycle. These are (1) a skills gap, (2) reduced transparency due to data analytics, (3) sources and tools, (4) standardization of methods and tools, (5) linking of policy experiments with impact assessments, and (6) enabling policy-makers to be informed about the tools that are developed and piloted. The highlighted themes give context to the issue of big data in policy. While we see the significant impacts being created by the use of big data in policy making, along with the subsequent adaptation of data analytics, we need to better explain and make transparent the utility and complementarity of big data-driven analyses for the policy cycle (Vydra & Klievink, 2019). The challenge highlighted by Poel et al. (2015) and Giest et al. (2017), is also reflected in Pencheva et al. (2020) and Ingram (2019). Both note that big data in public policy has focused more on the "techno-rational factors," dismissing the importance of interaction with the policy process. We know that technology adoption is dependent on the perceived usefulness by the user (Venkatesh & Davis, 2000) and that there is scepticism towards the use of big data and data analytics is public policy-making (Guenduez et al., 2020). This can be the result of a mismatch in practise and expectations. Durrant et al. (2018) show how there is a aspirational motivation to the use of big data and data analytics, not reflected by its everyday utility. While we know that by employing data drive approaches, there is significant potential for anticipatory governance, interaction among stakeholders is the key to draw value from big data and analytics (Maffei et al., 2020; Starke & Lünich, 2020). While the increased stakeholder involvement does not protect from big data and data analytics policy-making creating hard to detect inequalities (Giest & Samuels, 2020; van Veenstra et al., in press). However, engaging with a large pool of stakeholders in the public policy-making process will increase the complexities of adopting big data and data analytics (Janssen et al., 2017). In addition to stakeholder interaction, the ability to build capacity and also evaluate deficiencies is important (Okuyucu & Yavuz, 2020). Building capacity should not be merely seen as the technical capacity, while that is also important (Poel et al., 2018), but a more holistic capability to integrate big data and analytics into the policy cycle (Höchtl et al., 2016). Policy-making organizations, while being exceptionally technically capable, can be in a situation where the benefits from big data and data analytics remain small due to "applications" not fitting "their organizations and main statutory tasks" (Klievink et al., 2017). This to say that when we talk about big data and data analytics capabilities in public policy-making, the literature focuses on technical issues but also on the ability of big data and data analytics to produce policy-relevant applications. While we see an increasing and diverse set of research addressing different challenges of big data and data analytics, the current body of literature lacks holistic research agendas (Desouza & Jacob, 2017) addressing the issues highlighted from practice by Giest (2017) and Poel et al. (2015). While we can note emerging fields such as policy analytics (De Marchi et al., 2016; Tsoukias et al., 2013), there is a need to better understand the theoretical grounding and research gap of big data and data analytics in public policy-making. This study's methodological approach was based on a mixed method of quantitative and qualitative analyses of the bibliometric data and the publication's content. The selected four-step mixed-methods approach, described below, enables a holistic approach to comprehending the current state-of-the-art and allows us to propose an agenda for going forward. The first phase focused on retrieving the sample of relevant articles and their bibliometric data for analysis. The second phase involved the bibliometric analysis of the retrieved data, which was performed by analysing descriptive statistics, bibliographical coupling, network analytics, and community detection. By gaining a comprehensive view of the more extensive body of literature, we could implement more filtering process based on eigenvector centrality to get a shortlist of papers for the next phase. The third phase, qualitative analysis, continued the process with an in-depth review and coding of the articles' full text. Finally, in the fourth phase, Synthesis, we draw insights from the MAXQDA coding analysis and reporting. This four-phase process is shown in Figure 2. The data used in this study were retrieved from the Scopus database. Scopus is Elsevier's abstract and citation database that has over 1.7 billion cited references dating back to 1970. A central aspect of the quality of the results is that the query used to search for relevant articles was correctly designed. The study focuses on public policy-making and big data and data analytics. The study's scope is relatively narrow, focusing solely on publications that address policy-making and the policy process. The decision of the scope excludes articled that focus on, for example, big data or data analytics, but lack the specific aspect of policy-making. This was the key inclusion criteria of articles into the data set. To focus on this specific scope, we used an iterative approach where multiple search strings were tested, and after each search, the abstracts of the 10 most cited articles and the 10 most recent articles were reviewed to understand if the query results reflected the objectives of the study. In practice, the process started with a seed query of "big data" or "data analytics" and "public policy." The query results were reviewed to estimate which articles focused on big data or data analytics and policy-making. These articles were reviewed to see if new terms emerged through the titles, abstracts, and keywords that needed to be included in the analysis. The process adjusted based on a subjective evaluation of the number of false-positives in the 10 most recent and 10 most cited publications and the number of articles retrieved. This method of short-listing the important literature is known as the snowball method, and the process includes consulting the bibliographies of the key documents to find other relevant titles in the subject (Jalali & Wohlin, 2012). After multiple tests of a comprehensive query that also limited the number of false-positives, we downloaded the metadata for 538 documents. These documents were retrieved using the query "public policy," "policy analysis," "policy making," or "public administration," with the terms "big data," "data analytics," or "automated decision-making" in the title, abstract, or keywords of the document. To analyze the literature, we used the well-established bibliometric method of bibliographical coupling. Bibliographical coupling allows for analysis of the publications' shared intellectual background (Kessler, 1963), highlighting contemporary research (Youtie et al., 2013). It is an approach to analyzing the shared theoretical background of scientific publications where the link between documents is calculated by the number of references the two documents share. Kessler (1963) elaborates, "A single item of reference shared by two documents is defined as a unit of coupling between them," and if multiple references are shared, the weight of the coupling increases. Bibliographical coupling is able to highlight hot topics (Glanzel & Czerwon, 1996) and links documents with a similar research focus (Jarneving, 2007), ultimately creating a "contemporaneous representation of knowledge" (Youtie et al., 2013). This approach has been used in several research papers to form the basis for research agenda building (Suominen et al., 2019; Yuan et al., 2015). Using the retrieved publication metadata, the VOSviewer tool (van Eck & Waltman, 2009) was selected to calculate bibliographical coupling weights for all the documents in our data set. VOSviewer is a free tool used for bibliometrics and was selected due to the exports available in the software allowing for deeper network analysis in Gephi. The SCOPUS data export was used as an input to the VOSviewer. During the analysis process, we selected documents as the level of analysis, minimum number of citations for a document was set to zero and the full set was selected for the analysis. The full counting method, which assigns each researcher with full credit of one publication rather than a fractional share per the number of authors, was used for the calculation method. Finally, we accepted VOSviewer default to keep the most extensive set of related items, which limited the analysis to the largest, by created by the bibliographical coupling analysis. This limited the analysis to documents. Bibliographical coupling analysis of a data set a by a set of and a set of we calculated the link weight between each publication we created a This data, with the VOSviewer was to because it allows for more network and community detection. network analysis, network descriptive for example, for were calculated in Gephi. were identified using et al., The is one of the most methods to find of in The methods by each to a separate community there after the in by in a This process is continued through the ultimately creating a new network with to communities et al., for a In the the can be by a that the number of communities the This was to the number of We increased the value the community has one share of the documents. We also calculated the eigenvector centrality for each document in the data set. In eigenvector centrality the a has in A value is calculated to all based on an that to other important are more important than equal This centrality value a publication's to the network created by the bibliographical coupling analysis. for we identified the most central publications from each community to be selected for systematic the most central publications was to keep the sample of documents in the coding phase and the most eigenvector central publications was a method to the most important publications from each community for a deeper analysis. of study of study or empirical study findings The offer the as a to be to the study's at For the current we adopted the of elements in the creating a approach (1) the context of the (2) and (3) the scope defined as a research We the or empirical study and to be the (4) for the theoretical or methodological We also included a separate for (5) the method used to understand the specific approach used in the study. For the key we also created The first focused on (6) results and the second focused on discussion and This was to highlight other study results and to the analysis in the scholarly debate the research To add to the coding process, the qualitative analysis was using MAXQDA The software allows for coding the documents with the and drawing from the created The use of the software transparency and the of the analysis et al., & 2012). After the coding was the coding process was by one the eigenvector central documents from each and the documents based on the The second researcher the role of the The interaction between the was to that all of the in the coding were The second researcher through the papers and by the first researcher to make that each if was However, as publications can the our approach was to we not the information multiple per This of an approach for example, The MAXQDA analysis software used in the coding created of the documents and created a document about the of text. These were used in the phase. The of by MAXQDA was used to the The two researcher with the MAXQDA of the communities After review of the the on their There after the to draw insights from the coding of the core of The retrieved publications are with the first publication in the data set in as seen in Figure The data were retrieved in 2020; publications for the first and one should a in publication While Figure it is important to this into Figure the of the topics of big data and data analytics in the public policy related literature concerning the public policy literature for the The publication volumes are on 10 to be able to them in one It is clear that the body of literature focused on big data and data analytics in public policy is a much increase in interest in to the public policy body of literature. of the publications are from or seen in 1, the three have over highlights all with over publications in the data set. the are highlighting the use of big data and data analytics for example, the topics of and issues and The descriptive analyses of the data also highlight the different that are on the shown in the publication sources for the articles included in the data set are and information we should note our The study focused solely on big data or data analytics and its use in public policy-making. In the the number of articles focusing on single for example, big data is much We should also note that the search in SCOPUS at the title, abstract and making articles big data or data analytics and policy-making in by the It is that the publication sources with at publications all or of all This that publications are over different publication and an on-going debate on the subject is hard to to a specific The of the publications by with that of scientific publication with is significantly the with The two largest are followed by the and and with Looking at the of the between different seen in of has a significant of publications in the data and with the publication there were publications per To understand more in-depth the we the and keywords of the publications in the We the and fields by and in the Figure the terms are as focusing on big data and public policy. thematic terms emerging to the most terms relate to and into the articles' the descriptive analyses highlight that the data set recent among multiple sources and with relatively volumes by The terms in the publications with the search The 538 publications' bibliometric data were for bibliographic coupling using VOSviewer During the calculation process, the software first if of the publications were and from the network emerging from the data. VOSviewer the largest in the full and offers an to used the largest set. to use the largest set allows focusing on the core the In our data the largest set of creating a network was documents. The documents were from the as these articles not be included in the bibliometric coupling based clusters but remain throughout the analysis. with the sample of the VOSviewer created network was to software for further analysis. were and the bibliographic coupling network a with the of the network being This to say that there is significant between in the communities as is but the full network is not as the is the communities were using the et al., The was increased from its default value of the community an share of the documents. a of the in nine with the the community has of the The largest community of the A representation of the created can be seen in Figure In this the the clusters created using the The of a the number of citations of an article in the the the network was created by two large communities