A Survey on Recent Advances in Conversational Data Generation
Recent advancements in conversational systems have significantly enhanced human-machine interactions across various domains. However, training these systems is challenging due to the scarcity of specialized dialogue data. Traditionally, conversational datasets were created through crowdsourcing, but this method has proven costly, limited in scale, and labor-intensive. As a solution, the development of synthetic dialogue data has emerged, utilizing techniques to augment existing datasets or convert textual resources into conversational formats, providing a more efficient and scalable approach to dataset creation. In this survey, we offer a systematic and comprehensive review of multi-turn conversational data generation, focusing on three types of dialogue systems: open domain, task-oriented, and information-seeking. We categorize the existing research based on key components like seed data creation, utterance generation, and quality filtering methods, and introduce a general framework that outlines the main principles of conversation data generation systems. Additionally, we examine the evaluation metrics and methods for assessing synthetic conversational data, address current challenges in the field, and explore potential directions for future research. Our goal is to accelerate progress for researchers and practitioners by presenting an overview of state-of-the-art methods and highlighting opportunities to further research in this area.
- Research Article
- 10.7717/peerj-cs.3067
- Aug 21, 2025
- PeerJ Computer Science
Conversational recommender systems (CRS) facilitate natural language interactions for more effective item suggestions. While these systems show promise, they face challenges in effectively utilizing and integrating informative data with conversation history through semantic fusion. In this study we present an innovative framework for extracting social information from conversational datasets by inferring ratings and constructing user-item interaction and user-user relationship graphs. We introduce a social information sensitive semantic fusion (SISSF) method that employs contrastive learning (CL) to bridge the semantic gap between generated social information and conversation history. We evaluated the framework on two public datasets (ReDial and INSPIRED) using both automatic and human evaluation metrics. Our SISSF framework demonstrated significant improvements over baseline models across all metrics. For the ReDial dataset, SISSF achieved superior performance in recommendation tasks (R@1: 0.062, R@50: 0.437) and conversational quality metrics (Distinct-2: 4.223, Distinct-3: 5.595, Distinct-4: 6.155). Human evaluation showed marked improvement in both fluency (1.81) and informativeness (1.63). We observed similar performance gains on the INSPIRED dataset, with notable improvements in recommendation accuracy (R@1: 0.046, R@10: 0.129, R@50: 0.269) and response diversity (Distinct-2: 2.061, Distinct-3: 4.293, Distinct-4: 6.242). The experimental results consistently validate the effectiveness of our approach in both recommendation and conversational tasks. These findings suggest that incorporating social context through CL can significantly improve the personalization and relevance of recommendations in conversational systems.
- Research Article
33
- 10.1109/tai.2022.3229289
- Jan 1, 2024
- IEEE Transactions on Artificial Intelligence
Synthetic tabular data generation becomes crucial when real data is limited, expensive to collect, or simply cannot be used due to privacy concerns. However, producing good quality synthetic data is challenging. Several probabilistic, statistical, generative adversarial networks (GANs), and variational auto-encoder (VAEs) based approaches have been presented for synthetic tabular data generation. Once generated, evaluating the quality of the synthetic data is quite challenging. Some of the traditional metrics have been used in the literature but there is lack of a common, robust, and single metric. This makes it difficult to properly compare the effectiveness of different synthetic tabular data generation methods. In this paper we propose a new universal metric, TabSynDex, for robust evaluation of synthetic data. The proposed metric assesses the similarity of synthetic data with real data through different component scores which evaluate the characteristics that are desirable for “high quality” synthetic data. Being a single score metric and having an implicit bound, TabSynDex can also be used to observe and evaluate the training of neural network based approaches. This would help in obtaining insights that was not possible earlier. We present several baseline models for comparative analysis of the proposed evaluation metric with existing generative models. We also give a comparative analysis between TabSynDex and existing synthetic tabular data evaluation metrics. This shows the effectiveness and universality of our metric over the existing metrics. Source Code: <uri xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">https://github.com/vikram2000b/tabsyndex</uri>
- Research Article
7
- 10.1186/s12911-024-02731-9
- Feb 18, 2025
- BMC Medical Informatics and Decision Making
BackgroundThe exponential growth in patient data collection by healthcare providers, governments, and private industries is yielding large and diverse datasets that offer new insights into critical medical questions. Leveraging extensive computational resources, Machine Learning and Artificial Intelligence are increasingly utilized to address health-related issues, such as predicting outcomes from Electronic Health Records and detecting patterns in multi-omics data. Despite the proliferation of medical devices based on Artificial Intelligence, data accessibility for research is limited due to privacy concerns. Efforts to de-identify data have met challenges in maintaining effectiveness, particularly with large datasets. As an alternative, synthetic data, that replicate main statistical properties of real patient data, are proposed. However, the lack of standardized evaluation metrics complicates the selection of appropriate synthetic data generation methods. Effective evaluation of synthetic data must consider resemblance, utility and privacy, tailored to specific applications. Despite available metrics, benchmarking efforts remain limited, necessitating further research in this area.ResultsWe present SynthRO (Synthetic data Rank and Order), a user-friendly tool for benchmarking health synthetic tabular data across various contexts. SynthRO offers accessible quality evaluation metrics and automated benchmarking, helping users determine the most suitable synthetic data models for specific use cases by prioritizing metrics and providing consistent quantitative scores. Our dashboard is divided into three main sections: (1) Loading Data section, where users can locally upload real and synthetic datasets; (2) Evaluation section, in which several quality assessments are performed by computing different metrics and measures; (3) Benchmarking section, where users can globally compare synthetic datasets based on quality evaluation.ConclusionsSynthetic data mitigate concerns about privacy and data accessibility, yet lacks standardized evaluation metrics. SynthRO provides an accessible dashboard helping users select suitable synthetic data models, and it also supports various use cases in healthcare, enhancing prognostic scores and enabling federated learning. SynthRO’s accessible GUI and modular structure facilitate effective data evaluation, promoting reliability and fairness. Future developments will include temporal data evaluation, further broadening its applicability.
- Research Article
4
- 10.13053/cys-25-1-3899
- Feb 15, 2021
- Computación y Sistemas
Question answering (QA), one of the important applications of Natural Language Processing (NLP) aims to take the user questions and returned to the user with the answers. An open domain QA system deals with a set of questions that can be of any domain. The other type of QA is close-domain where it deals with the questions under a specific domain e.g., agriculture, medicine, education, tourism, etc. Our cooking question answering system is an example of a closed domain QA system. Here, users can ask the cooking related questions and the system returns the actual answer to the user. In the present article, we have proposed different modules of a cooking QA system. In addition to dataset preparation, the development of a cooking ontology, the classification of questions as well as the extraction of candidate answers are also treated as other important aspects which are discussed in this paper in details. In the cooking QA system, automatic evaluation metrics are like precision, recall, f-score, and c@1 have been used for the evaluation of precise answers. Also, human evaluation is used on the basis of a rating scale. Moreover, the recommendation of recipes has also been attempted and the evaluation metrics show satisfactory performances of the systems.
- Research Article
3
- 10.3390/s24092750
- Apr 25, 2024
- Sensors
Biometric authentication plays a vital role in various everyday applications with increasing demands for reliability and security. However, the use of real biometric data for research raises privacy concerns and data scarcity issues. A promising approach using synthetic biometric data to address the resulting unbalanced representation and bias, as well as the limited availability of diverse datasets for the development and evaluation of biometric systems, has emerged. Methods for a parameterized generation of highly realistic synthetic data are emerging and the necessary quality metrics to prove that synthetic data can compare to real data are open research tasks. The generation of 3D synthetic face data using game engines' capabilities of generating varied realistic virtual characters is explored as a possible alternative for generating synthetic face data while maintaining reproducibility and ground truth, as opposed to other creation methods. While synthetic data offer several benefits, including improved resilience against data privacy concerns, the limitations and challenges associated with their usage are addressed. Our work shows concurrent behavior in comparing semi-synthetic data as a digital representation of a real identity with their real datasets. Despite slight asymmetrical performance in comparison with a larger database of real samples, a promising performance in face data authentication is shown, which lays the foundation for further investigations with digital avatars and the creation and analysis of fully synthetic data. Future directions for improving synthetic biometric data generation and their impact on advancing biometrics research are discussed.
- Conference Article
- 10.1145/3184558.3186341
- Jan 1, 2018
To establish an automatic conversation system between human and computer is regarded as one of the most hardcore problems in computer science. It requires interdisciplinary techniques of information retrieval, natural language processing, data management as well as artificial intelligence. The arrival of big data era reveals the feasibility to create a conversation system empowered by data-driven approaches. Now we are able to collect extremely large conversational data on Web, and organize them to launch a human-computer conversation system. Owing to the diversity of Web resources available, a retrieval-based conversation system will be able to find at least some responses from the massive data repository for any user inputs. Given a human issued utterance, i.e., a query, a retrieval-based conversation system will search for appropriate replies, conduct a relevance ranking, and then output the highly relevant one as the response. In this paper, we propose a novel retrieval model named NeuRetrieval for short text understanding, representation and semantic matching. The proposed model is general and unified for both single-turn and multi-turn conversation scenarios in open domain. In the experiments, we investigate the effectiveness of the proposed deep neural network model for human-computer conversations. We demonstrate performance improvement against a series of baseline methods in several evaluation metrics. In contrast with previously proposed methods, NeuRetrieval is tailored for conversation scenarios and demonstrated to be more effective.
- Research Article
- 10.2196/78082
- Oct 24, 2025
- JMIR Formative Research
BackgroundThe rise of artificial intelligence and accessible audio equipment has led to a proliferation of recorded conversation transcripts datasets across various fields. However, automatic mass recording and transcription often produce noisy, unstructured data that contain unintended recordings such as hallway conversations, media (eg, TV, radio), or transcription inaccuracies as speaker misattribution or misidentified words. As a result, large conversational transcript datasets require careful preprocessing and filtering to ensure their research utility. This challenge is particularly relevant in behavioral health contexts (eg, therapy, counseling) where deriving meaningful insights, specifically dynamic processes, depends on accurate conversation representation.ObjectiveWe present a framework for preprocessing large datasets of conversational transcripts and filtering out non-sessions—transcripts that do not reflect a behavioral treatment session but instead capture unrelated conversations or background noise. This framework is applied to a large dataset of behavioral health transcripts from community mental health clinics across the United States.MethodsOur approach integrated basic feature extraction, human annotation, and advanced applications of large language models (LLMs). We began by mapping transcription errors and assessing the number of non-sessions. Next, we extracted statistical and structural features to characterize transcripts and detect outliers. Notably, we used LLM perplexity as a measure of comprehensibility to assess transcript noise levels. Finally, we used zero-shot prompting with an LLM to classify transcripts as sessions or non-sessions, validating its output against expert annotations. Throughout, we prioritized data security by selecting tools that preserve anonymity and minimize the risk of data breaches.ResultsInitial assessment revealed that transcription errors—such as incomprehensible segments, unusually short transcripts, and speaker diarization issues—were present in approximately one-third (n=36 out of 100) of a manually reviewed sample. Statistical outliers revealed that high speaking rate (>3.5 words per second) is associated with short transcripts and answering machine messages, while short conversation duration (<15 min) was an indicator for case management sessions. The 75th percentile of LLM perplexity scores was significantly higher in non-sessions than sessions (permutation test mean difference = −258, P =.02), although this feature alone offered only moderate classification performance (precision =0.63, recall =0.23 after outlier removal). In contrast, zero-shot LLM prompting effectively distinguished sessions from non-sessions with high agreement to expert ratings (κ=0.71) while also capturing the nature of the meeting.ConclusionsThis study’s hybrid approach effectively characterizes errors, evaluates content, and distinguishes text types within unstructured conversational dataset. It provides a foundation for research on conversational data, key methods, and practical guidelines that serve as crucial first steps in ensuring data quality and usability, particularly in the context of mental health sessions. We highlight the importance of integrating clinical experts with artificial intelligence tools while prioritizing data security throughout the process.
- Research Article
2
- 10.1142/s0218194025500032
- Jan 27, 2025
- International Journal of Software Engineering and Knowledge Engineering
The paper addresses the limitations of traditional evaluation metrics for Question Answering (QA) systems that primarily focus on syntax and n-gram similarity. We propose a novel model-based evaluation metric, MQA-metric, and create a human-judgment-based dataset, squad-qametric and marco-qametric, to validate our approach. The research aims to solve several key problems: the objectivity in dataset labeling, the effectiveness of metrics when there is no syntax similarity, the impact of answer length on metric performance, and the influence of real answer quality on metric results. To tackle these challenges, we designed an interface for dataset labeling and conducted extensive experiments with human reviewers. Our analysis shows that the MQA-metric outperforms traditional metrics like BLEU, ROUGE and METEOR. Unlike existing metrics, MQA-metric leverages semantic comprehension through large language models (LLMs), enabling it to capture contextual nuances and synonymous expressions more effectively. This approach sets a standard for evaluating QA systems by prioritizing semantic accuracy over surface-level similarities. The proposed metric correlates better with human judgment, making it a more reliable tool for evaluating QA systems. Our contributions include the development of a robust evaluation workflow, creation of high-quality datasets, and an extensive comparison with existing evaluation methods. The results indicate that our model-based approach provides a significant improvement in assessing the quality of QA systems, which is crucial for their practical application and trustworthiness.
- Dissertation
- 10.26686/wgtn.27014419
- Sep 13, 2024
<p><strong>Because of its wide range of applications, generative artificial intelligence (Generative AI) has received a lot of attention in academia and industry. Despite significant advances in computer vision and natural language processing, the use of Generative AI in tabular data is still underexplored. This gap is especially significant given the prevalence of tabular data as the primary data modality. This thesis seeks to close this gap by focusing on the efficient generation, evaluation, refinement, and application of tabular data, while addressing the challenges inherent in its heterogeneous nature.</strong></p><p>The primary goal of this thesis is to develop efficient algorithms for tabular data synthesis, advancing the field of Generative AI in the tabular domain. To achieve this goal, four specific research objectives have been outlined.</p><p>First, this thesis addresses gaps in evaluation metrics by unifying a framework for the comprehensive and consistent assessment of synthetic data. The proposed framework is designed to meet diverse business requirements across various downstream tasks by incorporating a wide range of advanced metrics. These metrics cover various types of evaluations, including univariate, bivariate, multivariate, cluster, and record-level evaluations. Additionally, standardized visualizations are provided to facilitate qualitative assessments. The results indicate that the proposed evaluation framework not only enables consistent comparisons, rankings, and the selection of different data synthesis approaches but also acts as a valuable tool to promptly assess the reliability of results obtained from synthetic data.</p><p>Second, this thesis presents innovative algorithms for the generation of tabular data, aiming to significantly enhance the quality of data synthesis. One of the contributions is the development of a reversible feature engineering pipeline designed to automatically represent tabular data in an efficient format while also ensuring that the transformed data can be easily converted back to its original format. Additionally, the thesis proposes novel deep learning-based tabular data generation models that are capable of learning the joint distribution of multivariate datasets without relying on predefined distribution assumptions. Experimental results highlight the effectiveness of the proposed algorithms in handling heterogeneous data, demonstrating superior performance in most scenarios within the same training duration when compared to similar alternatives.</p><p>Third, this thesis introduces innovative synthetic data prototype selection algorithms with the goal of refining the generated samples. This approach aims to leverage the advantages of Deep Generative Models, which, once trained, can generate unlimited and diverse synthetic data. By carefully identifying and selecting high-quality samples or removing unrealistic ones, the quality of the synthetic data can be enhanced from a post-processing perspective. Building on this hypothesis, we recognize the iterative nature inherent in the data synthesis procedure, where the processes of data generation, evaluation, and refinement should operate repeatedly in a cyclical flow. During the data generation phase, the evaluation framework assists in identifying potential issues and risks based on use cases, while the refinement (post-processing) step iteratively improves synthetic data in alignment with the evaluation outcomes.</p><p>Last, this thesis applies and validates the proposed models across various domains, including business, healthcare, and government, each requiring distinct downstream tasks. Specifically, we assess our algorithms from the standpoint of data balancing using churn data, evaluate our models focusing on data augmentation with health data, and test our algorithms from the perspective of data representation considering data privacy concerns within sensitive citizen data. These applications serve as practical illustrations, demonstrating the effectiveness and utility of the proposed Generative AI models in real-world scenarios.</p><p>In summary, this thesis makes a substantial contribution to the advancement of Generative AI within the domain of tabular data. It not only presents innovative algorithms and evaluation methods but also introduces practical frameworks applicable to real-world scenarios. The findings demonstrate the considerable capacity of Generative AI to fundamentally transform tabular data in various fields, ultimately leading to improved data availability, quality, and quantity.</p>
- Book Chapter
9
- 10.1007/978-3-030-62365-4_3
- Jan 1, 2020
Differentially private algorithmic synthetic data generation (SDG) solutions take input datasets \(D_p\) consisting of sensitive, private data and generate synthetic data \(D_s\) with similar qualities. The importance of such solutions is increasing both because more and more people realize how much data is collected about them and used in machine learning contexts, as well as a consequence of newly introduced data privacy regulations, e.g. the EU’s General Data Protection Regulation (GDPR). We aim to develop a novel and composite SDG evaluation metric which takes into account macro-statistical dataset similarities and data utility in machine learning tasks against privacy boundaries of the synthetic data. We formalize the mathematical foundations for quantitatively measuring both the statistical similarities and the data utility of synthetic data. We use two well-known datasets containing (potentially) personally identifiable information as inputs (\(D_p\)) and existing SDG algorithms PrivBayes and DPGroupFields to generate synthetic data (\(D_s\)) based on them. We then test our evaluation metric for different values of privacy budget \(\epsilon \). Based on our experiments we conclude that the proposed composite evaluation metric is appropriate for quantitatively measuring the quality of synthetic data generated by different SDG solutions and possesses an expected sensitivity to various privacy budget values.KeywordsSynthetic data generationDifferential privacyEvaluation metrics
- Book Chapter
9
- 10.1007/978-981-19-7447-2_43
- Jan 1, 2023
Artificial Intelligence (AI) has become the key driving force in Industrial Automation. Machine learning (ML) and Deep Learning (DL) can be considered to be the components of AI which rely on data for model training. Data generation has increased due to the Internet, connected devices, mobile devices and social networking which in turn have also given rise to cybercrime and cyber thefts. To prevent those and preserve the identity of individuals in the public data, government and policymakers have put stringent privacy-preserving laws. The economy of data collection, quality of data in the public domain, and data bias have made data accessibility and its usage a challenge for AI/ML training for research work or industrial purposes. This has forced researchers to look into the alternative. Synthetic Data offers a promising solution to overcome the data challenges. The last few years have seen many studies conducted to verify the utility and privacy protection capability of synthetic data. However, all of these have been exploratory. This paper focuses on various methods of synthetic data generation and their validation metrics. It opens up a few questions that need further study before we conclude that synthetic data offers a universal solution for AI and ML.KeywordsArtificial Intelligence (AI)Machine Learning (ML) synthetic dataStatistical disclosure limitationDifferential privacy
- Research Article
22
- 10.1093/jamia/ocae015
- Feb 16, 2024
- Journal of the American Medical Informatics Association : JAMIA
Question answering (QA) systems have the potential to improve the quality of clinical care by providing health professionals with the latest and most relevant evidence. However, QA systems have not been widely adopted. This systematic review aims to characterize current medical QA systems, assess their suitability for healthcare, and identify areas of improvement. We searched PubMed, IEEE Xplore, ACM Digital Library, ACL Anthology, and forward and backward citations on February 7, 2023. We included peer-reviewed journal and conference papers describing the design and evaluation of biomedical QA systems. Two reviewers screened titles, abstracts, and full-text articles. We conducted a narrative synthesis and risk of bias assessment for each study. We assessed the utility of biomedical QA systems. We included 79 studies and identified themes, including question realism, answer reliability, answer utility, clinical specialism, systems, usability, and evaluation methods. Clinicians' questions used to train and evaluate QA systems were restricted to certain sources, types and complexity levels. No system communicated confidence levels in the answers or sources. Many studies suffered from high risks of bias and applicability concerns. Only 8 studies completely satisfied any criterion for clinical utility, and only 7 reported user evaluations. Most systems were built with limited input from clinicians. While machine learning methods have led to increased accuracy, most studies imperfectly reflected real-world healthcare information needs. Key research priorities include developing more realistic healthcare QA datasets and considering the reliability of answer sources, rather than merely focusing on accuracy.
- Conference Article
2
- 10.1109/icpeca56706.2023.10076240
- Jan 29, 2023
Enterprises produce massive amounts of data every day. Data records generated in various formats are usually classified into structured, semi-structured, and unstructured data [1]. Many transactional data records are crucial and need to be exchanged between systems, thus data conversion becomes necessary and even tedious. Moreover, decision-makers always make "ad hoc" requests that require searching within large volumes of data. Therefore, an intelligent system is needed to respond rapidly to the demands of modern enterprises. In this paper, we purpose a novel method to build a question-answering (QA) system from a transactional system. We use object models that are translated from a star schema to represent the transactional system, such that it can generate questions (Natural Language, NL) and answers (SQL statements) as the training data. Then, we use an end-to-end (E2E) neural network to train a QA system with the generated data. Our experiments show that the Long Short-Term Memory (LSTM) network with a 0.95 BLEU value is more accurate than the Gated Recurrent Unit (GRU) network with a 0.90 BLEU value. Consequently, the proposed method can automatically generate training data from object models, and the trained artificial intelligence (AI) model can further become a QA system for us to ask questions directly.
- Research Article
- 10.1093/nargab/lqag012
- Feb 11, 2026
- NAR Genomics and Bioinformatics
Synthetic data (SD) has become an increasingly important asset in the life sciences, helping address data scarcity, privacy concerns, and barriers to data access. Creating artificial datasets that mirror the characteristics of real data allows researchers to develop and validate computational methods in controlled environments. Despite its promise, the adoption of SD in life sciences hinges on rigorous evaluation metrics designed to assess their fidelity and reliability. To explore the current landscape of SD evaluation metrics in distinct life sciences domains, the ELIXIR Machine Learning Focus Group performed a systematic review of the scientific literature following the PRISMA guidelines. Six critical domains were examined to identify current practices for assessing SD. Findings reveal that, while generation methods are rapidly evolving, systematic evaluation is often overlooked, limiting researchers’ ability to compare, validate, and trust synthetic datasets across different domains. This systematic review underscores the urgent need for robust, standardized evaluation approaches that not only bolster confidence in SD but also guide its effective and responsible implementation. By laying the groundwork for establishing domain-specific yet interoperable standards, this scoping review paves the way for future initiatives aimed at enhancing the role of SD in scientific discovery, clinical practice and beyond.
- Book Chapter
1
- 10.1093/oxfordhb/9780199573691.013.003
- May 1, 2014
A question-answering (QA) system is an application program which takes a user’s natural-language input question and attempts to return a precise answer. This chapter describes the current state of the art of the field of QA. We start with an analysis of the space of questions, and discuss which ones typically are suitable for automatic processing by current systems, how they are studied, and the types of QA system that process them. We look at the composition of a typical QA system, some of the recurring linguistic and semantic problems that QA systems must overcome, and a variety of specific approaches that have been developed to address them. We also describe the typical evaluation metrics used to measure the performance of QA systems. An overview of one specific system, IBM’s Watson, is provided at the end.