Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Retrieval-augmented Generation (RAG): What is There for Data Management Researchers? A discussion on research from a panel at LLM+Vector Data Workshop @ IEEE ICDE 2025

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

Large language models (LLMs) enable the state-ofthe- art in language processing by framing diverse tasks- from code synthesis and healthcare to finance, digital assistance, and scientific discovery-as next-token prediction problems [38, 53, 60, 65, 72, 20, 68, 76, 32]. In addition, LLMs enable automation in data science and engineering, optimizing processes such as data analysis, manipulation, querying, interpretation, research, and education [33, 7, 8, 22, 24, 42, 77, 37, 43, 44, 66, 49]. LLMs encode probabilistic token patterns instead of maintaining explicit knowledge structures, which (1) constrains multi-step reasoning under the next-token prediction paradigm; (2) ties outputs to static, pre-cutoff training data-undermining performance on evolving knowledge tasks; and (3) lacks a built-in factual verification mechanism, resulting in hallucinations [25].

Similar Papers
  • Research Article
  • Cite Count Icon 11
  • 10.1287/ijds.2023.0007
How Can IJDS Authors, Reviewers, and Editors Use (and Misuse) Generative AI?
  • Apr 1, 2023
  • INFORMS Journal on Data Science
  • Galit Shmueli + 7 more

How Can <i>IJDS</i> Authors, Reviewers, and Editors Use (and Misuse) Generative AI?

  • Research Article
  • 10.1093/ndt/gfae069.792
#2924 Comparison of large language models and traditional natural language processing techniques in predicting arteriovenous fistula failure
  • May 23, 2024
  • Nephrology Dialysis Transplantation
  • Suman Lama + 6 more

Background and Aims Large language models (LLMs) have gained significant attention in the field of natural language processing (NLP), marking a shift from traditional techniques like Term Frequency-Inverse Document Frequency (TF-IDF). We developed a traditional NLP model to predict arteriovenous fistula (AVF) failure within next 30 days using clinical notes. The goal of this analysis was to investigate whether LLMs would outperform traditional NLP techniques, specifically in the context of predicting AVF failure within the next 30 days using clinical notes. Method We defined AVF failure as the change in status from active to permanently unusable status or temporarily unusable status. We used data from a large kidney care network from January 2021 to December 2021. Two models were created using LLMs and traditional TF-IDF technique. We used “distilbert-base-uncased”, a distilled version of BERT base model [1], and compared its performance with traditional TF-IDF-based NLP techniques. The dataset was randomly divided into 60% training, 20% validation and 20% test dataset. The test data, comprising of unseen patients’ data was used to evaluate the performance of the model. Both models were evaluated using metrics such as area under the receiver operating curve (AUROC), accuracy, sensitivity, and specificity. Results The incidence of 30 days AVF failure rate was 2.3% in the population. Both LLMs and traditional showed similar overall performance as summarized in Table 1. Notably, LLMs showed marginally better performance in certain evaluation metrics. Both models had same AUROC of 0.64 on test data. The accuracy and balanced accuracy for LLMs were 72.9% and 59.7%, respectively, compared to 70.9% and 59.6% for the traditional TF-IDF approach. In terms of specificity, LLMs scored 73.2%, slightly higher than the 71.2% observed for traditional NLP methods. However, LLMs had a lower sensitivity of 46.1% compared to 48% for traditional NLP. However, it is worth noting that training on LLMs took considerably longer than TF-IDF. Moreover, it also used higher computational resources such as utilization of graphics processing units (GPU) instances in cloud-based services, leading to higher cost. Conclusion In our study, we discovered that advanced LLMs perform comparably to traditional TF-IDF modeling techniques in predicting the failure of AVF. Both models demonstrated identical AUROC. While specificity was higher in LLMs compared to traditional NLP, sensitivity was higher in traditional NLP compared to LLMs. LLM was fine-tuned with a limited dataset, which could have influenced its performance to be similar to that of traditional NLP methods. This finding suggests that while LLMs may excel in certain scenarios, such as performing in-depth sentiment analysis of patient data for complex tasks, their effectiveness is highly dependent on the specific use case. It is crucial to weigh the benefits against the resources required for LLMs, as they can be significantly more resource-intensive and costly compared to traditional TF-IDF methods. This highlights the importance of a use-case-driven approach in selecting the appropriate NLP technique for healthcare applications.

  • Research Article
  • Cite Count Icon 27
  • 10.1162/daed_e_01897
Getting AI Right: Introductory Notes on AI &amp; Society
  • May 1, 2022
  • Daedalus
  • James Manyika

This dialogue is from an early scene in the 2014 film Ex Machina, in which Nathan has invited Caleb to determine whether Nathan has succeeded in creating artificial intelligence.1 The achievement of powerful artificial general intelligence has long held a grip on our imagination not only for its exciting as well as worrisome possibilities, but also for its suggestion of a new, uncharted era for humanity. In opening his 2021 BBC Reith Lectures, titled "Living with Artificial Intelligence," Stuart Russell states that "the eventual emergence of general-purpose artificial intelligence [will be] the biggest event in human history."2Over the last decade, a rapid succession of impressive results has brought wider public attention to the possibilities of powerful artificial intelligence. In machine vision, researchers demonstrated systems that could recognize objects as well as, if not better than, humans in some situations. Then came the games. Complex games of strategy have long been associated with superior intelligence, and so when AI systems beat the best human players at chess, Atari games, Go, shogi, StarCraft, and Dota, the world took notice. It was not just that Als beat humans (although that was astounding when it first happened), but the escalating progression of how they did it: initially by learning from expert human play, then from self-play, then by teaching themselves the principles of the games from the ground up, eventually yielding single systems that could learn, play, and win at several structurally different games, hinting at the possibility of generally intelligent systems.3Speech recognition and natural language processing have also seen rapid and headline-grabbing advances. Most impressive has been the emergence recently of large language models capable of generating human-like outputs. Progress in language is of particular significance given the role language has always played in human notions of intelligence, reasoning, and understanding. While the advances mentioned thus far may seem abstract, those in driverless cars and robots have been more tangible given their embodied and often biomorphic forms. Demonstrations of such embodied systems exhibiting increasingly complex and autonomous behaviors in our physical world have captured public attention.Also in the headlines have been results in various branches of science in which AI and its related techniques have been used as tools to advance research from materials and environmental sciences to high energy physics and astronomy.4 A few highlights, such as the spectacular results on the fifty-year-old protein-folding problem by AlphaFold, suggest the possibility that AI could soon help tackle science's hardest problems, such as in health and the life sciences.5While the headlines tend to feature results and demonstrations of a future to come, AI and its associated technologies are already here and pervade our daily lives more than many realize. Examples include recommendation systems, search, language translators - now covering more than one hundred languages - facial recognition, speech to text (and back), digital assistants, chatbots for customer service, fraud detection, decision support systems, energy management systems, and tools for scientific research, to name a few. In all these examples and others, AI-related techniques have become components of other software and hardware systems as methods for learning from and incorporating messy real-world inputs into inferences, predictions, and, in some cases, actions. As director of the Future of Humanity Institute at the University of Oxford, Nick Bostrom noted back in 2006, "A lot of cutting-edge AI has filtered into general applications, often without being called AI because once something becomes useful enough and common enough it's not labeled AI anymore."6As the scope, use, and usefulness of these systems have grown for individual users, researchers in various fields, companies and other types of organizations, and governments, so too have concerns when the systems have not worked well (such as bias in facial recognition systems), or have been misused (as in deepfakes), or have resulted in harms to some (in predicting crime, for example), or have been associated with accidents (such as fatalities from self-driving cars).7Dædalus last devoted a volume to the topic of artificial intelligence in 1988, with contributions from several of the founders of the field, among others. Much of that issue was concerned with questions of whether research in AI was making progress, of whether AI was at a turning point, and of its foundations, mathematical, technical, and philosophical-with much disagreement. However, in that volume there was also a recognition, or perhaps a rediscovery, of an alternative path toward AI - the connectionist learning approach and the notion of neural nets-and a burgeoning optimism for this approach's potential. Since the 1960s, the learning approach had been relegated to the fringes in favor of the symbolic formalism for representing the world, our knowledge of it, and how machines can reason about it. Yet no essay captured some of the mood at the time better than Hilary Putnam's "Much Ado About Not Very Much." Putnam questioned the Dædalus issue itself: "Why a whole issue of Dædalus? Why don't we wait until AI achieves something and then have an issue?" He concluded:This volume of Dædalus is indeed the first since 1988 to be devoted to artificial intelligence. This volume does not rehash the same debates; much else has happened since, mostly as a result of the success of the machine learning approach that was being rediscovered and reimagined, as discussed in the 1988 volume. This issue aims to capture where we are in AI's development and how its growing uses impact society. The themes and concerns herein are colored by my own involvement with AI. Besides the television, films, and books that I grew up with, my interest in AI began in earnest in 1989 when, as an undergraduate at the University of Zimbabwe, I undertook a research project to model and train a neural network.9 I went on to do research on AI and robotics at Oxford. Over the years, I have been involved with researchers in academia and labs developing AI systems, studying AI's impact on the economy, tracking AI's progress, and working with others in business, policy, and labor grappling with its opportunities and challenges for society.10The authors of the twenty-five essays in this volume range from AI scientists and technologists at the frontier of many of AI's developments to social scientists at the forefront of analyzing AI's impacts on society. The volume is organized into ten sections. Half of the sections are focused on AI's development, the other half on its intersections with various aspects of society. In addition to the diversity in their topics, expertise, and vantage points, the authors bring a range of views on the possibilities, benefits, and concerns for society. I am grateful to the authors for accepting my invitation to write these essays.Before proceeding further, it may be useful to say what we mean by artificial intelligence. The headlines and increasing pervasiveness of AI and its associated technologies have led to some conflation and confusion about what exactly counts as AI. This has not been helped by the current trend-among researchers in science and the humanities, startups, established companies, and even governments-to associate anything involving not only machine learning, but data science, algorithms, robots, and automation of all sorts with AI. This could simply reflect the hype now associated with AI, but it could also be an acknowledgment of the success of the current wave of AI and its related techniques and their wide-ranging use and usefulness. I think both are true; but it has not always been like this. In the period now referred to as the AI winter, during which progress in AI did not live up to expectations, there was a reticence to associate most of what we now call AI with AI.Two types of definitions are typically given for AI. The first are those that suggest that it is the ability to artificially do what intelligent beings, usually human, can do. For example, artificial intelligence is:The human abilities invoked in such definitions include visual perception, speech recognition, the capacity to reason, solve problems, discover meaning, generalize, and learn from experience. Definitions of this type are considered by some to be limiting in their human-centricity as to what counts as intelligence and in the benchmarks for success they set for the development of AI (more on this later). The second type of definitions try to be free of human-centricity and define an intelligent agent or system, whatever its origin, makeup, or method, as:This type of definition also suggests the pursuit of goals, which could be given to the system, self-generated, or learned.13 That both types of definitions are employed throughout this volume yields insights of its own.These definitional distinctions notwithstanding, the term AI, much to the chagrin of some in the field, has come to be what cognitive and computer scientist Marvin Minsky called a "suitcase word."14 It is packed variously, depending on who you ask, with approaches for achieving intelligence, including those based on logic, probability, information and control theory, neural networks, and various other learning, inference, and planning methods, as well as their instantiations in software, hardware, and, in the case of embodied intelligence, systems that can perceive, move, and manipulate objects.Three questions cut through the discussions in this volume: 1) Where are we in AI's development? 2) What opportunities and challenges does AI pose for society? 3) How much about AI is really about us?Notions of intelligent machines date all the way back to antiquity.15 Philosophers, too, among them Hobbes, Leibnitz, and Descartes, have been dreaming about AI for a long time; Daniel Dennett suggests that Descartes may have even anticipated the Turing Test.16 The idea of computation-based machine intelligence traces to Alan Turing's invention of the universal Turing machine in the 1930s, and to the ideas of several of his contemporaries in the mid-twentieth century. But the birth of artificial intelligence as we know it and the use of the term is generally attributed to the now famed Dartmouth summer workshop of 1956. The workshop was the result of a proposal for a two-month summer project by John McCarthy, Marvin Minsky, Nathaniel Rochester, and Claude Shannon whereby "An attempt will be made to find how to make machines use language, form abstractions and concepts, solve kinds of problems now reserved for humans, and improve themselves."17In their respective contributions to this volume, "From So Simple a Beginning: Species of Artificial Intelligence" and "If We Succeed," and in different but complementary ways, Nigel Shadbolt and Stuart Russell chart the key ideas and developments in AI, its periods of excitement as well as the aforementioned AI winters. The current AI spring has been underway since the 1990s, with headline-grabbing breakthroughs appearing in rapid succession over the last ten years or so: a period that Jeffrey Dean describes in the title of his essay as a "golden decade," not only for the pace of AI development but also its use in a wide range of sectors of society, as well as areas of scientific research.18 This period is best characterized by the approach to achieve artificial intelligence through learning from experience, and by the success of neural networks, deep learning, and reinforcement learning, together with methods from probability theory, as ways for machines to learn.19A brief history may be useful here: In the 1950s, there were two dominant visions of how to achieve machine intelligence. One vision was to use computers to create a logic and symbolic representation of the world and our knowledge of it and, from there, create systems that could reason about the world, thus exhibiting intelligence akin to the mind. This vision was most espoused by Allen Newell and Hebert Simon, along with Marvin Minsky and others. Closely associated with it was the "heuristic search" approach that supposed intelligence was essentially a problem of exploring a space of possibilities for answers. The second vision was inspired by the brain, rather than the mind, and sought to achieve intelligence by learning. In what became known as the connectionist approach, units called perceptrons were connected in ways inspired by the connection of neurons in the brain. At the time, this approach was most associated with Frank Rosenblatt. While there was initial excitement about both visions, the first came to dominate, and did so for decades, with some successes, including so-called expert systems.Not only did this approach benefit from championing by its advocates and plentiful funding, it came with the suggested weight of a long intellectual tradition-exemplified by Descartes, Boole, Frege, Russell, and Church, among others-that sought to manipulate symbols and to formalize and axiomatize knowledge and reasoning. It was only in the late 1980s that interest began to grow again in the second vision, largely through the work of David Rumelhart, Geoffrey Hinton, James McClelland, and others. The history of these two visions and the associated philosophical ideas are discussed in Hubert Dreyfus and Stuart Dreyfus's 1988 Dædalus essay "Making a Mind Versus Modeling the Brain: Artificial Intelligence Back at a Branchpoint."20 Since then, the approach to intelligence based on learning, the use of statistical methods, back-propagation, and training (supervised and unsupervised) has come to characterize the current dominant approach.Kevin Scott, in his essay "I Do Not Think It Means What You Think It Means: Artificial Intelligence, Cognitive Work & Scale," reminds us of the work of Ray Solomonoff and others linking information and probability theory with the idea of machines that can not only learn, but compress and potentially generalize what they learn, and the emerging realization of this in the systems now being built and those to come. The success of the machine learning approach has benefited from the boon in the availability of data to train the algorithms thanks to the growth in the use of the Internet and other applications and services. In research, the data explosion has been the result of new scientific instruments and observation platforms and data-generating breakthroughs, for example, in astronomy and in genomics. Equally important has been the co-evolution of the software and hardware used, especially chip architectures better suited to the parallel computations involved in data- and compute-intensive neural networks and other machine learning approaches, as Dean discusses.Several authors delve into progress in key subfields of AI.21 In their essay, "Searching for Computer Vision North Stars," Fei-Fei Li and Ranjay Krishna chart developments in machine vision and the creation of standard data sets such as ImageNet that could be used for benchmarking performance. In their respective essays "Human Language Understanding & Reasoning" and "The Curious Case of Commonsense Intelligence," Chris Manning and Yejin Choi discuss different eras and ideas in natural language processing, including the recent emergence of large language models comprising hundreds of billions of parameters and that use transformer architectures and self-supervised learning on vast amounts of data.22 The resulting pretrained models are impressive in their capacity to take natural language prompts for which they have not been trained specifically and generate human-like outputs, not only in natural language, but also images, software code, and more, as Mira Murati discusses and illustrates in "Language & Coding Creativity." Some have started to refer to these large language models as foundational models in that once they are trained, they are adaptable to a wide range of tasks and outputs.23 But despite their unexpected performance, these large language models are still early in their development and have many shortcomings and limitations that are highlighted in this volume and elsewhere, including by some of their developers.24In "The Machines from Our Future," Daniela Rus discusses the progress in robotic systems, including advances in the underlying technologies, as well as in their integrated design that enables them to operate in the physical world. She highlights the limitations in the "industrial" approaches used thus far and suggests new ways of conceptualizing robots that draw on insights from biological systems. In robotics, as in AI more generally, there has always been a tension as to whether to copy or simply draw inspiration from how humans and other biological organisms achieve intelligent behavior. Elsewhere, AI researcher Demis Hassabis and colleagues have explored how neuroscience and AI learn from and inspire each other, although so far more in one than the other, as and have the success of the current approaches to AI, there are still many shortcomings and as well as problems in It is useful to on one such as when AI does not as or or or that can to or when it on or information about the world, or when it has such as of all of which can to a of public shortcomings have captured the attention of the wider public and as well as among there is an on AI and In recent years, there has been a of to principles and approaches to AI, as well as involving and such as the on AI, that to best important has been the of with to and - in the and developing AI in both and as has been well in recent This is an important in its own but also with to the of the resulting AI and, in its intersections with more the other there are limitations and problems associated with the that AI is not capable of if could to more more or more general AI. In their Turing deep learning and Geoffrey took of where deep learning and highlighted its current such as the with In the case of natural language processing, Manning and Choi the challenges in and despite the of large language Elsewhere, and have the notion that large language models do anything learning, or In & of in a and discuss the problems in systems, the as how to reason about other their systems, and well as challenges in both and especially when the include both humans and Elsewhere, and others a useful of the problems in there is a growing among many that we do not have for the of AI systems, especially as they become more capable and the of use although AI and its related techniques are to be powerful tools for research in science, as examples in this volume and recent examples in which AI not only help results but also by design and become what some have AI to science and and to and challenges for the possibility that more powerful AI could to new in science, as well as progress in some of challenges and has long been a key for many at the frontier of AI research to more capable the of each of AI, the of more general problems that to the possibility of more capable AI learning, reasoning, of and and of these and other problems that could to more capable systems the of whether current characterized by deep learning, the of and and more foundational and and reinforcement or whether different approaches are in such as cognitive agent approaches or or based on logic and probability theory, to name a few. whether and what of approaches be the AI is but many the current along with of and learning architectures have to their about the of the current approaches is associated with the of whether artificial general intelligence can be and if how and Artificial general intelligence is in to what is called that AI and for tasks and goals, such as The development of on the other aims for more powerful AI - at as powerful as is generally to problem or and, in some the capacity to and improve as well as set and its own and the of and when will be is a for most that its achievement have and as is often in and such as A through and The to Ex and it is or there is growing among many at the frontier of AI research that we for the possibility of powerful with to and and with humans, its and use, and the possibility that of could and that we these into how we approach the development of of the research and development, and in AI is of the AI and in its what Nigel Shadbolt the of AI. This is given the for useful and applications and the for in sectors of the However, a few have made the development of their the most of these are and each of which has demonstrated results of increasing still a long way from the most discussed impact of AI and automation is on and the future of This is not In in the of the excitement about AI and and concerns about their impact on a on and the was that such technologies were important for growth and and "the that but not Most recent of this including those I have been involved have and that over time, more are than are that it is the and the and the of will the In their essay AI & and John discuss these for work and further, in & the of & to discuss the with to and and as well as the opportunities that are especially in developing In "The Turing The & of Artificial Intelligence," discusses how the use of human benchmarks in the development of AI the of AI that rather than human He that the AI's development will take in this and resulting for will on the for companies, and a that the that more will be than too much from of the and does not far enough into the future and at what AI will be capable The for AI could from of that in the is and labor and ability to are and and until automation has mostly physical and but that AI will be on more cognitive and tasks based on and, if early examples are even tasks are not of the In other are now in the world machines that that learn and that their ability to do these is to a range of problems they can will be with the range to which the human has been This was and Allen Newell in that this time could be different usually two that new labor will in which will by other humans for their own even when machines may be capable of these as well as or even better than The other is that AI will create so much and all without the for human and the of will be to for when that will the that once the first time since his creation will be with his his to use his from how to the which science and interest will have for to live and and However, most researchers that we are not to a future in which the of will and that until then, there are other and that be in the labor now and in the such as and other and how humans work increasingly capable that and John and discuss in this are not the only of the by AI. Russell a of the potentially from artificial general intelligence, once a of or ten But even we to general-purpose AI, the opportunities for companies and, for the and growth as well as from AI and its related technologies are more than to pursuit and by companies and in the development, and use of AI. At the many the is it is generally that is a in AI, as by its growth in AI research, and as highlighted in several will have for companies and given the of such technologies as discussed by and others the may in the way of approaches to AI and (such as whether they are companies or as and have have the to to in AI. The role of AI in intelligence, systems, autonomous even and other of increasingly In &

  • Research Article
  • Cite Count Icon 23
  • 10.1136/bmjhci-2023-100978
Performance of large language models on advocating the management of meningitis: a comparative qualitative study
  • Feb 1, 2024
  • BMJ Health & Care Informatics
  • Urs Fisch + 3 more

ObjectivesWe aimed to examine the adherence of large language models (LLMs) to bacterial meningitis guidelines using a hypothetical medical case, highlighting their utility and limitations in healthcare.MethodsA simulated clinical scenario...

  • Research Article
  • 10.55041/ijsrem36608
Exploring Vulnerabilities and Threats in Large Language Models: Safeguarding Against Exploitation and Misuse
  • Aug 10, 2024
  • INTERANTIONAL JOURNAL OF SCIENTIFIC RESEARCH IN ENGINEERING AND MANAGEMENT
  • Mr Aarush Varma + 1 more

This research paper delves into the inherent vulnerabilities and potential threats posed by large language models (LLMs), focusing on their implications across diverse applications such as natural language processing and data privacy. The study aims to identify and analyze these risks comprehensively, emphasizing the importance of mitigating strategies to prevent exploitation and misuse in LLM deployments. In recent years, LLMs have revolutionized fields like automated content generation, sentiment analysis, and conversational agents, yet their immense capabilities also raise significant security concerns. Vulnerabilities such as bias amplification, adversarial attacks, and unintended data leakage can undermine trust and compromise user privacy. Through a systematic examination of these challenges, this paper proposes safeguarding measures crucial for responsibly harnessing the potential of LLMs while minimizing associated risks. It underscores the necessity of rigorous security protocols, including robust encryption methods, enhanced authentication mechanisms, and continuous monitoring frameworks. Furthermore, the research discusses regulatory implications and ethical considerations surrounding LLM usage, advocating for transparency, accountability, and stakeholder engagement in policy- making and deployment practices. By synthesizing insights from current literature and real-world case studies, this study provides a comprehensive framework for stakeholders—developers, policymakers, and users—to navigate the complex landscape of LLM security effectively. Ultimately, this research aims to inform future advancements in LLM technology, ensuring its safe and beneficial integration into various domains while mitigating potential risks to individuals and society as a whole. Keywords— Adversarial attacks on LLMs, Bias in LLMs, Data privacy in LLMs, Ethical considerations LLMs, Exploitation of LLMs, Large Language Models (LLMs), Misuse of LLMs, Mitigation strategies for LLMs, Natural Language Processing (NLP), Regulatory frameworks LLMs, Responsible deployment of LLMs, Risks of LLMs, Security implications of LLMs, Threats to LLMs, Vulnerabilities in LLMs.

  • Discussion
  • Cite Count Icon 2
  • 10.1111/ans.18720
Response to: Investigating the impact of innovative AI chatbot on post-pandemic medical education and clinical assistance: a comprehensive analysis.
  • Oct 2, 2023
  • ANZ Journal of Surgery
  • Yi Xie + 4 more

We are grateful for the thoughtful commentary provided by Kleebayoon and Wiwanitkit on our recent article published in the ANZ Journal of Surgery.1 While we welcome this scholarly engagement, we find it imperative to address their concerns in order to elucidate the robustness of our study's design, methodology, and findings. First, the concern regarding the sample size and scope appears to overlook our study's qualitative nature, aimed at understanding the foundational capabilities of large language models (LLMs) in a controlled medical environment. It is important to underscore that our study serves as an exploratory assessment, where a large sample size is not the primary focus. Second, our study deliberately restricts its application to controlled settings as a foundational step, with future work intending to address real-world efficacy, as explicitly stated in our conclusions. Third, although our evaluation criteria focus on readability, reliability, and consistency with clinical guidelines, these are integral components that inherently contribute to clinical accuracy and patient safety, with ethical considerations forming an underlying theme. Fourth, the critique regarding potential bias in the evaluation process seems to underestimate the diversity of expertise among our panel of evaluators, which included three plastic surgeons and two junior doctors, thereby bringing multiple perspectives to the assessment. Fifth, while the study does not compare LLM performance to human expertise, it is not designed to propose LLMs as substitutes for human clinicians but rather as supplementary tools. Lastly, we acknowledge the ethical dimensions surrounding AI deployment in clinical settings and emphasize in our study the need for ongoing human supervision and algorithmic auditing to mitigate risks and biases. While LLM's have existed in the academic and scientific space for some time, it was the introduction of ChatGPT that captured the public's attention and imagination.2, 3 Since then, multiple papers have been published exploring the role of ChatGPT and other LLMs in the academic and clinical setting. Initial inquiries into the utility of LLMs have included topics such as medical education, clinical management, and scientific research.2-4 As interest in LLMs continued to grow, various other models have arisen, each with their advantages and drawbacks when compared with ChatGPT. Current literature demonstrates that LLMs show significant deficits in referencing, a tangible information inaccuracy rate, and are susceptible to bias. We agree with our readers that our study only reinforces these concerns, while the true extent to which LLMs can be developed into safe, ethical and clinically viable tools remains to be seen. However, only by asking the right questions and raising the relevant issue can we further drive research and lead to a greater understanding of this field. Consequently, we believe that our study offers valuable preliminary insights into the role of LLMs in medical education and clinical assistance. We advocate for a responsible integration of AI into clinical practice that adheres to stringent ethical and safety standards, and our paper mentions the need for future studies that scrutinize these ethical concerns in greater detail. Moreover, our study underscores the importance of interdisciplinary collaborations involving clinicians, data scientists, and ethicists to ensure that AI systems are both effective and ethical. Yi Xie: Conceptualization; project administration; writing – original draft; writing – review and editing. Ishith Seth: Conceptualization; investigation; writing – original draft; writing – review and editing. David J. Hunter-Smith: Supervision, writing - original draft; writing - review and editing. Marc A. Seifman: Supervision; writing – original draft; writing – review and editing. Warren M. Rozen: Supervision; writing – original draft; writing – review and editing.

  • Conference Article
  • Cite Count Icon 62
  • 10.1145/3649217.3653584
CS1-LLM: Integrating LLMs into CS1 Instruction
  • Jul 3, 2024
  • Annapurna Vadaparty + 6 more

The recent, widespread availability of Large Language Models (LLMs) like ChatGPT and GitHub Copilot may impact introductory programming courses (CS1) both in terms of what should be taught and how to teach it. Indeed, recent research has shown that LLMs are capable of solving the majority of the assignments and exams we previously used in CS1. In addition, professional software engineers are often using these tools, raising the question of whether we should be training our students in their use as well. This experience report describes a CS1 course at a large research-intensive university that fully embraces the use of LLMs from the beginning of the course. To incorporate the LLMs, the course was intentionally altered to reduce emphasis on syntax and writing code from scratch. Instead, the course now emphasizes skills needed to successfully produce software with an LLM. This includes explaining code, testing code, and decomposing large problems into small functions that are solvable by an LLM. In addition to frequent, formative assessments of these skills, students were given three large, open-ended projects in three separate domains (data science, image processing, and game design) that allowed them to showcase their creativity in topics of their choosing. In an end-of-term survey, students reported that they appreciated learning with the assistance of the LLM and that they interacted with the LLM in a variety of ways when writing code. We provide lessons learned for instructors who may wish to incorporate LLMs into their course.

  • Research Article
  • Cite Count Icon 4
  • 10.2118/0324-0034-jpt
As Hype Fades, LLMs Gaining Acceptance in Upstream as New Age Research and Coding Tool
  • Mar 1, 2024
  • Journal of Petroleum Technology
  • Trent Jacobs

After using ChatGPT-4 for more than a year, Clinton Lott has made familiarity with the large language model (LLM), a requirement for new software programming hires. The director of Houston-based Spotter Tech gave this reasoning: “Taking 2 hours to do something that ChatGPT can do in 15 seconds is unacceptable today.” It’s a line you’re likely to hear from the most ardent champions of LLMs which have simultaneously reignited the world’s interest in and ire for artificial intelligence (AI). And while Spotter Tech, a small hydraulic fracturing data specialist, may be an exception with its hiring practices, it is far from the only upstream company embracing the new technology. Both BP and Shell have recently given thousands of employees access to Microsoft’s new cloud service called Copilot, which at its core is powered by GPT-4, OpenAI’s latest and most capable LLM. Others including ExxonMobil have placed use of LLMs under lockdown while still signaling a long-term interest in their embrace. As chair of SPE’s Data Science &amp; Engineering Analytics Technical Section, Pushpesh Sharma has spent the past several months speaking with peers in the industry about how and where LLMs are being used. The product manager for field automation firm Aspen Technology told JPT that it’s clear from those conversations that a good number of energy firms are interested in tailoring the technology to their needs, albeit many are doing so quietly as it remains early days. “The big focus right now is on what they call fine tuning,” he said. “You take an existing base model that one of the big tech companies have created and make it more of a fit for your business. It’s kind of like taking a graduate student and then letting them work on their PhD with very specific subject matter.” While OpenAI’s ChatGPT has become the popular symbol of LLMs, there are a growing number of alternatives in the mix. In February, Google introduced its latest model called Gemini 1.5 which seeks to rival GPT-4. There are also open source LLMs that oil and gas companies can use as templates for internal projects. Or they may turn to other big tech firms such as NVIDIA and Amazon Web Services (AWS) which are in the business of helping industrials build and train their own private, custom LLMs. SPE joined the trend after inking a Memorandum of Understanding (MoU) in October with AI firm i2k Connect and Aramco Americas. The collaboration is expected to result in a new domain-specific LLM that industry engineers and researchers can use to answer questions with more technical depth than can be done with the current slate of commercial offerings. One enabling aspect of this current phase of experimentation with LLM technology is retrieval-augmented generation (RAG). The method is considered essential for models to provide more-accurate and actionable responses by extending their scope beyond initial training datasets to include more authoritative and even real-time data.

  • Research Article
  • 10.1016/j.jbi.2026.105034
A Study of Large Language Models for Patient Information Extraction: Model Architecture, Fine-Tuning Strategy, and Multi-task Instruction Tuning
  • Mar 27, 2026
  • Journal of biomedical informatics
  • Cheng Peng + 5 more

A Study of Large Language Models for Patient Information Extraction: Model Architecture, Fine-Tuning Strategy, and Multi-task Instruction Tuning

  • Research Article
  • Cite Count Icon 12
  • 10.1016/j.procs.2023.09.086
A Large and Diverse Arabic Corpus for Language Modeling
  • Jan 1, 2023
  • Procedia Computer Science
  • Abbas Raza Ali + 3 more

A Large and Diverse Arabic Corpus for Language Modeling

  • Research Article
  • Cite Count Icon 11
  • 10.12793/tcp.2024.32.e8
Data science through natural language with ChatGPT's Code Interpreter.
  • Jan 1, 2024
  • Translational and clinical pharmacology
  • Sangzin Ahn

Large language models (LLMs) have emerged as a powerful tool for biomedical researchers, demonstrating remarkable capabilities in understanding and generating human-like text. ChatGPT with its Code Interpreter functionality, an LLM connected with the ability to write and execute code, streamlines data analysis workflows by enabling natural language interactions. Using materials from a previously published tutorial, similar analyses can be performed through conversational interactions with the chatbot, covering data loading and exploration, model development and comparison, permutation feature importance, partial dependence plots, and additional analyses and recommendations. The findings highlight the significant potential of LLMs in assisting researchers with data analysis tasks, allowing them to focus on higher-level aspects of their work. However, there are limitations and potential concerns associated with the use of LLMs, such as the importance of critical thinking, privacy, security, and equitable access to these tools. As LLMs continue to improve and integrate with available tools, data science may experience a transformation similar to the shift from manual to automatic transmission in driving. The advancements in LLMs call for considering the future directions of data science and its education, ensuring that the benefits of these powerful tools are utilized with proper human supervision and responsibility.

  • Supplementary Content
  • Cite Count Icon 120
  • 10.2196/22769
Applications and Concerns of ChatGPT and Other Conversational Large Language Models in Health Care: Systematic Review
  • Nov 7, 2024
  • Journal of Medical Internet Research
  • Leyao Wang + 7 more

BackgroundThe launch of ChatGPT (OpenAI) in November 2022 attracted public attention and academic interest to large language models (LLMs), facilitating the emergence of many other innovative LLMs. These LLMs have been applied in various fields, including health care. Numerous studies have since been conducted regarding how to use state-of-the-art LLMs in health-related scenarios.ObjectiveThis review aims to summarize applications of and concerns regarding conversational LLMs in health care and provide an agenda for future research in this field.MethodsWe used PubMed, ACM, and the IEEE digital libraries as primary sources for this review. We followed the guidance of PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) to screen and select peer-reviewed research articles that (1) were related to health care applications and conversational LLMs and (2) were published before September 1, 2023, the date when we started paper collection. We investigated these papers and classified them according to their applications and concerns.ResultsOur search initially identified 820 papers according to targeted keywords, out of which 65 (7.9%) papers met our criteria and were included in the review. The most popular conversational LLM was ChatGPT (60/65, 92% of papers), followed by Bard (Google LLC; 1/65, 2% of papers), LLaMA (Meta; 1/65, 2% of papers), and other LLMs (6/65, 9% papers). These papers were classified into four categories of applications: (1) summarization, (2) medical knowledge inquiry, (3) prediction (eg, diagnosis, treatment recommendation, and drug synergy), and (4) administration (eg, documentation and information collection), and four categories of concerns: (1) reliability (eg, training data quality, accuracy, interpretability, and consistency in responses), (2) bias, (3) privacy, and (4) public acceptability. There were 49 (75%) papers using LLMs for either summarization or medical knowledge inquiry, or both, and there are 58 (89%) papers expressing concerns about either reliability or bias, or both. We found that conversational LLMs exhibited promising results in summarization and providing general medical knowledge to patients with a relatively high accuracy. However, conversational LLMs such as ChatGPT are not always able to provide reliable answers to complex health-related tasks (eg, diagnosis) that require specialized domain expertise. While bias or privacy issues are often noted as concerns, no experiments in our reviewed papers thoughtfully examined how conversational LLMs lead to these issues in health care research.ConclusionsFuture studies should focus on improving the reliability of LLM applications in complex health-related tasks, as well as investigating the mechanisms of how LLM applications bring bias and privacy issues. Considering the vast accessibility of LLMs, legal, social, and technical efforts are all needed to address concerns about LLMs to promote, improve, and regularize the application of LLMs in health care.

  • Research Article
  • Cite Count Icon 1
  • 10.1200/jco.2025.43.16_suppl.11160
Evaluation of large language model (LLM)-based clinical abstraction of electronic health records (EHRs) for non-small cell lung cancer (NSCLC) patients.
  • Jun 1, 2025
  • Journal of Clinical Oncology
  • Kabir Manghnani + 10 more

11160 Background: Abstraction is a critical step for converting clinical data from unstructured EHRs into a structured format suitable for real-world data analyses. Typically this is a manual, labor-intensive activity requiring substantial training. While prior work has shown that abstraction by humans is reliable, advances in LLMs may improve the efficiency of abstraction. We aim to measure the performance of LLMs in abstracting a diverse set of oncology data elements. Methods: Two clinical abstractors independently abstracted unstructured records of 222 advanced or metastatic NSCLC patients (mean: 248 pages per case). A two-stage LLM system balancing cost and comprehensiveness was used to abstract clinical elements for demographics, diagnosis, third-party lab biomarker testing, and first line (1L) treatment. The initial stage extracted 16 documents semantically similar to the abstraction query and input them, along with abstraction rules, into an LLM (GPT-4o). The LLM was instructed to provide both the abstracted field and a completeness assessment of provided context. If the first phase resulted in a low completeness score, the entire patient record was then input into a long-context LLM (Gemini-Pro-1.5) to re-attempt abstraction. Gwet’s agreement coefficient (AC) was the primary measure of agreement between the LLM and each abstractor. Date agreement was calculated within ±30 days. Results: The LLM system yielded abstracted values for 90% of elements where both abstractors provided non-missing values. In these cases, the LLM also demonstrated high agreement with each abstractor (≥0.81 across all categories). Agreement was highest in demographic and diagnosis domains and lower for 1L treatment domain, which require deeper understanding of a patient's temporal journey. For elements where neither abstractor provided values, the LLM sometimes provided outputs (frequency: 4.9% for non-biomarker elements; 38.5% for biomarker elements). These discrepancies were primarily driven by nuances in abstraction rules; the LLM often included Tempus-tested biomarkers, while abstractors were more rigorous in abstracting only third-party biomarker results. Conclusions: LLMs show high completion rates and high agreement with human abstractors across a variety of critical abstraction fields. The use of LLMs may significantly reduce the burden of human abstraction and allow for large-scale curation of oncology records. Challenges in handling nuanced contexts underscore the need for careful refinement and evaluation prior to deployment. Domain LLM agreement with abstractors (AC, min-max) Demographic (birth date, sex, race, smoking status) 0.96-1 Diagnosis (stage, histology, year of diagnosis) 0.92-0.98 Third Party Biomarker (EGFR, ALK, ROS1, PDL1, BRAF, RET, NTRK) 0.87-1 1L Treatment (agents, initiation date) 0.81-0.86

  • Research Article
  • Cite Count Icon 1
  • 10.1200/jco.2025.43.16_suppl.e23161
Aiding data retrieval in clinical trials with large language models: The APOLLO 11 Consortium in advanced lung cancer patients.
  • Jun 1, 2025
  • Journal of Clinical Oncology
  • Federica Corso + 19 more

e23161 Background: Data retrieval is challenging in clinical research and traditional methods for data collection are often time-consuming and may be error-prone. Large Language Models (LLMs) have shown zero-shot capabilities in converting unstructured clinical text into structured data. These technologies could support the retrieval stage of clinical trials by leveraging the information reported in Electronic Health Records (EHRs) without relying any longer on manual curation. APOLLO 11 Consortium (NCT05550961) is a multicentric Italian trial which leverages a federated infrastructure for the analysis of advanced lung cancer patient data across Italy. Methods: We conducted a pilot study using Llama 3.1 8B on 358 Non-Small Cell Lung Cancer patients from the IRCCS Istituto Nazionale dei Tumori, leader of the APOLLO 11 Consortium. Anonymized EHRs have been analyzed within the LLM pipeline for feature extraction by Wiest et al. A combination of zero/few shot prompting techniques both in English and Italian languages was used. We selected smoking, histology, PD-L1 and staging as multiclass variables and bone/brain/liver metastases as binary variables. The ground truth collection involved a first Manual Data Entry (1-MDE) and a final full-revised MDE (2-MDE). The LLM accuracy was calculated only for the comparison LLM vs 2-MDE. In addition, we calculated the percentage of Missing Information (% MI) in 1-MDE, 2-MDE and LLM extraction. Results: Compared to 2-MDE, LLM achieved feature-specific accuracies of 0.78 for PD-L1, 0.85 for BONE METASTASIS, 0.83 for BRAIN METASTASIS, 0.89 for LIVER METASTASIS and 0.96 for TUMOUR STAGING. For smoking and staging, LLM extraction also reduced % MI relative to 1-MDE (Table 1). Only for PD-L1, we further analyzed the 12.8% of MI and found that 91.3% resulted from hallucinations (i.e., PD-L1 was misclassified as missing). Evaluations using English prompts confirmed the pipeline’s adaptability and high tasks accuracy. Conclusions: This study confirms the feasibility of LLMs for data retrieval in clinical trials demonstrating strong performance across diverse clinical features with minimal prompt optimization. LLMs could assist clinicians and data entry personnel in the 1-MDE process, streamlining initial data structuring and saving time. The 2-MDE step can remain as a quality check to address any discrepancies. Further improvements could focus on prompt optimization and integrating human feedback to reduce hallucination rates. Clinical trial information: NCT05550961 . %MI in 1-MDE, 2-MDE and LLM extraction. Accuracy refers only to LLM vs 2-MDE. Histology and metastasis sites were collected only in 2-MDE. NA = not available. Smoking PD-L1 Histology Bone Met Brain Met Liver Met T N M Stage % MI 1-MDE 6.4 8.9 NA NA NA NA 22.5 22.5 23.11 98.3 % MI 2-MDE 6.6 3 0 0 0 0 0 0 0 0 % MI LLM 2.7 12.8 10.3 0 0 0 0 0 0 6.9 % accuracy (LLM vs 2-MDE) 67 78 91 85 83 89 39 52 70 96

  • Research Article
  • 10.1007/s13755-026-00432-3
Exploring the potential of large language models in healthcare: a focus on cardiovascular disease analysis.
  • Dec 1, 2026
  • Health information science and systems
  • Aihua Li + 4 more

With the rapid development of big data and artificial intelligence technologies, large language models (LLMs) are increasingly being applied across multiple fields. In the healthcare domain, efficient utilization of extensive patient data and medical records considerably enhances diagnostic accuracy and enables disease-risk prediction. This study aims to establish a research framework for evaluating the application of assessment models in structured and unstructured data, exploring the potential applications of LLMs in healthcare. This study proposes the HLLM-Potential Framework (Healthcare Large Language Model Potential Evaluation Framework), a comprehensive evaluation framework designed to assess the applicability and performance of LLMs in medical data analysis through comparative experiments with traditional models. Publicly available and standardized cardiovascular datasets were adopted, covering both structured and unstructured data. Existing LLMs were utilized through training and task-specific configuration on medical data to perform disease prediction and health-risk assessment. In addition, various LLMs were systematically compared with traditional machine-learning and deep-learning models to quantify the differences in their predictive performance. Existing LLMs can process structured and unstructured medical data and use them to predict diseases and evaluate health risks. For structured cardiovascular disease prediction tasks focusing on heart failure, all considered LLMs achieved accuracy and recall rates above 80%. Meanwhile, in unstructured cardiovascular data analysis for ECG image classification, the multimodal LLM Janus Pro 7B attained an overall accuracy rate of 85%. Compared with traditional machine learning and deep learning models, the considered LLMs exhibit stronger interactivity and generalization capabilities as well as less dependence on data quality and feature engineering. Overall, LLMs offer notable advantages, namely, high efficiency, high flexibility, and multimodal integration, for use in the healthcare field; thus, LLMs are expected to be widely applied in the medical and healthcare domains. The online version contains supplementary material available at 10.1007/s13755-026-00432-3.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant