Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Generative artificial intelligence in optimizing the quality of cancer care: potential, limitations, and future directions of development for Large Language Models. A narrative literature review

  • TL;DR
  • Abstract
  • Literature Map
  • Similar Papers
TL;DR

This review examines the potential of Large Language Models (LLMs) to enhance cancer care by improving patient education, supporting clinical decisions, promoting prevention, and reducing disparities, while highlighting limitations like model hallucination, personalization issues, legal concerns, and ethical dilemmas, and advocating for integration into e-health systems and specialized development.

Abstract
Translate article icon Translate Article Star icon

This literature review included scientific articles, published between 2022 and 2025, indexed in PubMed, Scopus, and Proquest. The article investigates the potential of generative artificial intelligence (GenAI), particularly Large Language Models (LLMs), to improve the quality of cancer care. LLMs have demonstrated effectiveness in patient education by simplifying complex medical terminology and tailoring content to the user’s level of understanding. LLMs also assist physicians in clinical decision-making by analyzing medical data and supporting adherence to the latest guidelines. However, expert oversight is still necessary due to the risk of error. For cancer prevention, LLMs promote healthy lifestyle adoption, participation in screening programs, and vaccination. They also play an important role in reducing inequities in access to information. Another key feature of LLMs is their ability to translate complex diagnostic reports into patient-friendly language. LLMs have also shown promise in counteracting cancer-related misinformation. However, the article identifies certain LLMs limitations, such as model hallucination, incomplete personalization, and unresolved legal liability concerns. The article emphasizes ethical dilemmas, particularly those related to patient autonomy and the risk of dehumanizing care. For future progress, the article emphasizes the need to integrate LLMs into e-health systems and to develop specialized models supported by interdisciplinary teams.

Similar Papers
  • Research Article
  • Cite Count Icon 11
  • 10.1287/ijds.2023.0007
How Can IJDS Authors, Reviewers, and Editors Use (and Misuse) Generative AI?
  • Apr 1, 2023
  • INFORMS Journal on Data Science
  • Galit Shmueli + 7 more

How Can <i>IJDS</i> Authors, Reviewers, and Editors Use (and Misuse) Generative AI?

  • Research Article
  • 10.1007/s10006-026-01514-y
Large language model use in oral and maxillofacial surgery training: a national resident survey.
  • Feb 21, 2026
  • Oral and maxillofacial surgery
  • Nolan Kranc + 7 more

Large language models (LLMs) are advanced artificial intelligence (AI) tools capable of generating human-like text and are increasingly used in education, clinical care, and research. Little is known about their use within oral and maxillofacial surgery (OMFS) training. This study investigates LLM usage trends, perceived value, and educational integration among OMFS residents in the United States. A national, anonymous cross-sectional survey was distributed to OMFS residents via program directors. It gathered demographic data, LLM usage patterns, applications, perceived limitations, and attitudes toward incorporating LLMs into formal education. Eighty-one residents responded, 79.0% (64/81) reported having used an LLM, and of that group, 96.9% (62/64) use ChatGPT. 51.9% (42/81) of respondents used LLMs at least monthly in residency; however, 97.5% (79/81) reported having received no formal LLM education during residency. Residents used LLMs for clinical decision support, board preparation, research, and career planning. Free-text responses revealed a wide spectrum of views. Some advocated for curricular integration and patient education applications, while others questioned the need for formal instruction. Some respondents supported integrating LLMs into curriculums and patient education while others questioned the need for formal instruction. LLMs are used frequently by OMFS residents for a variety of purposes. As AI and LLMs become embedded in healthcare, understanding how OMFS residents interact with LLMs is vital. These findings may guide curriculum development, fostering responsible and effective use of LLMs in surgical training and practice.

  • Research Article
  • Cite Count Icon 184
  • 10.3389/fmed.2024.1477898
Large language models in patient education: a scoping review of applications in medicine.
  • Oct 29, 2024
  • Frontiers in medicine
  • Serhat Aydin + 3 more

Large Language Models (LLMs) are sophisticated algorithms that analyze and generate vast amounts of textual data, mimicking human communication. Notable LLMs include GPT-4o by Open AI, Claude 3.5 Sonnet by Anthropic, and Gemini by Google. This scoping review aims to synthesize the current applications and potential uses of LLMs in patient education and engagement. Following the PRISMA-ScR checklist and methodologies by Arksey, O'Malley, and Levac, we conducted a scoping review. We searched PubMed in June 2024, using keywords and MeSH terms related to LLMs and patient education. Two authors conducted the initial screening, and discrepancies were resolved by consensus. We employed thematic analysis to address our primary research question. The review identified 201 studies, predominantly from the United States (58.2%). Six themes emerged: generating patient education materials, interpreting medical information, providing lifestyle recommendations, supporting customized medication use, offering perioperative care instructions, and optimizing doctor-patient interaction. LLMs were found to provide accurate responses to patient queries, enhance existing educational materials, and translate medical information into patient-friendly language. However, challenges such as readability, accuracy, and potential biases were noted. LLMs demonstrate significant potential in patient education and engagement by creating accessible educational materials, interpreting complex medical information, and enhancing communication between patients and healthcare providers. Nonetheless, issues related to the accuracy and readability of LLM-generated content, as well as ethical concerns, require further research and development. Future studies should focus on improving LLMs and ensuring content reliability while addressing ethical considerations.

  • Research Article
  • 10.1016/j.jsurg.2026.103884
Testing the Implementation and Acceptance of Generative Artificial Intelligence to Augment Vascular Surgery Journal Club.
  • May 1, 2026
  • Journal of surgical education
  • Rhea Puthumana + 4 more

Testing the Implementation and Acceptance of Generative Artificial Intelligence to Augment Vascular Surgery Journal Club.

  • Research Article
  • Cite Count Icon 3
  • 10.1109/mwc.2025.3600789
AGI and LLM-Driven Spectrum Intelligence in Future Wireless Networks
  • Feb 1, 2026
  • IEEE Wireless Communications
  • Shumaila Javaid + 3 more

Artificial General Intelligence (AGI) and Large Language Models (LLMs) are gaining attention for their transformative potential across various fields. While LLMs have significantly advanced Natural Language Processing (NLP), they face challenges in reasoning, adaptability, and bias. AGI, with its human-like cognitive functions, offers a promising solution by enhancing the flexibility and context-awareness of LLMs. This paper explores the integration of AGI with LLMs to address complex, dynamic problems, focusing on advancements in Cognitive Radio (CR) and Spectrum Intelligence (SI) technologies. Spectrum sensing, a cornerstone of CR and SI, is critical for identifying underutilized frequency bands and mitigating interference. Traditional methods often struggle in dynamic environments due to their reliance on static models. By combining AGI’s adaptive decision-making with LLMs’ context-aware understanding, the integrated system can enhance the accuracy and efficiency of spectrum sensing. This integration enables better processing of diverse data, prediction of spectrum usage, and dynamic adaptation to changing conditions, paving the way for intelligent spectrum management. As the demand for efficient communication grows with the proliferation of connected devices, AGI-augmented LLMs offer scalable, context-aware solutions to modern communication challenges. AGI with LLMs has the potential to transform spectrum sensing and management into a more adaptive, efficient paradigm, ensuring the performance of next-generation wireless networks.

  • Research Article
  • Cite Count Icon 2
  • 10.1016/j.jtha.2025.09.004
Large language models vs thrombosis experts: a comparative study on patient education and clinical decision-making in venous thromboembolism.
  • Mar 1, 2026
  • Journal of thrombosis and haemostasis : JTH
  • Nikola Vladic + 6 more

Large language models (LLMs) have demonstrated remarkable capabilities in various medical fields, yet their performance in thrombosis and hemostasis, particularly in patient education and complex clinical decision making, is unexplored. We aimed to compare the quality of responses from LLMs vs thrombosis experts and assess clinician's ability to distinguish between them. Three experts on thrombosis and hemostasis and 3 LLMs (Le Chat Pixtral Large, DeepSeek-R1, and ChatGPT-4.5) answered 3 patient education and 3 clinical decision-making queries. Thirty-seven physicians rated responses for adequacy (1 = very poor; 10 = excellent) and estimated their origin (1 = certainly LLM; 10 = certainly human). Mean differences were assessed via t-tests, medians via Wilcoxon tests, and correlations via Spearman test. All P values were adjusted with Bonferroni correction. LLMs provided significantly better patient education responses than experts. Mean adequacy score differences were as follows: Le Chat Pixtral Large +1.6 (95% CI, 1.3-2.0; P < .01), DeepSeek-R1 +1.7 (95% CI, 1.3-2.1; P < .001), and ChatGPT-4.5 +1.9 (95% CI, 1.6-2.3; P < .001). In clinical decision-making, DeepSeek-R1 outperformed experts (+1.4; 95% CI, 1.1-1.8; P < .001), whereas Le Chat Pixtral Large (-0.3; 95% CI, -0.8-0.1; P = .96) and ChatGPT-4.5 (+0.5; 95% CI, 0.0-0.9; P = .18), performed comparably with experts. Evaluators could not distinguish between expert (median, 6.0; IQR, 3.0-8.0) and LLM-generated responses (median, 6.0; IQR, 4.0-8.0). LLMs outperform experts in venous thromboembolism-related patient education and match or exceed them in clinical decision making, providing responses indistinguishable from experts. Although major barriers need to be addressed, LLMs have strong potential to support clinical management and patient education in the field of thrombosis and hemostasis.

  • Research Article
  • Cite Count Icon 10
  • 10.1007/s00345-024-05423-1
The interaction of structured data using openEHR and large Language models for clinical decision support in prostate cancer.
  • Jan 13, 2025
  • World journal of urology
  • Philippe Kaiser + 8 more

Multidisciplinary teams (MDTs) are essential for cancer care but are resource-intensive. Decision-making processes within MDTs, while critical, contribute to increased healthcare costs due to the need for specialist time and coordination. The recent emergence of large language models (LLMs) offers the potential to improve the efficiency and accuracy of clinical decision-making processes, potentially reducing costs associated with traditional MDT models. We conducted a retrospective study of 171 consecutively treated patients with newly diagnosed prostate cancer. Relevant structured clinical data and the European Association of Urology (EAU) pocket guidelines were provided to two LLMs (chatGPT-4, Claude-3-Opus). LLM treatment recommendations were compared to actual treatment recommendations of the MDT meeting (MDM). Both LLMs demonstrated an overall adherence of 93% with the MDT treatment recommendations. Discrepancies between LLM and MDT recommendations were observed in 15 cases (9%), primarily due to lack of clinical information that could be provided to the LLMs. In 5 cases (3%), the LLM recommendations were not in line with EAU guidelines despite having access to all relevant information. Our findings provide evidence that LLMs can provide accurate treatment recommendations for newly diagnosed prostate cancer patients. LLMs have the potential to streamline MDT workflows, enabling specialists to focus on complex cases and patient-centered discussions. In this study, we explored the potential of artificial intelligence models called large language models (LLMs) to assist in treatment decision-making for prostate cancer patients. We found that LLMs, when provided with patient information and clinical guidelines, can recommend treatments that closely match those made by a team of cancer specialists, suggesting that LLMs could help streamline the decision-making process and potentially reduce healthcare costs.

  • Research Article
  • Cite Count Icon 10
  • 10.1007/s41669-025-00580-4
Using Generative Artificial Intelligence in Health Economics and Outcomes Research: A Primer on Techniques and Breakthroughs.
  • Apr 29, 2025
  • PharmacoEconomics - open
  • Tim Reason + 7 more

The emergence of generative artificial intelligence (GenAI) offers the potential to enhance health economics and outcomes research (HEOR) by streamlining traditionally time-consuming and labour-intensive tasks, such as literature reviews, data extraction, and economic modelling. To effectively navigate this evolving landscape, health economists need a foundational understanding of how GenAI can complement their work. This primer aims to introduce health economists to the essentials of using GenAI tools, particularly large language models (LLMs), in HEOR projects. For health economists new to GenAI technologies, chatbot interfaces like ChatGPT offer an accessible way to explore the potential of LLMs. For more complex projects, knowledge of application programming interfaces (APIs), which provide scalability and integration capabilities, and prompt engineering strategies, such as few-shot and chain-of-thought prompting, is necessary to ensure accurate and efficient data analysis, enhance model performance, and tailor outputs to specific HEOR needs. Retrieval-augmented generation (RAG) can further improve LLM performance by incorporating current external information. LLMs have significant potential in many common HEOR tasks, such as summarising medical literature, extracting structured data, drafting report sections, generating statistical code, answering specific questions, and reviewing materials to enhance quality. However, health economists must also be aware of ongoing limitations and challenges, such as the propensity of LLMs to produce inaccurate information ('hallucinate'), security concerns, issues with reproducibility, and the risk of bias. Implementing LLMs in HEOR requires robust security protocols to handle sensitive data in compliance with the European Union's General Data Protection Regulation (GDPR) and the United States' Health Insurance Portability and Accountability Act (HIPAA). Deployment options such as local hosting, secure API use, or cloud-hosted open-source models offer varying levels of control and cost, each with unique trade-offs in security, accessibility, and technical demands. Reproducibility and transparency also pose unique challenges. To ensure the credibility of LLM-generated content, explicit declarations of the model version, prompting techniques, and benchmarks against established standards are recommended. Given the 'black box' nature of LLMs, a clear reporting structure is essential to maintain transparency and validate outputs, enabling stakeholders to assess the reliability and accuracy of LLM-generated HEOR analyses. The ethical implications of using artificial intelligence (AI) in HEOR, including LLMs, are complex and multifaceted, requiring careful assessment of each use case to determine the necessary level of ethical scrutiny and transparency. Health economists must balance the potential benefits of AI adoption against the risks of maintaining current practices, while also considering issues such as accountability, bias, intellectual property, and the broader impact on the healthcare system. As LLMs and AI technologies advance, their potential role in HEOR will become increasingly evident. Key areas of promise include creating dynamic, continuously updated HEOR materials, providing patients with more accessible information, and enhancing analytics for faster access to medicines. To maximise these benefits, health economists must understand and address challenges such as data ownership and bias. The coming years will be critical for establishing best practices for GenAI in HEOR. This primer encourages health economists to adopt GenAI responsibly, balancing innovation with scientific rigor and ethical integrity to improve healthcare insights and decision-making.

  • Research Article
  • Cite Count Icon 4
  • 10.1016/j.cjco.2025.02.012
A Primer on Large Language Models (LLMs) and ChatGPT for Cardiovascular Healthcare Professionals.
  • May 1, 2025
  • CJC open
  • Muneeb Ahmed + 3 more

A Primer on Large Language Models (LLMs) and ChatGPT for Cardiovascular Healthcare Professionals.

  • Research Article
  • 10.1177/17562872261441968
Generative AI in urology: rethinking patient counselling and shared decision-making - a scoping review from the European Association of Urology Patient Office.
  • Mar 1, 2026
  • Therapeutic advances in urology
  • Clara Cerrato + 4 more

Shared decision-making (SDM) in urology faces challenges including limited health literacy, language barriers, and time constraints that can compromise informed consent and treatment adherence. Generative artificial intelligence (GAI), particularly large language models, offers opportunities to personalise patient education and enhance SDM. To evaluate the role of GAI applications in SDM for patients with urological conditions. Peer-reviewed observational studies, validation studies, or mixed-methods studies evaluating GAI (e.g., large language models, AI chatbots) in patient communication, education, counselling, or SDM for urological conditions were included. Editorials, opinion pieces, conference abstracts, and non-English language publications were excluded. PubMed, Embase, Cochrane Library, and Web of Science databases were comprehensively searched through June 2025. Study assessments: Newcastle-Ottawa Scale, the STROBE or the AGREE II as per study type. Charting methods was performed by using a standardised form. Outcomes of interest included accuracy of GAI-generated information, patient understanding, satisfaction, and decisional conflict. Following the Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews guidelines, 18 observational studies (2023-2025) were included, comprising 310 patients in real-world settings plus hundreds of simulated queries across diverse urological conditions. GAI demonstrated moderate to high accuracy (52%-95%) for guideline-based information, with optimal performance in disease-specific patient education. A prospective comparative study showed 27% reduction in consultation time and improved patient understanding with ChatGPT-4 assistance. Limitations emerged including poor performance in emergencies and complex oncological counselling, and readability issues with content written at a college level (mean Flesch-Kincaid Grade Level 13.5). Most studies evaluated ChatGPT versions, limiting generalizability. GAI could enhance and potentially transform SDM in urology with appropriate clinical oversight and human-in-the-loop governance. Currently, GAI is useful for consultation preparation and patient education, while maintaining physician expertise for complex scenarios. Future implementation should prioritise patient safety, equitable access, and environmental sustainability while developing speciality-specific models and clinician education programmes.

  • Research Article
  • 10.30574/wjarr.2025.26.1.1086
From large language models to artificial general intelligence: Evolution pathways in clinical healthcare
  • Apr 30, 2025
  • World Journal of Advanced Research and Reviews
  • Indraneel Borgohain

This article examines the trajectory and challenges of evolving current large language models (LLMs) toward artificial general intelligence (AGI) capabilities within clinical healthcare environments. The text analyzes the gaps between contemporary LLMs' pattern recognition abilities and the robust reasoning, causal understanding, and contextual adaptation required for true medical AGI. Through a systematic review of current clinical applications and limitations of LLMs, the article identifies three critical areas requiring advancement: dynamic integration of multi-modal medical data streams, consistent medical reasoning across novel scenarios, and autonomous learning from clinical interactions while maintaining safety constraints. A novel architectural framework is proposed that combines LLM capabilities with symbolic reasoning, causal inference, and continual learning mechanisms specifically designed for clinical environments. The article suggests that while LLMs provide a promising foundation, achieving AGI in clinical systems requires fundamental breakthroughs in areas including knowledge representation, uncertainty quantification, and ethical decision-making. The article concludes by outlining a roadmap for research priorities and safety considerations essential for progressing toward clinical AGI while maintaining patient safety and care quality.

  • Research Article
  • Cite Count Icon 206
  • 10.1186/s41073-023-00133-5
Fighting reviewer fatigue or amplifying bias? Considerations and recommendations for use of ChatGPT and other large language models in scholarly peer review
  • May 18, 2023
  • Research Integrity and Peer Review
  • Mohammad Hosseini + 1 more

BackgroundThe emergence of systems based on large language models (LLMs) such as OpenAI’s ChatGPT has created a range of discussions in scholarly circles. Since LLMs generate grammatically correct and mostly relevant (yet sometimes outright wrong, irrelevant or biased) outputs in response to provided prompts, using them in various writing tasks including writing peer review reports could result in improved productivity. Given the significance of peer reviews in the existing scholarly publication landscape, exploring challenges and opportunities of using LLMs in peer review seems urgent. After the generation of the first scholarly outputs with LLMs, we anticipate that peer review reports too would be generated with the help of these systems. However, there are currently no guidelines on how these systems should be used in review tasks.MethodsTo investigate the potential impact of using LLMs on the peer review process, we used five core themes within discussions about peer review suggested by Tennant and Ross-Hellauer. These include 1) reviewers’ role, 2) editors’ role, 3) functions and quality of peer reviews, 4) reproducibility, and 5) the social and epistemic functions of peer reviews. We provide a small-scale exploration of ChatGPT’s performance regarding identified issues.ResultsLLMs have the potential to substantially alter the role of both peer reviewers and editors. Through supporting both actors in efficiently writing constructive reports or decision letters, LLMs can facilitate higher quality review and address issues of review shortage. However, the fundamental opacity of LLMs’ training data, inner workings, data handling, and development processes raise concerns about potential biases, confidentiality and the reproducibility of review reports. Additionally, as editorial work has a prominent function in defining and shaping epistemic communities, as well as negotiating normative frameworks within such communities, partly outsourcing this work to LLMs might have unforeseen consequences for social and epistemic relations within academia. Regarding performance, we identified major enhancements in a short period and expect LLMs to continue developing.ConclusionsWe believe that LLMs are likely to have a profound impact on academia and scholarly communication. While potentially beneficial to the scholarly communication system, many uncertainties remain and their use is not without risks. In particular, concerns about the amplification of existing biases and inequalities in access to appropriate infrastructure warrant further attention. For the moment, we recommend that if LLMs are used to write scholarly reviews and decision letters, reviewers and editors should disclose their use and accept full responsibility for data security and confidentiality, and their reports’ accuracy, tone, reasoning and originality.

  • Research Article
  • Cite Count Icon 16
  • 10.1016/j.jaci.2025.02.004
Evaluating large language model performance to support the diagnosis and management of patients with primary immune disorders.
  • Jul 1, 2025
  • The Journal of allergy and clinical immunology
  • Nicholas L Rider + 8 more

Evaluating large language model performance to support the diagnosis and management of patients with primary immune disorders.

  • Research Article
  • Cite Count Icon 1
  • 10.2196/75452
GPT-4o and OpenAI o1 Performance on the 2024 Spanish Competitive Medical Specialty Access Examination: Cross-Sectional Quantitative Evaluation Study
  • Jan 12, 2026
  • JMIR Medical Education
  • Pau Benito + 9 more

BackgroundIn recent years, generative artificial intelligence and large language models (LLMs) have rapidly advanced, offering significant potential to transform medical education. Several studies have evaluated the performance of chatbots on multiple-choice medical examinations.ObjectiveThe study aims to assess the performance of two LLMs—GPT-4o and OpenAI o1—on the Médico Interno Residente (MIR) 2024 examination, the Spanish national medical test that determines eligibility for competitive medical specialist training positions.MethodsA total of 176 questions from the MIR 2024 examination were analyzed. Each question was presented individually to the chatbots to ensure independence and prevent memory retention bias. No additional prompts were introduced to minimize potential bias. For each LLM, response consistency under verification prompting was assessed by systematically asking, “Are you sure?” after each response. Accuracy was defined as the percentage of correct responses compared to the official answers provided by the Spanish Ministry of Health. It was assessed for GPT-4o, OpenAI o1, and, as a benchmark, for a consensus of medical specialists and for the average MIR candidate. Subanalyses included performance across different medical subjects, question difficulty (quintiles based on the percentage of examinees correctly answering each question), and question types (clinical cases vs theoretical questions; positive vs negative questions).ResultsOverall accuracy was 89.8% (158/176) for GPT-4o and 90% (160/176) after verification prompting, 92.6% (163/176) for OpenAI o1 and 93.2% (164/176) after verification prompting, 94.3% (166/176) for the consensus of medical specialists, and 56.6% (100/176) for the average MIR candidate. Both LLMs and the consensus of medical specialists outperformed the average MIR candidate across all 20 medical subjects analyzed, with ≥80% LLMs’ accuracy in most domains. A performance gradient was observed: LLMs’ accuracy gradually declined as question difficulty increased. Slightly higher accuracy was observed for clinical cases compared to theoretical questions, as well as for positive questions compared to negative ones. Both models demonstrated high response consistency, with near-perfect agreement between initial responses and those after the verification prompting.ConclusionsThese findings highlight the excellent performance of GPT-4o and OpenAI o1 on the MIR 2024 examination, demonstrating consistent accuracy across medical subjects and question types. The integration of LLMs into medical education presents promising opportunities and is likely to reshape how students prepare for licensing examinations and change our understanding of medical education. Further research should explore how the wording, language, prompting techniques, and image-based questions can influence LLMs’ accuracy, as well as evaluate the performance of emerging artificial intelligence models in similar assessments.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 119
  • 10.2196/56764
Large Language Models and User Trust: Consequence of Self-Referential Learning Loop and the Deskilling of Health Care Professionals
  • Apr 25, 2024
  • Journal of Medical Internet Research
  • Avishek Choudhury + 1 more

As the health care industry increasingly embraces large language models (LLMs), understanding the consequence of this integration becomes crucial for maximizing benefits while mitigating potential pitfalls. This paper explores the evolving relationship among clinician trust in LLMs, the transition of data sources from predominantly human-generated to artificial intelligence (AI)–generated content, and the subsequent impact on the performance of LLMs and clinician competence. One of the primary concerns identified in this paper is the LLMs’ self-referential learning loops, where AI-generated content feeds into the learning algorithms, threatening the diversity of the data pool, potentially entrenching biases, and reducing the efficacy of LLMs. While theoretical at this stage, this feedback loop poses a significant challenge as the integration of LLMs in health care deepens, emphasizing the need for proactive dialogue and strategic measures to ensure the safe and effective use of LLM technology. Another key takeaway from our investigation is the role of user expertise and the necessity for a discerning approach to trusting and validating LLM outputs. The paper highlights how expert users, particularly clinicians, can leverage LLMs to enhance productivity by off-loading routine tasks while maintaining a critical oversight to identify and correct potential inaccuracies in AI-generated content. This balance of trust and skepticism is vital for ensuring that LLMs augment rather than undermine the quality of patient care. We also discuss the risks associated with the deskilling of health care professionals. Frequent reliance on LLMs for critical tasks could result in a decline in health care providers’ diagnostic and thinking skills, particularly affecting the training and development of future professionals. The legal and ethical considerations surrounding the deployment of LLMs in health care are also examined. We discuss the medicolegal challenges, including liability in cases of erroneous diagnoses or treatment advice generated by LLMs. The paper references recent legislative efforts, such as The Algorithmic Accountability Act of 2023, as crucial steps toward establishing a framework for the ethical and responsible use of AI-based technologies in health care. In conclusion, this paper advocates for a strategic approach to integrating LLMs into health care. By emphasizing the importance of maintaining clinician expertise, fostering critical engagement with LLM outputs, and navigating the legal and ethical landscape, we can ensure that LLMs serve as valuable tools in enhancing patient care and supporting health care professionals. This approach addresses the immediate challenges posed by integrating LLMs and sets a foundation for their maintainable and responsible use in the future.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant