A review of multimodal large language models for smart grids management and control
The worldwide effort to reach carbon peak and neutrality objectives alongside energy market expansion has sped up renewable energy integration, like wind and solar power. The shift towards renewable energy integration introduces substantial uncertainties in power system scheduling and control processes, which test the limits of existing theoretical methods. The advanced reasoning and data-processing capabilities of Large Language Models (LLMs), with particular reference to their ability to analyze multimodal data, provide transformative potential for managing and controlling smart grids. This review examines how LLMs can tackle modern power system challenges while confirming their fit with the power sector’s expanding dependency on Artificial Intelligence (AI) technologies. We assess the requirements of modern power systems for such AI-based solutions, while evaluating how LLMs shape grid management and exploring their enabling technologies, such as model architecture and training methods, along with necessary data. Our review investigates how multimodal LLM technology serves different smart grids’ functions, including generation, transmission, distribution, consumption, and equipment management, to exhibit its adaptable nature in strengthening grid resilience and efficiency. • This review explores the role of multimodal Large Language Models (LLMs) in smart grid management, showing how their ability to integrate and process different types of data, including sensor readings, text logs, weather forecasts, and equipment images, can significantly improve decision-making, fault diagnosis, and operational planning in power systems. • The study analyzes the architectural and training aspects of multimodal LLMs, including the use of pretrained modular encoders, efficient fine-tuning methods such as Low-Rank adaptation (LoRA), and specialized loss functions, highlighting how these techniques enable adaptation to the specific needs of smart grid applications without lengthy retraining. • Practical considerations for industrial implementation are examined, covering multimodal data collection and preprocessing, domain-specific knowledge integration, intelligent task decomposition, and system-level integration, illustrating how LLMs can be seamlessly integrated into power system operating environments. • The review highlights the potential of multimodal LLMs to improve the resilience of the power grid, optimize the integration of renewable energy, and support human-machine collaboration, while outlining future research directions, such as domain-specific base models, physics-based architectures, and human-in-the-loop feedback, in order to further improve reliability and interpretability in critical infrastructure applications.
- Research Article
- 10.3348/kjr.2025.1045
- Jan 1, 2026
- Korean journal of radiology
To evaluate the accuracy and reasoning capabilities of large multimodal language models compared with those of neuroradiology subspecialty-trained radiologists in neuroradiology case interpretation. This experimental study used custom-made 401 radiologic quizzes derived from articles published in RadioGraphics covering neuroradiology and head and neck topics (October 2020 to February 2024). We prompted the GPT-4 Turbo with Vision (GPT-4V), GPT-4 Omni, Gemini Flash, and Claude models to provide the top three differential diagnoses with a rationale and describe examination characteristics such as imaging modality, sequence, use of contrast, image plane, and body part. The temperature was adjusted to 0 and 1 (T1). Two neuroradiologists answered the same questions. The accuracies of the large language models (LLMs) and the neuroradiologists were compared using generalized estimating equations. Three neuroradiologists assessed the rationale provided by the LLMs for their differential diagnoses using four-point scales, separately for specific lesion locations and imaging findings, and evaluated the presence of hallucinations and the overall acceptability of the responses. Top-3 accuracy (i.e., correct answers present among top-3 differential diagnoses) of LLMs ranged from 29.9% (120 of 401) to 49.4% (198 of 401, obtained with GPT-4V in the T1 setting), while radiologists achieved 80.3% (322 of 401) and 68.3% (274 of 401), respectively (P < 0.001). Regarding the rationale for differential diagnoses, GPT-4V (T1) accurately identified both the specific lesion location and imaging findings in 30.7% (123 of 401) and 12.9% (16 of 124) of cases without textual clinical history. Hallucinations occurred in 4.5% (18 of 401), and only 29.4% (118 of 401) of the LLM-generated analyses were deemed acceptable. GPT-4V (T1) demonstrated high accuracy in identifying the imaging modality (97.4% [800 of 821]) and scanned body parts (92.2% [756 of 820]). LLMs remarkably underperformed compared with neuroradiologists and showed unsatisfactory reasoning for their differential diagnoses, with performance declining further in cases without textual input of clinical history. These findings highlight the limitations of current multimodal LLMs in neuroradiological interpretation and their reliance on text input.
- Research Article
6
- 10.34133/icomputing.0110
- Jan 1, 2025
- Intelligent Computing
Light curves serve as a valuable source of information on stellar formation and evolution. With the rapid advancement of machine learning techniques, they can be effectively processed to extract astronomical patterns and information. In this study, we present a comprehensive evaluation of models based on deep learning and large language models (LLMs) for the automatic classification of variable star light curves, using large datasets from the Kepler and K2 missions. Special emphasis is placed on Cepheids, RR Lyrae, and eclipsing binaries, examining the influence of observational cadence and phase distribution on classification precision. Employing automated deep learning optimization, we achieve striking performance using 2 architectures: one that combines one-dimensional convolution (Conv1D) with bidirectional long short-term memory (BiLSTM) and another called the Swin Transformer. These achieved accuracies of 94% and 99%, respectively, with the latter demonstrating a notable 83% accuracy in discerning the elusive type II Cepheids that comprise merely 0.02% of the total dataset. We unveil StarWhisper LightCurve (LC), a series of 3 LLM models based on an LLM, a multimodal large language model (MLLM), and a large audio language model (LALM). Each model is fine-tuned with strategic prompt engineering and customized training methods to explore the emergent abilities of these models for astronomical data. Remarkably, StarWhisper LC series models exhibit high accuracies of around 90%, considerably reducing the need for explicit feature engineering, thereby paving the way for streamlined parallel data processing and the progression of multifaceted multimodal models in astronomical applications. The study furnishes 2 detailed catalogs illustrating the impacts of phase and sampling intervals on deep learning classification accuracy, showing that a substantial decrease of up to 14% in observation duration and 21% in sampling points can be realized without compromising accuracy by more than 10%.
- Research Article
10
- 10.1016/j.egyr.2025.06.051
- Dec 1, 2025
- Energy Reports
Large Language Models (LLMs) are changing the way we operate our society and will undoubtedly impact power systems as well—but how exactly? By integrating various data streams—including real-time grid data, market dynamics, and consumer behaviors—LLMs have the potential to make power system operations more adaptive, enhance proactive security measures, and deliver personalized energy services. This paper provides a comprehensive analysis of 30 real-world applications across eight key categories: Grid Operations and Management, Energy Markets and Trading, Personalized Energy Management and Customer Engagement, Grid Planning and Education, Grid Security and Compliance, Advanced Data Analysis and Knowledge Discovery, Emerging Applications and Societal Impact, and LLM-Enhanced Reinforcement Learning. Critical technical hurdles, such as data privacy and model reliability, are examined, along with possible solutions. Ultimately, this review illustrates how LLMs can significantly contribute to building more resilient, efficient, and sustainable energy infrastructures, underscoring the necessity of their responsible and equitable deployment. • LLMs can enhance smart grid operations through natural language capabilities. • AI-driven analysis can improve grid security and reliability. • LLMs could personalize energy management and empower consumers. • Ethical AI integration is crucial for equity and trust in smart grids. • LLMs can unlock new optimization opportunities via reinforcement learning.
- Research Article
- 10.1182/blood-2025-6131
- Nov 3, 2025
- Blood
AI multimodal large language model on CAR-T pre-leukapheresis evaluation to predict monitoring needs post infusion for early dismissal planning
- Research Article
27
- 10.1016/j.oneear.2021.04.018
- May 1, 2021
- One Earth
Multiscale design for system-wide peer-to-peer energy trading
- Research Article
7
- 10.1016/j.ecmx.2025.101329
- Oct 1, 2025
- Energy Conversion and Management: X
Artificial intelligence and machine learning for smart grids: from foundational paradigms to emerging technologies with digital twin and large language model-driven intelligence
- Research Article
7
- 10.2214/ajr.25.32729
- Jul 1, 2025
- AJR. American journal of roentgenology
BACKGROUND. The American College of Radiology (ACR) Incidental Findings Committee (IFC) algorithm provides guidance for pancreatic cystic lesion (PCL) management. Its implementation using plain-text large language model (LLM) solutions is challenging given that key components include multimodal data (e.g., figures and tables). OBJECTIVE. The purpose of the study is to evaluate a multimodal LLM approach incorporating knowledge retrieval using flowchart embedding for forming follow-up recommendations for PCL management. METHODS. This retrospective study included patients who underwent abdominal CT or MRI from September 1, 2023, to September 1, 2024, and whose report mentioned a PCL. The reports' Findings sections were inputted to a multimodal LLM (GPT-4o). For task 1 (198 patients: mean age, 69.0 ± 13.0 [SD] years; 110 women, 88 men), the LLM assessed PCL features (presence of PCL, PCL size and location, presence of main pancreatic duct communication, presence of worrisome features or high-risk stigmata) and formed a follow-up recommendation using three knowledge retrieval methods (default knowledge, plain-text retrieval-augmented generation [RAG] from the ACR IFC algorithm PDF document, and flowchart embedding using the LLM's image-to-text conversion for in-context integration of the document's flowcharts and tables). For task 2 (85 patients: mean initial age, 69.2 ± 10.8 years; 48 women, 37 men), an additional relevant prior report was inputted; the LLM assessed for interval PCL change and provided an adjusted follow-up schedule accounting for prior imaging using flowchart embedding. Three radiologists assessed LLM accuracy in task 1 for PCL findings in consensus and follow-up recommendations independently; one radiologist assessed accuracy in task 2. RESULTS. For task 1, the LLM with flowchart embedding had accuracy for PCL features of 98.0-99.0%. The accuracy of the LLM follow-up recommendations based on default knowledge, plain-text RAG, and flowchart embedding for radiologist 1 was 42.4%, 23.7%, and 89.9% (p < .001), respectively; radiologist 2 was 39.9%, 24.2%, and 91.9% (p < .001); and radiologist 3 was 40.9%, 25.3%, and 91.9% (p < .001). For task 2, the LLM using flowchart embedding showed an accuracy for interval PCL change of 96.5% and for adjusted follow-up schedules of 81.2%. CONCLUSION. Multimodal flowchart embedding aided the LLM's automated provision of follow-up recommendations adherent to a clinical guidance document. CLINICAL IMPACT. The framework could be extended to other incidental findings through the use of other clinical guidance documents as the model input.
- Research Article
- 10.1007/s10266-025-01283-2
- Dec 10, 2025
- Odontology
This study aimed to evaluate the diagnostic accuracy of multimodal large language models in classifying superior labial frenulum attachments from intraoral photographs using expert consensus as the reference standard. Five experts (two periodontists and three orthodontists) established the consensus standard by classifying frenulum attachments in 117 intraoral images as mucosal, gingival, papillary, and papilla penetrating. The same photographs were then presented to three multimodal large language models (ChatGPT 4o, Gemini 2.5 Pro, and Microsoft Copilot GPT-4), and their diagnostic performance was evaluated using accuracy, sensitivity, specificity, and F1 score. Reliability was assessed using Fleiss' and Cohen's Kappa, and diagnostic performances were compared using Cochran's Q test. Human raters demonstrated almost perfect agreement (Îş = 0.838, p < 0.001), whereas large language models showed poor inter-model agreement (Îş = -0.124, p < 0.001). ChatGPT achieved slight significant agreement with the consensus (Îş = 0.114, p = 0.019), although its clinical relevance was negligible. Gemini (Îş = 0.099) and Copilot (Îş = 0.027) showed no significant agreement (p > 0.05). Copilot yielded the highest overall accuracy (46.2%), followed by Gemini (44.5%) and ChatGPT (35.0%). The performance of the large language models varied across frenulum types. Current multimodal large language models demonstrate inconsistent and clinically insufficient accuracy in the classification of superior labial frenulum attachments from photographs. Domain-specific training is essential before large language models can be considered reliable diagnostic tools in dentistry.
- Research Article
1
- 10.1371/journal.pone.0329590
- Aug 11, 2025
- PloS one
Enabling large language models (LLMs) to have multi-modal capabilities, such as vision-language learning, has become a current research hotspot and the next milestone in LLM development with the advent of models like GPT4. The basic structure of current multi-modal LLMs usually includes three parts: the image encoder for extracting visual features, the semantic space transformation network ST for aligning the multi-modal semantic spaces, and LLM for generating text. Current works on multi-modal LLMs primarily focus on enhancing performance by utilizing larger image encoders and LLMs, and designing more complex fine-tuning methods and STs, which results in an escalation of model parameters. In this paper, we propose EIM, a novel effective solution for improving the performance of multi-modal large language models from the perspective of training process which reduces the need to introduce new parameters and modify the model structure, and is ignored and less explored in current research. EIM includes corresponding improvement measures in the image encoder, ST, and LLM. To validate EIM, we first apply it to ClipCap and conduct experiments on the COCO Caption dataset. Secondly, we extend EIM to the multi-modal LLMs, such as LLaMA-Adapter and LaVIN, and evaluate them on the ScienceQA dataset. Finally, we also conduct multi-modal chatbot experiments with the EIM enhanced LaVIN and evaluate it on the MME benchmark. The COCO Caption dataset experimental results of [Formula: see text], which is a model that applies EIM on the [Formula: see text], show the 1.75% performance improvement when compared to those of [Formula: see text], which has 3.13 times the number of parameters of [Formula: see text]. The experimental results on the ScienceQA dataset and MME benchmark show that EIM can achieve competitive performance with 7B model parameters when compared to the 13B multi-modal LLMs, which confirms the effective performance improvement of EIM for multi-modal LLMs.
- Research Article
10
- 10.1016/j.csbj.2024.12.019
- Jan 1, 2025
- Computational and structural biotechnology journal
Visual data from images is essential for many medical diagnoses. This study evaluates the performance of multimodal Large Language Models (LLMs) in integrating textual and visual information for diagnostic purposes. We tested GPT-4o and Claude Sonnet 3.5 on 120 clinical vignettes with and without accompanying images. Each vignette included patient demographics, a chief concern, and relevant medical history. Vignettes were paired with either clinical or radiological images from two sources: 100 images from the OPENi database and 20 images from recent NEJM challenges, ensuring they were not in the LLMs' training sets. Three primary care physicians served as a human benchmark. We analyzed diagnostic accuracy and the models' explanations for a subset of cases. LLMs outperformed physicians in text-only scenarios (GPT-4o: 70.8 %, Claude Sonnet 3.5: 59.5 %, Physicians: 39.5 %, p < 0.001, Bonferroni-adjusted). With image integration, all improved, but physicians showed the largest gain (GPT-4o: 84.5 %, p < 0.001; Claude Sonnet 3.5: 67.3 %, p = 0.060; Physicians: 78.8 %, p < 0.001, all Bonferroni-adjusted). LLMs altered their explanatory reasoning in 45-60 % of cases when images were provided. Multimodal LLMs showed higher diagnostic accuracy than physicians in text-only scenarios, even in cases designed to require visual interpretation, suggesting that while images can enhance diagnostic accuracy, they may not be essential in every instance. Although adding images further improved LLM performance, the magnitude of this improvement was smaller than that observed in physicians. These findings suggest that enhanced visual data processing may be needed for LLMs to achieve the degree of image-related performance gains seen in human examiners.
- Research Article
1
- 10.1016/j.dld.2025.11.009
- Dec 1, 2025
- Digestive and liver disease : official journal of the Italian Society of Gastroenterology and the Italian Association for the Study of the Liver
Performance of gastroenterologists and multimodal LLMs in endoscopic EREFS scoring of Eosinophilic Esophagitis.
- Research Article
296
- 10.1038/s41368-023-00239-y
- Jul 28, 2023
- International Journal of Oral Science
The ChatGPT, a lite and conversational variant of Generative Pretrained Transformer 4 (GPT-4) developed by OpenAI, is one of the milestone Large Language Models (LLMs) with billions of parameters. LLMs have stirred up much interest among researchers and practitioners in their impressive skills in natural language processing tasks, which profoundly impact various fields. This paper mainly discusses the future applications of LLMs in dentistry. We introduce two primary LLM deployment methods in dentistry, including automated dental diagnosis and cross-modal dental diagnosis, and examine their potential applications. Especially, equipped with a cross-modal encoder, a single LLM can manage multi-source data and conduct advanced natural language reasoning to perform complex clinical operations. We also present cases to demonstrate the potential of a fully automatic Multi-Modal LLM AI system for dentistry clinical application. While LLMs offer significant potential benefits, the challenges, such as data privacy, data quality, and model bias, need further study. Overall, LLMs have the potential to revolutionize dental diagnosis and treatment, which indicates a promising avenue for clinical application and research in dentistry.
- Research Article
- 10.7759/cureus.106486
- Apr 1, 2026
- Cureus
Introduction Large language models (LLMs) have demonstrated promising performance on standardized medical examinations, yet systematic comparisons of contemporary multimodal and text-only models on radiology-specific assessments remain limited. Updated and newly released LLMs, including Grok 4.1 (xAI, San Francisco, USA), Bing Copilot GPT-5 (Microsoft, Redmond, USA), DeepSeek V3.2 (DeepSeek AI, Beijing, China), and OpenEvidence (Chalmers University of Technology, Gothenburg, Sweden), have not been evaluated on the American College of Radiology Diagnostic Imaging In-Training (ACR DXIT) examination. This study aimed to compare the performance of seven contemporary LLMs on the 2022 DXIT examination, stratified by question format and radiology subject domain. Methods Seven LLMs were evaluated on all 106 multiple-choice questions from the 2022 DXIT examination, comprising 42 written-only and 64 image-based questions. Five multimodal models [ChatGPT-5.1 (OpenAI, San Francisco, USA), Gemini 3 Pro (Google, Mountain View,USA), Claude Sonnet 4.5 (Anthropic, San Francisco, USA), Grok 4.1, and Bing Copilot GPT-5] were assessed on all questions. Two text-only models (DeepSeek V3.2 and OpenEvidence) were evaluated on written-only questions. A standardized orientation prompt was applied uniformly across all models. Statistical comparisons accounted for the paired nature of the data, as all models answered identical questions; Cochran's Q test was used for comparisons across three or more models, and McNemar's test for two-model comparisons. Ninety-five percent confidence intervals for accuracy proportions were calculated using the Wilson score method. For subgroups with fewer than 10 questions, p-values were not reported, and descriptive statistics only are presented. Results Overall accuracy among multimodal models ranged from 65.1% [Claude Sonnet 4.5; 95% confidence interval (CI): 55.6%-73.5%] to 76.4% (Gemini 3 Pro; 95% CI: 67.5%-83.5%), with no statistically significant differences among models (Cochran's Q=5.07, df=4, p=0.281). All multimodal models performed substantially better on written-only questions (88.1%-95.2%) than on image-based questions (46.9%-64.1%), representing an average gap of approximately 35 percentage points. Neither written-only nor image-based comparisons reached significance (p=0.813 and p=0.226, respectively). Domain-level analysis identified consistent strengths in ultrasound (80%-90%; p=0.948) and chest radiology (70%-90%; p=0.870), and persistent weakness in musculoskeletal imaging (40%-60%; p=0.898). Among text-only models, OpenEvidence and DeepSeek V3.2 achieved overall accuracies of 83.3% (95% CI: 69.4%-91.7%) and 88.1% (95% CI: 75.0%-94.8%), respectively, with no significant difference between them (McNemar's p=0.773). Conclusion Contemporary multimodal LLMs achieve moderately high accuracy on radiology in-training examination questions, exceeding earlier-generation model benchmarks and junior resident performance levels, yet no single model demonstrated statistically significant superiority. A consistent and substantial performance gap between written and image-based questions persists across all architectures, underscoring unresolved limitations in radiologic image interpretation. These findings suggest that current LLMs may support circumscribed roles in radiology education, particularly for conceptual and non-interpretive content, but remain unsuitable for tasks requiring visual diagnostic reasoning.
- Research Article
23
- 10.1007/s00261-024-04708-8
- Dec 2, 2024
- Abdominal radiology (New York)
Large language models (LLMs) and multi-modal large language models (MLLMs) represent the cutting-edge in artificial intelligence. This review provides a comprehensive overview of their capabilities and potential impact on radiology. Unlike most existing literature reviews focusing solely on LLMs, this work examines both LLMs and MLLMs, highlighting their potential to support radiology workflows such as report generation, image interpretation, EHR summarization, differential diagnosis generation, and patient education. By streamlining these tasks, LLMs and MLLMs could reduce radiologist workload, improve diagnostic accuracy, support interdisciplinary collaboration, and ultimately enhance patient care. We also discuss key limitations, such as the limited capacity of current MLLMs to interpret 3D medical images and to integrate information from both image and text data, as well as the lack of effective evaluation methods. Ongoing efforts to address these challenges are introduced.
- Research Article
6
- 10.3390/bdcc9050132
- May 16, 2025
- Big Data and Cognitive Computing
This study presents a comparative analysis of several multimodal large language models (LLMs) for no-reference image quality assessment, with a particular focus on images containing authentic distortions. We evaluate three models developed by OpenAI and three models from Claude.AI, comparing their performance in estimating image quality without reference images. Our results demonstrate that these LLMs outperform traditional methods based on hand-crafted features. However, more advanced deep learning models, especially those based on deep convolutional networks, surpass LLMs in performance. Notably, we make a unique contribution by publishing the processed outputs of the LLMs, providing a transparent and direct comparison of their quality assessments based solely on the predicted quality scores. This work underscores the potential of multimodal LLMs in image quality evaluation, while also highlighting the continuing advantages of specialized deep learning approaches.