Discovery Logo
Sign In
Search
Paper
Search Paper
R Discovery for Libraries Pricing Sign In
  • Home iconHome
  • My Feed iconMy Feed
  • Search Papers iconSearch Papers
  • Library iconLibrary
  • Explore iconExplore
  • Ask R Discovery iconAsk R Discovery Star Left icon
  • Literature Review iconLiterature Review NEW
  • Chat PDF iconChat PDF Star Left icon
  • Citation Generator iconCitation Generator
  • Chrome Extension iconChrome Extension
    External link
  • Use on ChatGPT iconUse on ChatGPT
    External link
  • iOS App iconiOS App
    External link
  • Android App iconAndroid App
    External link
  • Contact Us iconContact Us
    External link
  • Paperpal iconPaperpal
    External link
  • Mind the Graph iconMind the Graph
    External link
  • Journal Finder iconJournal Finder
    External link
Discovery Logo menuClose menu
  • Home iconHome
  • My Feed iconMy Feed
  • Search Papers iconSearch Papers
  • Library iconLibrary
  • Explore iconExplore
  • Ask R Discovery iconAsk R Discovery Star Left icon
  • Literature Review iconLiterature Review NEW
  • Chat PDF iconChat PDF Star Left icon
  • Citation Generator iconCitation Generator
  • Chrome Extension iconChrome Extension
    External link
  • Use on ChatGPT iconUse on ChatGPT
    External link
  • iOS App iconiOS App
    External link
  • Android App iconAndroid App
    External link
  • Contact Us iconContact Us
    External link
  • Paperpal iconPaperpal
    External link
  • Mind the Graph iconMind the Graph
    External link
  • Journal Finder iconJournal Finder
    External link
features
  • Audio Papers iconAudio Papers
  • Paper Translation iconPaper Translation
  • Chrome Extension iconChrome Extension
Content Type
  • Journal Articles iconJournal Articles
  • Conference Papers iconConference Papers
  • Preprints iconPreprints
  • Seminars by Cassyni iconSeminars by Cassyni
More
  • R Discovery for Libraries iconR Discovery for Libraries
  • Research Areas iconResearch Areas
  • Topics iconTopics
  • Resources iconResources

Related Topics

  • Flesch Reading Ease Score
  • Flesch Reading Ease Score
  • Simple Measure Of Gobbledygook
  • Simple Measure Of Gobbledygook
  • Reading Grade Level
  • Reading Grade Level

Articles published on Flesch-Kincaid Grade Level

Authors
Select Authors
Journals
Select Journals
Duration
Select Duration
2439 Search results
Sort by
Recency
  • New
  • Research Article
  • Cite Count Icon 1
  • 10.1111/evj.70074
Cost of referral treatment for colic in the United Kingdom-What has changed in the last 5 years?
  • Jul 1, 2026
  • Equine veterinary journal
  • F E Wilson + 2 more

Referral treatment costs and insurance status impact treatment decisions for colic. To evaluate changes in the cost of referral treatment for colic, and insurance cover and premiums in the United Kingdom between 2018 and 2023. Cross sectional study. Thirty UK equine referral hospitals were contacted in January 2024 and asked about their colic caseload and costs of the last three cases across six categories (surgical +/- resection, euthanasia before, during or after surgery, and medical treatment), using similar methodology to a 2018 study. Data are reported as mean/median (range). A standardised case was used to retrieve data on veterinary fees, insurance cover, and monthly premiums from five companies. Findings were compared with actual and inflation-adjusted 2018 data. Readability of insurance documents were assessed using the Flesch Kincaid Reading Ease (FKRE) score and the Gunning Fog Score (GFS). The FKRE is ranked from 0 to 100 (easy to read-hard to read); FKRE scores below are 65 recommended. The GFS estimates the years of formal education needed to understand text; GFS scores higher than 12 are too complex for most people to read. Eighteen hospitals responded, contributing costings for 248 cases in total. Mean/median (range) costs for cases euthanised without surgery (n = 41) were £1200 (£500-£4389), for medical cases (n = 44) were £2379 (£683-£13,762), and for all surgical cases that survived surgery (n = 122) were £7905 (£3023-£20,343). When compared with inflation-adjusted 2018 data, medical treatment and euthanasia without surgery costs had increased; surgery costs had decreased. Maximum insurance cover was between £5000 and £7500. The actual cover value had not changed for 3/5 companies since 2018, and was reduced for 4/5 companies after inflation adjustment. Monthly premiums ranged from £42.76 to £97.23, and were all increased compared with 2018 inflation-adjusted data (£34.01-£59.39). Insurance document FKRE Scores ranged from 31.2 to 54.8, and GFS ranged from 13.6 to 20.6. All were outside the recommended range. Small case numbers, UK population only. Costs of referral treatment have largely risen in line with inflation, and now frequently exceed maximum insurance cover. Insurance premiums have increased above inflation, and insurance documents remain complex and hard to read.

  • New
  • Research Article
  • 10.1161/strokeaha.126.055985
Improving Readability of Stroke Clinical Trial Consent Forms Using Artificial Intelligence.
  • Jul 1, 2026
  • Stroke
  • Rohan Arora + 6 more

Informed consent forms (ICFs) for clinical trials are often written above the recommended eighth-grade level. We aimed to compare the readability of original ICFs used for National Institutes of Health-funded stroke-related clinical trials with ICFs edited for readability using artificial intelligence. Publicly available ICFs associated with National Institutes of Health-funded stroke-related clinical trials were accessed through ClinicalTrials.gov (search period: inception to August 12, 2025). Using ChatGPT-4o, we created a customized Generative Pre-Trained Transformer (GPT) designed to lower the reading level to eighth grade or below while maintaining ICF content. We processed each ICF using this GPT to create edited ICFs. Standard readability metrics, including the Flesch-Kincaid grade level (primary outcome), were compared between original and edited ICFs using paired t tests or the McNemar test (cross-sectional design). We also assessed semantic similarity using the MPNet language model, which produced continuous scores from 0 (no similarity) to 1 (perfect similarity). ICFs were available for 46 stroke trials, including behavioral (n=21), device (n=15), drug (n=5), and other (n=5) intervention types. Mean reading levels were 11.52 for the original and 9.47 for the GPT-edited ICFs using the Flesch-Kincaid grade level (P<0.001). Only 1 (2%) of the original ICFs and 18 (39%) of the GPT-edited ICFs had a Flesch-Kincaid reading level at or below eighth grade (P<0.001). Both the Simple Measure of Gobbledygook and Gunning Fog Index favored the GPT-edited ICFs by 1 to 2 grade levels. The Flesch Reading Ease score favored the GPT-edited ICFs by about 8 points. The mean similarity score was 0.85 (SD=0.04). GPT-edited ICFs achieved a readability reduction of approximately 2 grade levels compared with the original ICFs while preserving high semantic similarity. Customized GPTs may be a useful tool to improve the readability of clinical trial ICFs.

  • New
  • Research Article
  • 10.1016/j.ijporl.2026.112850
Exploring one aspect of family-centered practice: Readability of written documents to support parental understanding and decision-making regarding cochlear implants.
  • Jul 1, 2026
  • International journal of pediatric otorhinolaryngology
  • Shani Dettman + 4 more

Exploring one aspect of family-centered practice: Readability of written documents to support parental understanding and decision-making regarding cochlear implants.

  • New
  • Research Article
  • 10.1093/milmed/usag296
Artificial Intelligence Answering Femoroacetabular Impingement Patient Questions: Helpful Tool or Harmful Risk? Evaluating NIPRGPT Answers to Frequently Asked Questions About Femoroacetabular Impingement.
  • Jun 30, 2026
  • Military medicine
  • William J Ferris + 6 more

As patients increasingly seek medical information online, artificial intelligence (AI) chatbots like NIPRGPT-the most widely available AI tool for Department of Defense (DOD) computer users-offer a novel resource for addressing queries about femoroacetabular impingement (FAI). To date, there have not been any studies evaluating NIPRGPT responses to orthopedic medical questions. The primary objective of this study was to evaluate the accuracy, comprehensiveness, and readability of NIPRGPT's responses to common FAI-related questions. Twelve frequently asked questions (FAQs) regarding FAI were selected from a curated list and posed to NIPRGPT. The accuracy and adequacy of the responses were graded by a panel of board-certified surgeons as excellent (not requiring clarification), satisfactory (requiring minimal clarification), satisfactory (requiring moderate clarification), or unsatisfactory (requiring substantial clarification). Additionally, readability was assessed using the Flesch-Kincaid readability score. Of the 12 responses, four (33.3%) were excellent, requiring no clarification, seven (58.3%) were satisfactory, requiring minimal clarification, and one (8.3%) was satisfactory, requiring moderate clarification. No responses were deemed unsatisfactory. The average quality score was 3.38/4.0. However, the average Flesch-Kincaid readability score was a 19.6 Grade Level, indicating a reading level suited for postgraduate or specialized academic backgrounds. Interobserver agreement was low, with a Krippendorff's alpha of 0.046. NIPRGPT provides answers to FAQs about FAI that are generally accurate and reliable. However, the responses are generated at a complexity level far exceeding the recommended reading level for patient education. While a potentially useful adjunct in military healthcare settings where access may be limited, clinicians must be aware of the high literacy demand placed on patients using this tool.

  • New
  • Research Article
  • 10.1186/s12911-026-03662-3
Performance of large language models in mitral valve surgery patient education: a comparative analysis.
  • Jun 29, 2026
  • BMC medical informatics and decision making
  • Banu Bahriye Akdag + 9 more

Large language models (LLMs), a form of artificial intelligence, are increasingly being utilized in healthcare to support patient education and information delivery. The aim of this study was to perform a comparative analysis of five different LLMs (i.e., ChatGPT-4o, Claude 3.7 Sonnet, Gemini 2.5 Pro Preview, DeepSeek-V3, and Microsoft Copilot) in terms of accuracy, completeness, and readability, based on their responses to frequently asked questions in preoperative patient education for mitral valve surgery (MVS). A standardized questionnaire comprising seven frequently asked questions by patients prior to MVS was developed. Prompting procedures and model parameters were fully reported to support reproducibility. These questions were presented to each LLM in an identical manner. The responses were evaluated by two academic experts in cardiac surgery using structured assessment criteria across three main dimensions: accuracy, completeness, and readability. For the readability analysis, the Simplified Measure of Gobbledygook (SMOG) Index and the Flesch-Kincaid Grade Level (FKGL) scale were utilized. The ChatGPT-4o and Gemini 2.5 Pro Preview models received statistically significantly higher scores than Claude 3.7 Sonnet and Microsoft Copilot for both accuracy (median 5 for ChatGPT-4o and Gemini 2.5 Pro Preview vs. 4 for Claude 3.7 Sonnet and Microsoft Copilot, p < 0.001) and completeness (median 5 for Gemini 2.5 Pro Preview vs. 3 for Claude 3.7 Sonnet, p < 0.001). Claude 3.7 Sonnet achieved the highest readability scores, with significantly lower SMOG (10.90 for Claude 3.7 Sonnet vs. 12.24 for ChatGPT-4o, p = 0.006) and FKGL (8.0 for Claude 3.7 Sonnet vs. 9.04 for ChatGPT-4o, p = 0.004) scores, indicating simpler and more comprehensible sentence structures. Significant differences were observed among the evaluated models across all three assessment dimensions (p < 0.001 for all comparisons). The LLMs represent valuable supplementary tools in patient education processes. However, their implementation in clinical practice must be carefully evaluated, particularly with regard to accuracy and completeness. This study highlights the potential applicability of ChatGPT-4o and Claude 3.7 Sonnet models for preoperative patient education in MVS, while emphasizing that all LLMs should be used under the supervision and guidance of healthcare professionals. For LLMs to be reliably utilized in the medical field, improvement in medical accuracy and standardization are essential.

  • New
  • Research Article
  • 10.1016/j.jogc.2026.103442
Assessment of the accuracy and readability of artificial generative intelligence for patient questions on heavy menstrual bleeding.
  • Jun 29, 2026
  • Journal of obstetrics and gynaecology Canada : JOGC = Journal d'obstetrique et gynecologie du Canada : JOGC
  • Olivia Carere + 7 more

To evaluate the quality, patient-centeredness, clinician endorsement, and readability of generative artificial intelligence (AI) responses to patient questions about heavy menstrual bleeding (HMB) compared with high-quality patient-facing websites. A cross-sectional study compared responses generated by ChatGPT and Google Gemini to excerpts from the five highest-quality HMB patient resources identified through a Google Trends-informed search and the QUality Evaluation Scoring Tool (QUEST). Five layperson-style questions representing HMB subtopics (definition, causes, investigations, management, and safety) were submitted to each model. Responses were de-identified and independently evaluated by five gynecologists, blinded to source, using 5-point Likert scales for accuracy/comprehensiveness (quality), empathetic and validating communication style (patient-centeredness), and expert comfort with patient use of the resource to guide understanding (clinician endorsement). Readability was assessed using the Flesch-Kincaid Grade Level (FKGL). Inter-rater reliability was measured using intraclass correlation coefficients (ICCs). Thirty-six responses were reviewed (16 AI-generated and 20 web-based). Compared with web-based excerpts, AI responses showed similar quality (3.94/5±0.70 vs. 3.53/5±1.21; p=0.21), lower patient-centeredness (2.81/5±0.82 vs. 3.52/5±1.13; p=0.04), and comparable clinician endorsement (3.56/5±0.92 vs. 3.44/5±1.35; p=0.52). Without impacting content quality, AI responses were written at a lower grade level (7.6±2.4) than web-based excerpts (11.1±5.2; p=0.042). Inter-rater reliability was high (ICC=0.82-0.84). Compared with high-quality web-based resources, AI-generated responses to HMB questions were more inclusive for varying health literacy levels, comparable in quality and clinician endorsement, but performed worse in patient-centeredness. In an evolving digital landscape for health information acquisition, AI-generated responses represent a valuable and safe supplementary or first-line resource for patients.

  • New
  • Research Article
  • 10.1515/jpm-2026-0007
Readability of patient educational materials on ultrasound: a cross-sectional study.
  • Jun 25, 2026
  • Journal of perinatal medicine
  • Yinka Oyelese + 2 more

To evaluate the readability and quality of publicly available patient information pamphlets on ultrasound and assess their accessibility for patients with varying literacy levels. This was a cross-sectional descriptive study using the publicly available online International Society of Ultrasound in Obstetrics and Gynecology (ISUOG) patient information library. A total of 155English-language patient information materials ("pamphlets") on pregnancy and gynecology topics available in early 2025 were analyzed. Readability was assessed using Readability Studio™ software and four validated indices: Gunning Fog, SMOG, Coleman-Liau, and Flesch Reading Ease (FRE). The DISCERN instrument, a validated 16-item tool, was applied independently to evaluate reliability, clarity, and balance of treatment information. The main outcomes measures were the grade level of readability and DISCERN quality scores. Only one pamphlet (1 %) met the recommended eighth-grade readability standard. Most pamphlets (124; 80 %) were written at or above the 11th-grade level (mean Gunning Fog 14.8, SMOG 13.5, Coleman-Liau 12.9, FRE 45.2). The LIX index classified the majority as "difficult to technical." Despite the high reading level, DISCERN scores were uniformly high (4-5/5), indicating strong reliability, clarity, and balance of information, but poor accessibility for the average patient. ISUOG patient information materials are accurate, reliable, and evidence-based but written well above recommended readability standards, limiting comprehension for many patients. Simplifying language, shortening sentences, and involving health-literacy and cultural experts may improve accessibility and promote global equity in patient education.

  • New
  • Research Article
  • 10.1044/2026_jslhr-26-00031
Preclinical Dialogue Simulation: Evaluating Response Accessibility in Conversational Artificial Intelligence for Aphasia Therapy.
  • Jun 24, 2026
  • Journal of speech, language, and hearing research : JSLHR
  • Gerald C Imaezue + 4 more

Large language model (LLM)-driven conversational agents are increasingly considered for use in clinical contexts, yet systematic approaches for evaluating their behavior in impairment-rich, speech-based therapeutic interactions remain limited. This study extends the Agent-Based Conversational Dialogue (ABCD) simulation method as a preclinical testbed to evaluate how LLMs generate accessible clinician language when responding to characteristic aphasic speech during Response Elaboration Training. ABCD was used to simulate multi-turn spoken therapeutic dialogues between an LLM-driven clinician and an artificial intelligence (AI)-simulated aphasic patient, enabling controlled manipulation of impairment profiles, prompting strategies, and reasoning modes without human participants. Three LLM families (Claude, GPT, and Gemini) were benchmarked under zero-shot and few-shot prompting and standard versus advanced reasoning. Response accessibility was quantified using established readability metrics (Flesch Reading Ease, Dale-Chall) and a composite score derived from 16 standardized readability measures. Distinct accessibility signatures emerged across model architectures and configurations. Few-shot prompting and advanced reasoning generally yielded more accessible clinician responses, whereas Gemini demonstrated superior accessibility under zero-shot, standard reasoning. LLMs differ systematically in their capacity to adapt clinician language to impaired speech. ABCD provides a scalable, preclinical dialogue simulation framework for benchmarking conversational AI in clinically oriented, impairment-rich dialogue across multidimensional constraints. It offers guidance for model selection and configuration prior to clinical translation in communication rehabilitation. https://doi.org/10.23641/asha.32736078.

  • New
  • Research Article
  • 10.3389/fonc.2026.1757933
Development and expert radiologist validation of a custom pipeline for simplification of oncology radiology reports using large language model
  • Jun 23, 2026
  • Frontiers in Oncology
  • Prerna Garg + 9 more

Purpose Radiology reports are often written for clinicians and contain complex medical terminology, which patients struggle to understand. This comprehension gap may lead to anxiety and misinformed decisions. This is especially important in cancer care. Large language models (LLMs) provide an opportunity to bridge this gap. Direct use of LLM by patients may potentially be misleading and may not ensure clinical fidelity. We developed an LLM- based tool to automatically translate radiology findings into clear, patient-friendly language which could improve patient-centered care. Objective To develop and validate a bilingual LLM driven tool that simplifies oncology radiology reports into patient-understandable English and Hindi, while maintaining diagnostic fidelity and emotional tone. Methods This study was approved by the Institute’s ethics committee. A retrospective corpus of 100 computed tomography (CT) reports (April 2025-July 2025) of patients with colo-rectal cancers were used for the development of the pipeline. Five large language models—GPT-4o, Gemini 2.5 Pro, Claude Opus, LLaMA-3.1-8B and Phi-3.5-mini—were tested. Five iterative prompt versions guided LLMs through successive refinements to ensure medical accuracy, clarity, and inclusion of a standard disclaimer. A custom pipeline; Vernacular Language Coverter(VLC), based on Gemini 2.5 Pro that takes a radiologist’s report as input and generates two patient-facing outputs: (1) a simplified English and (2) a Hindi explanation was developed. The tool thus developed was then prospectively validated using 100 de-identified reports. Simplified outputs were reviewed by two radiologists, assessing accuracy, language clarity/Terminology and readability/tone. Completeness was assessed in terms of core diagnostic completeness and minor incompleteness. Flesch Reading Ease (FRE) was calculated for a fraction of reports. Results Mean rubric scores were high: Accuracy 4.77 ± 0.62, language clarity/Terminology 4.78 ± 0.56 and readability/tone 4.9 ± 0.32. Core diagnostic completeness was attained in all patients. 92% percent of reports were released ‘as is. Readability improved markedly (Flesch Reading Ease 49.2→73). Empathy phrases in English and Hindi were appropriate. Conclusions An AI-driven bilingual framework significantly enhances the clarity, tone, and readability of oncology radiology reports while retaining diagnostic precision. This tool demonstrates the feasibility of safe, patient-centered communication aligned with the goals of personalized medicine.

  • New
  • Research Article
  • 10.1177/17504589261454163
Comparative evaluation of traditional versus generative AI patient education material on prehabilitation.
  • Jun 23, 2026
  • Journal of perioperative practice
  • Aritra Kundu + 3 more

Effective patient education materials on prehabilitation are essential for optimising patients before surgery. With the growing use of generative artificial intelligence (AI) chatbots in health care communication, it is important to evaluate their suitability compared with established human-generated resources. We aimed to compare patient education materials on prehabilitation, generated by artificial intelligence chatbots (ChatGPT-4o, Gemini 2.5, and DeepSeek V3) with a National Health Service leaflet, assessing factual accuracy, readability, and emotional tone. A comparative observational study design was used. All four patient education materials were blinded and evaluated by ten experts using a 10-point Likert-type scale. Readability was assessed using the Flesch Reading Ease and the Flesch-Kincaid Grade Level. Sentiment analysis was done using an online tool. Patient Education Materials Assessment Tool for Printable Materials (PEMAT-P) scores were calculated. National Health Service patient education material showed the highest mean ± SD accuracy scores 9.8 ± 0.3 from experts outperforming all artificial intelligence models (p = 0.000). Among artificial intelligence, Gemini scored highest. For readability, ChatGPT and the National Health Service were comparable. Sentiment analysis showed varying tones across all models. Patient Education Materials Assessment Tool for Printable Materials scores showed high understandability across all patient education materials (>75%), but actionability was highest for the National Health Service (93.3%). Artificial intelligence chatbots can generate readable and promising patient education materials. Traditional materials remain superior in accuracy and completeness. A hybrid 'human-in-the-loop' approach is recommended for effective patient education.

  • New
  • Research Article
  • 10.36141/svdld.2026.18399
Quality of AI chatbot-generated information on hypersensitivity pneumonitis for clinical and patient use.
  • Jun 22, 2026
  • Sarcoidosis, vasculitis, and diffuse lung diseases : official journal of WASOG
  • Derya Yenibertiz + 1 more

Hypersensitivity pneumonitis (HP) is a complex, immüne mediated interstitial lung disease in which accurate diagnosis and long term management require integration of clinical, radiologic, and exposure-related information. Patients increasingly use artificial intelligence (AI) based chatbots to obtain disease related information; however, the quality, readability, and patient usability of such content remain unclear. This study aimed to evaluate the quality, reliability, readability, and patient-centered usability of AI chatbot generated information on HP. Using Google Trends, we identified four of the most frequently searched patient-oriented questions regarding HP: (1) What is HP and what causes it? (2) What are the clinical features of HP? (3) How is HP treated? (4) How is HP diagnosed? These questions were submitted verbatim to eight AI chatbots (ChatGPT-5.1, Claude 3, Microsoft Copilot, DeepSeek V3, Gemini Pro, Grok 4, Kimi K2, Perplexity AI). A total of 32 responses were independently evaluated in a blinded fashion by four pulmonology professors specializing in interstitial lung diseases. Content quality and reliability were assessed using DISCERN; understandability and actionability with PEMAT-P; global written readability with the Written Readability Rating (WRR); and structural readability with the Flesch-Kincaid Grade Level (FKGL). All chatbot outputs required advanced literacy, with FKGL scores ranging from 20.17 to 29.07 and a mean of approximately 24-25, indicating college or postgraduate reading level. No chatbot produced content within the recommended patient-appropriate range (FKGL ≤ 8). WRR scores declined with increasing clinical complexity, from 67.85 for definitional content (Q1) to 51.227 for diagnostic explanations (Q4). DISCERN scores varied substantially across models (35.001-57.103), with most chatbots falling into the "fair-good" range, reflecting partially reliable but incomplete information. [..] Conclusion: AI chatbots can generate clinically rich explanations of HP but currently produce content that is too complex and insufficiently actionable for most patients. [..].

  • New
  • Research Article
  • 10.1016/j.ijotn.2026.101297
Portuguese adaptation and psychometric validation of the Osteoporosis Knowledge Assessment Tool in community-dwelling adults.
  • Jun 20, 2026
  • International journal of orthopaedic and trauma nursing
  • Tiago Silva + 5 more

Portuguese adaptation and psychometric validation of the Osteoporosis Knowledge Assessment Tool in community-dwelling adults.

  • New
  • Research Article
  • 10.2196/93054
Evaluation of Five Large Language Models for Parental Education in Pediatric Anesthesia: Reliability and Readability Study
  • Jun 18, 2026
  • JMIR Medical Informatics
  • Fulin Pu + 3 more

BackgroundAlthough large language models (LLMs) show potential for patient education, their accuracy, usability, and comprehensibility lack validation in high-risk pediatric anesthesia. Rigorous evaluation is therefore essential prior to widespread clinical use in perioperative parental anesthesia education.ObjectiveThis study aims to evaluate the accuracy, reliability, and readability of responses generated by 5 LLMs to parental inquiries regarding pediatric anesthesia, and to assess their suitability for clinical use in perioperative caregiver education.MethodsTwo expert anesthesiologists identified 33 parental questions on pediatric anesthesia by screening authoritative resources and Google Trends. On December 14, 2025, these questions were submitted to 5 LLMs (DeepSeek-V3.2, ChatGPT-5, Gemini 2.5 Flash, Copilot, and Perplexity) via official web interfaces with default settings and zero-shot prompting, with each query in a separate conversation. Responses were standardized for blinded assessment. Two pediatric anesthesiologists with ≥10 years of clinical experience independently evaluated accuracy and reliability using the 4-point Likert accuracy scale, DISCERN, Ensuring Quality Information for Patients (EQIP), Journal of the American Medical Association (JAMA) benchmark, and Global Quality Score (GQS). After text preprocessing, readability was evaluated using 6 algorithms (Automated Readability Index [ARI], Flesch Reading Ease Score [FRES], Gunning Fog Index [GFI], Flesch-Kincaid Grade Level [FKGL], Coleman-Liau Index [CL], and the Simple Measure of Gobbledygook [SMOG]) via an online calculator. Interrater reliability was analyzed using the intraclass correlation coefficient (ICC); differences across models were assessed with the Kruskal-Wallis H test; and deviations from the sixth-grade benchmark were evaluated using 1-sample Wilcoxon signed-rank tests (P<.05 considered significant).ResultsAll 5 LLMs demonstrated high clinical accuracy (>90%; P=.12), with Gemini reaching 100%. Nevertheless, safety risks and content hallucinations were still observed. Excluding Gemini and Copilot, the remaining 3 models (ChatGPT, DeepSeek, and Perplexity) each produced unsafe content in 3.03% (n=1) of the 33 queries. Hallucinations were detected in all models except Gemini, with DeepSeek and Perplexity showing the highest hallucination rate (3/33, 9.09%). Furthermore, Perplexity showed superior reliability on DISCERN (median 41; P<.05), yet no model achieved a “good” rating. Gemini achieved the highest EQIP (median 66.67%; P<.05) despite lower GQS (median 3). Transparency was universally poor (JAMA median ≤1), with DeepSeek and ChatGPT showing a “floor effect.” ChatGPT had superior readability, but all models exceeded the recommended 6-grade complexity level.ConclusionsIn this study, 5 LLMs generally provided clinically accurate information when responding to parental questions about pediatric anesthesia. However, limitations were also identified, including hallucinated content, safety-related deficiencies, limited source transparency, and readability levels exceeding recommended standards. Therefore, LLM-generated information should be interpreted with caution and should not replace clinician guidance.

  • New
  • Research Article
  • 10.1007/s13187-026-02926-w
Evaluation of Web-based Information on Phytotherapy for Cancer Patients: A Quality and Readability Analysis.
  • Jun 18, 2026
  • Journal of cancer education : the official journal of the American Association for Cancer Education
  • Mehmet Uzun + 3 more

Phytotherapy is widely used by cancer patients as a complementary and alternative medicine approach. With the increasing reliance on the internet for health-related information, concerns regarding the quality, reliability, and readability of online phytotherapy content have become more prominent. This study aimed to evaluate the readability and quality of web-based information on phytotherapy for cancer patients using validated assessment tools and to identify specific deficiencies in content quality. A descriptive cross-sectional analysis was conducted using the Google search engine with four predefined search terms related to phytotherapy and oncology. The first 50 websites for each term were screened, yielding 200 websites, of which 99 met the inclusion criteria. Websites were categorized by source type and visibility. Readability was assessed using the Flesch-Kincaid Grade Level (FKGL), Gunning Fog Index, SMOG, and Coleman-Liau Index. Content quality was evaluated using the JAMA benchmark criteria and the DISCERN instrument, including item-level analysis. Non-parametric statistical tests were applied where appropriate. The median FKGL score was 9.3, indicating that most content required a high reading level. The median JAMA score was 4, while the median DISCERN score was 55, reflecting moderate but variable quality. Item-level analysis revealed that critical aspects such as treatment risks, benefits, uncertainties, and consequences of no treatment were frequently insufficiently addressed. Commercial websites demonstrated lower DISCERN scores compared with non-commercial sources. No significant differences were observed between first-page and subsequent search results. Online phytotherapy information for cancer patients is characterized by moderate quality, high readability demands, and important deficiencies in key domains necessary for informed decision-making. In the evolving landscape of AI-assisted health information retrieval, these limitations may have broader implications, highlighting the need for accurate, evidence-based, and accessible online resources.

  • New
  • Research Article
  • 10.1007/s00192-026-06783-5
Evaluating Readability and Understandability of Patient-Reported Questionnaires for Bladder Pain Syndrome.
  • Jun 17, 2026
  • International urogynecology journal
  • Anna M Leone + 6 more

Patient outcome questionnaires help identify patients with bladder pain syndrome and determine treatment response. Our objective was to compare the readability and understandability of questionnaires assessing bladder pain syndrome. We assessed the readability and understandability of four questionnaires: The O'Leary-Sant, the Pelvic Pain and Urgency/Frequency (PUF) Symptom Score, the Female Genitourinary Pain Index (FGUPI), and the Bladder Pain/IC Symptom Score (BPIC-SS). We used four validated readability calculators: Flesh-Kincaid Grade Score (FKGS), Flesch Reading Ease Score (FRES), Fry Reading Graph Calculator (FRY), and Simple Measure of Gobbledygook (SMOG). The American Medical Association defines readability as a 6th grade reading level. We measured understandability using the Patient Education Materials Assessment Tool (PEMAT). Scores ≥ 70% represented acceptable understandability. Only the BPIC-SS met the 6th grade reading standard when evaluated using the FKGS assessment tool. The other questionnaires' scores ranged from 7th grade to college reading level. When assessing the question section only for each questionnaire, the PUF and BPIC-SS met the 6th grade readability standard. Reduced readability was due to use of polysyllabic words such as "urination" and "frequency," long questions and answer choices. None of the questionnaires met the PEMAT understandability threshold score of 70%. Reduced understandability was due to confusing structure, failure to define medical terms, and absence of visual aids. Questionnaires assessing bladder pain syndrome do not reliably meet recommended standards for patient readability and understandability and therefore may not accurately measure patient responses. These standards should be considered when designing future questionnaires to improve accessibility to all patients.

  • New
  • Research Article
  • 10.1016/j.fas.2026.06.006
Can artificial intelligence provide reliable patient education in minimally invasive bunion surgery? A comparative study of ChatGPT 5.2 and Gemini.
  • Jun 16, 2026
  • Foot and ankle surgery : official journal of the European Society of Foot and Ankle Surgeons
  • Berat Yüksel + 5 more

Can artificial intelligence provide reliable patient education in minimally invasive bunion surgery? A comparative study of ChatGPT 5.2 and Gemini.

  • New
  • Research Article
  • 10.1093/jamiaopen/ooag078
Enhancing the quality and trustworthiness of large language model-generated summaries of clinical oncology literature
  • Jun 16, 2026
  • JAMIA Open
  • Arnulf Stenzl + 12 more

ObjectivesThis study evaluated the quality and trustworthiness of large language model (LLM)-generated scientific and plain language summaries (PLS) from clinical oncology literature, focusing on faithfulness (absence of hallucinations), relevance, and readability.Materials and MethodsTen LLM-generated scientific summaries and PLS from the INSIDE (artificial INtelligence to Support Informed DEcision making) prostate cancer dataset. For comparison, expert-written PLS from the BioLaySumm dataset were used. A panel of 5 LLMs and 3 human experts verified faithfulness. Verification was performed on original facts and facts modified with varying levels of error (subtle, moderate, contradictory). Readability was assessed using Flesch-Kincaid Reading Ease (FRE) scores.ResultsFact verification against the summaries was ∼100%, confirming accurate fact extraction. LLM panel vs human panel agreement was substantial (kappa 0.67), outperforming agreement among the interhuman (0.43 [95% CI, 0.34–0.52]) and inter-LLM (0.40 [0.38–0.42]) panels. Large language model scientific summaries showed high faithfulness (88.9% [88.0–89.8]) and low hallucinations (9.6% [6.5–12.7]) compared to human-written PLS (61.6% [60.1–63.1] faithfulness; 40.6% [37.8– 43.4] hallucinations). The LLMs detected errors sensitively with scores decreasing as fact modifications became more severe. Finally, LLM-generated PLS were more readable than human-written versions (FRE 42.3 [interquartile range, IQR 35.27–49.41] vs 28.8 [IQR 21.02–36.18]).DiscussionA panel of LLMs reliably assessed the faithfulness of scientific summaries to their original source and thus can help increase reliability for clinical use. The lower faithfulness in human-written PLS likely reflects extrinsic hallucinations added for context.ConclusionThe study demonstrates a novel approach to automatically assess the quality and trustworthiness of LLM-generated scientific and PLS via faithfulness, relevance, and readability.

  • New
  • Research Article
  • 10.1016/j.msksp.2026.103601
The musculoskeletal pain literacy questionnaire (MSK-PLq) - Part 1: Development of a preliminary version through a systematic review and Delphi consensus.
  • Jun 16, 2026
  • Musculoskeletal science & practice
  • Pablo Bellosta-López + 17 more

The musculoskeletal pain literacy questionnaire (MSK-PLq) - Part 1: Development of a preliminary version through a systematic review and Delphi consensus.

  • New
  • Research Article
  • 10.1016/j.eprac.2026.06.003
Multimodal evaluation of differentiated thyroid cancer patient information booklets: strengths and limitations in content quality, readability, and patient-centered communication.
  • Jun 15, 2026
  • Endocrine practice : official journal of the American College of Endocrinology and the American Association of Clinical Endocrinologists
  • Grégoire Racine + 7 more

Multimodal evaluation of differentiated thyroid cancer patient information booklets: strengths and limitations in content quality, readability, and patient-centered communication.

  • New
  • Research Article
  • 10.1016/j.jfo.2026.104923
Intravitreal anti-VEGF therapy: Comparative evaluation of appropriateness and readability of large language model chatbots' responses to frequently asked patient questions.
  • Jun 15, 2026
  • Journal francais d'ophtalmologie
  • T Sezer + 2 more

Intravitreal anti-VEGF therapy: Comparative evaluation of appropriateness and readability of large language model chatbots' responses to frequently asked patient questions.

  • 1
  • 2
  • 3
  • 4
  • 5
  • 6
  • .
  • .
  • .
  • 10
  • 1
  • 2
  • 3
  • 4
  • 5

Popular topics

  • Latest Artificial Intelligence papers
  • Latest Nursing papers
  • Latest Psychology Research papers
  • Latest Sociology Research papers
  • Latest Business Research papers
  • Latest Marketing Research papers
  • Latest Social Research papers
  • Latest Education Research papers
  • Latest Accounting Research papers
  • Latest Mental Health papers
  • Latest Economics papers
  • Latest Education Research papers
  • Latest Climate Change Research papers
  • Latest Mathematics Research papers

Most cited papers

  • Most cited Artificial Intelligence papers
  • Most cited Nursing papers
  • Most cited Psychology Research papers
  • Most cited Sociology Research papers
  • Most cited Business Research papers
  • Most cited Marketing Research papers
  • Most cited Social Research papers
  • Most cited Education Research papers
  • Most cited Accounting Research papers
  • Most cited Mental Health papers
  • Most cited Economics papers
  • Most cited Education Research papers
  • Most cited Climate Change Research papers
  • Most cited Mathematics Research papers

Latest papers from journals

  • Scientific Reports latest papers
  • PLOS ONE latest papers
  • Journal of Clinical Oncology latest papers
  • Nature Communications latest papers
  • BMC Geriatrics latest papers
  • Science of The Total Environment latest papers
  • Medical Physics latest papers
  • Cureus latest papers
  • Cancer Research latest papers
  • Chemosphere latest papers
  • International Journal of Advanced Research in Science latest papers
  • Communication and Technology latest papers

Latest papers from institutions

  • Latest research from French National Centre for Scientific Research
  • Latest research from Chinese Academy of Sciences
  • Latest research from Harvard University
  • Latest research from University of Toronto
  • Latest research from University of Michigan
  • Latest research from University College London
  • Latest research from Stanford University
  • Latest research from The University of Tokyo
  • Latest research from Johns Hopkins University
  • Latest research from University of Washington
  • Latest research from University of Oxford
  • Latest research from University of Cambridge

Popular Collections

  • Research on Reduced Inequalities
  • Research on No Poverty
  • Research on Gender Equality
  • Research on Peace Justice & Strong Institutions
  • Research on Affordable & Clean Energy
  • Research on Quality Education
  • Research on Clean Water & Sanitation
  • Research on COVID-19
  • Research on Monkeypox
  • Research on Medical Specialties
  • Research on Climate Justice
Discovery logo
FacebookTwitterLinkedinInstagram

Download the FREE App

  • Play store Link
  • App store Link
  • Scan QR code to download FREE App

    Scan to download FREE App

  • Google PlayApp Store
FacebookTwitterTwitterInstagram
  • Universities & Institutions
  • Publishers
  • R Discovery PrimeNew
  • Ask R Discovery
  • Blog
  • Accessibility
  • Topics
  • Journals
  • Open Access Papers
  • Year-wise Publications
  • Recently published papers
  • Pre prints
  • Questions
  • FAQs
  • Contact us
Lead the way for us

Your insights are needed to transform us into a better research content provider for researchers.

Share your feedback here.

FacebookTwitterLinkedinInstagram
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.

Privacy PolicyCookies PolicyTerms of UseCareers