Discovery Logo
Sign In
Search
Paper
Search Paper
R Discovery for Libraries Pricing Sign In
  • Home iconHome
  • My Feed iconMy Feed
  • Search Papers iconSearch Papers
  • Library iconLibrary
  • Explore iconExplore
  • Ask R Discovery iconAsk R Discovery Star Left icon
  • Literature Review iconLiterature Review NEW
  • Chat PDF iconChat PDF Star Left icon
  • Citation Generator iconCitation Generator
  • Chrome Extension iconChrome Extension
    External link
  • Use on ChatGPT iconUse on ChatGPT
    External link
  • iOS App iconiOS App
    External link
  • Android App iconAndroid App
    External link
  • Contact Us iconContact Us
    External link
  • Paperpal iconPaperpal
    External link
  • Mind the Graph iconMind the Graph
    External link
  • Journal Finder iconJournal Finder
    External link
Discovery Logo menuClose menu
  • Home iconHome
  • My Feed iconMy Feed
  • Search Papers iconSearch Papers
  • Library iconLibrary
  • Explore iconExplore
  • Ask R Discovery iconAsk R Discovery Star Left icon
  • Literature Review iconLiterature Review NEW
  • Chat PDF iconChat PDF Star Left icon
  • Citation Generator iconCitation Generator
  • Chrome Extension iconChrome Extension
    External link
  • Use on ChatGPT iconUse on ChatGPT
    External link
  • iOS App iconiOS App
    External link
  • Android App iconAndroid App
    External link
  • Contact Us iconContact Us
    External link
  • Paperpal iconPaperpal
    External link
  • Mind the Graph iconMind the Graph
    External link
  • Journal Finder iconJournal Finder
    External link
features
  • Audio Papers iconAudio Papers
  • Paper Translation iconPaper Translation
  • Chrome Extension iconChrome Extension
Content Type
  • Journal Articles iconJournal Articles
  • Conference Papers iconConference Papers
  • Preprints iconPreprints
  • Seminars by Cassyni iconSeminars by Cassyni
More
  • R Discovery for Libraries iconR Discovery for Libraries
  • Research Areas iconResearch Areas
  • Topics iconTopics
  • Resources iconResources

Articles published on human-written-summaries

Authors
Select Authors
Journals
Select Journals
Duration
Select Duration
29 Search results
Sort by
Recency
  • Research Article
  • 10.1016/j.jbi.2026.105009
Enhancing adverse drug event extraction and summarization for cancer drugs through large language models.
  • Jun 1, 2026
  • Journal of biomedical informatics
  • Sofia Jamil + 2 more

Enhancing adverse drug event extraction and summarization for cancer drugs through large language models.

  • Research Article
  • 10.2196/71230
Clinical Summaries of Social Media Timelines for Mental Health Monitoring: Human Versus Large Language Model Comparative Evaluation Study.
  • Mar 27, 2026
  • JMIR formative research
  • Ayal Klein + 7 more

Social media timelines contain rich signals of users' mental states but are too voluminous for direct clinical review. Although large language models (LLMs) demonstrate robust linguistic and summarization capabilities in general‑purpose tasks, distilling clinically relevant insights demands deeper psychological analysis and sensitivity to each individual's unique personality and context. Accurately capturing subtle, personalized affective and behavioral patterns remains a significant challenge for current models. A thorough, systematic evaluation of LLM‑generated clinical summaries is therefore essential to understand their readiness for real‑world mental health monitoring. This study evaluates the ability of an LLM-based pipeline to generate clinically meaningful summaries of social media timelines, compared to summaries written by human clinicians. The summaries are structured along 3 key clinical aspects, including an overall mental health assessment, intrapersonal and interpersonal patterns, and mental state changes over time. We use a recent state-of-the-art approach that combines a hierarchical variational autoencoder (VAE) with an LLM (Large Language Model-Meta AI 2 13-billion-parameter version; LLaMA2 13B). This method first summarizes the patient's history using the VAE and then transforms this summary into a clinical narrative using the LLM. We also test both single-step and multistep LLM-prompting techniques and devise comprehensive clinical prompts. For 30 social media timelines, model outputs were evaluated against human-written summaries through human ratings and expert qualitative analysis. Linguistic diversity was automatically measured as a proxy for personalization. Human summaries scored highest for factual consistency (3.75) and general usefulness (3.63). The timeline-hierarchical variational autoencoder (TH-VAE) model outperformed LLaMA for factual consistency (3.35 vs 3.08) and general usefulness (3.28 vs 3.38). Both 2-step models were comparable to humans in describing interpersonal and intrapersonal patterns (3.45-3.48 vs 3.33) and changes over time (3.42 vs 3.35-3.30). The naive LLaMA baseline scored lower on all criteria except factual consistency. Furthermore, a qualitative analysis observed that human summaries provided more accurate, deep, and personalized insights, while LLMs offered more exhaustive but generic descriptions. Quantitatively, linguistic diversity was higher in human summaries both at the semantic level (mean Cohen d=1.19) and at the surface level (mean Cohen d=1.31). At this time medium-size LLMs can generate largely accurate and informative clinical summaries of social media timelines, and advanced prompting boosts performance modestly. However, at the time of this writing, they underperform human clinicians in capturing subtle psychological nuances and individual idiosyncrasies. Future work should integrate domain‑specific fine‑tuning and enhanced context modeling to improve LLM clinical fidelity.

  • Research Article
  • 10.4103/cmrp.cmrp_163_25
Readability and quality assessment of human versus artificial intelligence-generated plain language summaries across six large language models
  • Jan 1, 2026
  • Current Medicine Research and Practice
  • Aryan Jain + 1 more

Background: Plain language summaries (PLS) aim to make scientific research understandable to non-specialists, yet producing clear and accurate summaries remains challenging, especially for non-native English writers. With advances in large language models (LLMs), automated summarisation offers a potential solution. However, few studies have directly compared human-written PLS with outputs from multiple LLMs using the same dataset. Aims: To assess whether LLMs can support or augment human efforts to produce effective and accessible scientific communication. Materials and Methods: In this cross-sectional study, 30 human-written PLS were compared with 180 PLS generated by six LLMs. Readability was assessed using Flesch Reading Ease, Flesch–Kincaid Grade Level, sentence length and syllables per word. Three independent reviewers evaluated clarity, inclusiveness, interpretation and factual accuracy. Group differences were analysed using one-way ANOVA with post hoc Tukey testing. Results: LLM-generated summaries were significantly more readable than human-written summaries across all metrics ( P < 0.001). Human-authored PLS showed marginally higher factual accuracy, though overall reviewer-rated quality did not differ significantly. Among the LLMs, Gemini produced the simplest text, whereas Meta Artificial intelligence demonstrated the best balance of readability and quality. Conclusion: LLMs can generate PLS that are comparable in quality to human-written summaries while offering substantially improved readability. These tools may enhance accessibility for diverse audiences, though human oversight remains essential to ensure contextual accuracy and interpretive depth.

  • Research Article
  • Cite Count Icon 4
  • 10.1016/j.eswa.2025.129001
Stylometry recognizes human and LLM-generated texts in short samples
  • Jan 1, 2026
  • Expert Systems with Applications
  • Karol Przystalski + 3 more

• Application of Stylometry to differentiate texts generated by LLMs, • Insights into LLM and human text characteristics, • Using SHAP for LLM text generated stylometry analysis The paper explores stylometry as a method to distinguish between texts created by Large Language Models (LLMs) and humans, addressing issues of model attribution, intellectual property, and ethical AI use. Stylometry has been used extensively to characterise the style and attribute authorship of texts. By applying it to LLM-generated texts, we identify their emergent writing patterns. The paper involves creating a benchmark dataset based on Wikipedia, with (a) human-written term summaries, (b) texts generated purely by LLMs (GPT-3.5/4, LLaMa 2/3, Orca, and Falcon), (c) processed through multiple text summarisation methods (T5, BART, Gensim, and Sumy), and (d) rephrasing methods (Dipper, T5). The 10-sentence long texts were classified by tree-based models (decision trees and LightGBM) using human-designed (StyloMetrix) and n-gram-based (our own pipeline) stylometric features that encode lexical, grammatical, syntactic, and punctuation patterns. The cross-validated results reached a performance of up to.87 Matthews correlation coefficient in the multiclass scenario with 7 classes, and accuracy between.79 and 1. in binary classification, with the particular example of Wikipedia and GPT-4 reaching up to.98 accuracy on a balanced dataset. Shapley Additive Explanations pinpointed features characteristic of the encyclopaedic text type, individual overused words, as well as a greater grammatical standardisation of LLMs with respect to human-written texts. These results show – crucially, in the context of the increasingly sophisticated LLMs – that it is possible to distinguish machine- from human-generated texts at least for a well-defined text type

  • Research Article
  • 10.1177/08944393251409744
Generating the Past: How Artificial Intelligence Summaries of Historical Events Affect Knowledge
  • Dec 24, 2025
  • Social Science Computer Review
  • Daniel Karell + 3 more

Many people now use AI chatbots to obtain summaries of complex topics, yet we know little about how this affects knowledge acquisition, including how the effects might vary across different groups of people. We conducted two experiments comparing how well people recalled factual information after reading AI-generated or human-written historical summaries. Participants who read AI-generated summaries scored significantly higher on knowledge tests than those who read expert-written blog posts (Study 1) or Wikipedia articles (Study 2). These improvements were present regardless of whether readers knew the content was AI-generated or if the AI summaries were politically biased. Moreover, AI summaries improved recall across various demographic groups, including gender, race, income, education, and digital literacy levels. This suggets that using AI tools for everyday factual queries does not create new knowledge inequalities but could still amplify existing ones through differential access. Our findings indicate that the increasingly routine use of AI for information-seeking could enhance factual learning, with implications for education policy and addressing inequality.

  • Research Article
  • Cite Count Icon 1
  • 10.1016/j.jbi.2025.104949
Scalable Scientific Interest Profiling Using Large Language Models
  • Nov 1, 2025
  • Journal of biomedical informatics
  • Yilun Liang + 8 more

Scalable Scientific Interest Profiling Using Large Language Models

  • Research Article
  • 10.1080/10447318.2025.2574513
Skim or Swim? Investigating AI-Generated Summaries, Trust, and Comprehension in Chinese Mobile Reading
  • Oct 17, 2025
  • International Journal of Human–Computer Interaction
  • Guangrui Fan + 2 more

As mobile reading surges in China, platforms increasingly deploy AI-generated summaries to mitigate information overload. This study examines, among Chinese mobile readers, how summary type (AI vs. human vs. none) shapes engagement (reading time, click-through), comprehension, and trust, and when skepticism toward AI translates into opening the full article. We conducted an explanatory, mixed-methods, within-subjects experiment ( N = 47 ) in which adults in China read six Chinese-language news articles (technology, health, lifestyle) on their own smartphones under three conditions: AI summary + “Read Full,” Human summary + “Read Full,” and No summary (full text). We logged reading time, click-through, and scroll depth; administered five-item comprehension quizzes and trust/completeness scales after each article; analyzed outcomes using repeated-measures ANOVA, linear mixed-effects models, and logistic regression; and thematically analyzed open-ended responses. Human-written summaries received higher trust and perceived completeness than AI-generated summaries, yet click-through rates were comparable (AI: 81.9% vs. Human: 72.3%). Providing the full text upfront yielded the highest comprehension; choosing to open the full article after viewing a summary largely eliminated the comprehension deficit. Topic relevance moderated the relationship between trust and click-through (lower trust increased opening only when relevance was high), and domain knowledge attenuated the comprehension penalty of summary-only reading. In on-the-go mobile contexts, an attitude-behavior gap emerges: readers who express low trust in AI summaries often still rely on them when perceived time costs of full-article reading are high. For mobile HCI and news design, clearly label AI provenance with brief coverage disclaimers, offer adaptive/layered summaries that expand based on interest signals or in high-stakes domains, and provide strategic verification prompts that nudge readers to open the full article when depth and accuracy matter.

  • Research Article
  • 10.59934/jaiea.v5i1.1395
Automatic Criminal News Summarization System with Extractive Method Based on Latent
  • Oct 15, 2025
  • Journal of Artificial Intelligence and Engineering Applications (JAIEA)
  • Christ Chandra

The rapid growth of digital information demands automatic systems to help users efficiently extract the core of information, especially in criminal news which often attracts significant public attention. This study aims to design and develop an automatic summarization system for criminal news using an extractive method based on Latent Semantic Analysis (LSA). In the process, textual features are first extracted using the Term Frequency-Inverse Document Frequency (TF-IDF) method to weigh the importance of each word in the document. The resulting TF-IDF matrix is then used as input for LSA to model semantic relationships between sentences and identify those most representative of the document content. The dataset consists of Indonesian-language criminal news articles collected from various online news portals. The system is evaluated by comparing the automatically generated summaries with human-written summaries using the ROUGE metric. The experimental results show that the combination of TF-IDF and LSA can generate informative and relevant summaries, achieving a ROUGE-1 score of 0.72. This system is expected to help users understand news content quickly and efficiently.

  • Research Article
  • Cite Count Icon 2
  • 10.1002/jhm.70163
Quality assessment of artificial intelligence-generated versus human-written hospital summaries evaluating detail, usefulness, and continuity of care.
  • Sep 30, 2025
  • Journal of hospital medicine
  • Douglas Challener + 4 more

Hospital discharge summaries are critical for ensuring continuity of care, but their quality often varies. Large language models (LLMs) have the potential to standardize and enhance the efficiency of this documentation process. To evaluate the quality of hospital discharge summaries created by an LLM-based hospital course drafting tool created by Epic Systems compared with human-written summaries. Retrospective study at a single tertiary-care institution in 2024. The cohort included 100 adult hospitalizations lasting >72 h across medical and surgical dismissing services. No interventions were performed. Summaries (LLM-generated vs. human-written) were independently reviewed using a standardized rubric covering nine domains (e.g., comprehensiveness, clarity, relevance). Scores were normalized and compared. Readability was assessed using Flesch Reading Ease. LLM-generated summaries outperformed human-written summaries across all criteria (p < .05), with the greatest difference observed in comprehensiveness (LLM median 0.62 vs. human -0.23). Human-written summaries from surgical services scored lower than those from medical services, but LLM performance was consistent across both. Human summaries had higher Flesch Reading Ease scores (33.11 vs. 26.2; p < .05), reflecting simpler language. LLM-generated summaries demonstrated superior quality, consistency, and clinical utility compared with human-written summaries, highlighting their potential to improve documentation efficiency and standardization.

  • Research Article
  • 10.2196/68511
Collecting and Sharing Person-Centered AI Clinical Summaries Across Frailty Services Provided by the National Health Service and Voluntary, Community, and Social Enterprise: Protocol for a Co-Design and Feasibility Study
  • Sep 17, 2025
  • JMIR Research Protocols
  • Kieran Green + 4 more

BackgroundDue to its association with multimorbidity, frailty gives rise to multidimensional needs for different services. Too often, patient preferences and service encounter information are not adequately shared.ObjectiveThis developmental study aims to co-design, collect, and analyze encounter data from multiple community and primary-based multidisciplinary teams (MDTs) providing services for people with frailty to develop prototype large language models that can generate clinical and person-centered care summaries.MethodsEngaging stakeholders in 2 primary care networks, we will co-design the large language model to ensure it meets local needs and preferences as well as infrastructure, information governance, and regulation requirements. General practitioners will identify 50 patients with frailty requiring MDT engagement. Three consecutive encounters between the patients and different members of MDTs will then be audio-recorded. Recordings will be transcribed into text for concept design and model pretraining. These data combine stakeholder engagement insights to develop sensitive artificial intelligence (AI) models responding to stakeholders’ needs, workflows, and preferences. To generate the person-centered summaries, we will test 2 approaches to modeling the encounter data: graph-based modeling and hierarchical transformers. The AI-generated summaries will be compared to human-written summaries of the same encounter data and assessed for accuracy, quality, fluency, and person-centeredness. They will also be shared with the original MDT members for validation. We will capture inputs, processes, and outcomes across all key phases of the implementation journey to identify capability requirements, determinants of implementation (including key challenges and best practices to overcome them), and the value added by the technology.ResultsThis protocol aims to review implementation evidence and engage stakeholders in co-design. This work package will aid the development of contextually sensitive, longitudinal, and AI-generated person-centered summarization tools. Model development will aim to achieve longitudinal person-centered summaries tested against MDT standards. If deemed suitable for deployment, optimum ways of integrating these summaries into shared care records will be explored with local key system leaders. Model evaluations will provide conclusive insights into such technologies’ benefits and risks. As of August 2025, this study has not yet been funded, nor has ethical approval for the project been obtained. Consequently, dates of data collection and numbers of recruited participants are not applicable at this time.ConclusionsOur protocol provides a robust method of co-designing, evaluating, and implementing a longitudinal AI medical summary tool. Including key stakeholders at multiple stages facilitates an iterative development strategy that is designed to solve implementation challenges as they emerge. This project fits within our long-term vision to deliver a multimodal AI tool that saves clinicians time and deepens the health care professional–patient relationship. Future studies should include a larger patient sample, video-recorded health care professional–patient encounters, and a more extensive longitudinal evaluation.International Registered Report Identifier (IRRID)PRR1-10.2196/68511

  • Research Article
  • 10.1162/tacl.a.30
Explanatory Summarization with Discourse-Driven Planning
  • Sep 9, 2025
  • Transactions of the Association for Computational Linguistics
  • Dongqi Liu + 3 more

Abstract Lay summaries for scientific documents typically include explanations to help readers grasp sophisticated concepts or arguments. However, current automatic summarization methods do not explicitly model explanations, which makes it difficult to align the proportion of explanatory content with human-written summaries. In this paper, we present a plan-based approach that leverages discourse frameworks to organize summary generation and guide explanatory sentences by prompting responses to the plan. Specifically, we propose two discourse-driven planning strategies, where the plan is conditioned as part of the input or part of the output prefix, respectively. Empirical experiments on three lay summarization datasets show that our approach outperforms existing state-of-the-art methods in terms of summary quality, and it enhances model robustness, controllability, and mitigates hallucination. The project information is available at https://dongqi.me/projects/ExpSum.

  • Research Article
  • 10.11922/11-6035.csd.2025.0007.zh
BodSUM-6000: A dataset of abstractive Tibetan text summarization
  • Sep 6, 2025
  • China Scientific Data
  • Wuji Xia + 2 more

<p indent="0mm">The dataset is the foundation for automatic text summarization, and its quality directly determines the performance of summarization systems. Currently, research on Tibetan text summarization suffers from a lack of publicly available, high-quality open-source datasets, which has constrained progress in this field. To fill this gap, news texts were collected from various Tibetan websites, and human-written summaries were created for each text. Additionally, both subjective and objective evaluation methods were used to comprehensively assess the matching between Tibetan news texts and their corresponding human-written summaries, ensuring the quality of the data. Ultimately, BodSUM-6000, a dataset of high-quality ive Tibetan text summarization consisting of 6,000 text-summary pairs, was constructed. This dataset provides valuable data resources for training and evaluating abstractive Tibetan text summarization models, helping to address the limitations in the generalizability and accuracy of current Tibetan text summarization techniques, and promoting the further development of Tibetan automatic summarization research.

  • Research Article
  • Cite Count Icon 6
  • 10.1177/07342829251346623
Human vs. Machine: Comparing AI-Generated and Human-Written Psychological Reports
  • Jun 12, 2025
  • Journal of Psychoeducational Assessment
  • Adam Lockwood + 4 more

This study examines the effectiveness of artificial intelligence (AI) in psychological report writing by comparing reports generated by human psychologists with those produced by OpenAI’s Generative Pre-trained Transformer Version 4 (ChatGPT-4). A total of 249 licensed psychologists evaluated the reports based on overall quality, readability, writing style, organization, summary quality, recommendations, preference, and willingness to sign off on the reports. Although human-generated reports were generally rated more favorably and participants expressed greater comfort in approving them, effect sizes were typically small. Two exceptions were noted: moderate effect sizes were found in favor of human-written summaries, while AI-generated reports showed moderate effect sizes for the quality of their recommendations. These findings suggest that AI shows potential for augmenting report writing. Comprehensive guidelines are necessary for the ethical and effective integration of AI into psychological practice. Further research is needed to enhance our understanding of AI’s role and capabilities in psychological assessment and reporting.

  • Research Article
  • Cite Count Icon 3
  • 10.3390/electronics14112115
ChatGPT vs. Human Journalists: Analyzing News Summaries Through BERTScore and Moderation Standards
  • May 22, 2025
  • Electronics
  • Hui-Sang Kim + 2 more

Recent advances in natural language processing (NLP) have enabled the development of powerful language models such as Generative Pre-trained Transformers (GPTs). This study evaluates the performance of ChatGPT in generating news summaries by comparing them with summaries written by professional journalists at The New York Times. Using BERTScore as the primary metric, we assessed the semantic similarity between ChatGPT-generated and human-authored summaries. We further employed OpenAI’s moderation API to examine the extent to which each set of summaries contained potentially biased, inflammatory, or violent language. The results indicate that ChatGPT-generated summaries exhibit a high degree of contextual alignment with human-written summaries, achieving a BERTScore F1-score above 0.87. Moreover, ChatGPT outputs consistently omit language flagged as problematic by moderation algorithms, producing summaries that are less likely to include harmful or polarizing content—a feature we define as moderation-friendly summarization. These findings suggest that ChatGPT can serve as a valuable tool for automated news summarization, offering content that is both contextually accurate and aligned with content moderation standards, thereby supporting more objective and responsible news dissemination.

  • Abstract
  • Cite Count Icon 1
  • 10.1017/cts.2024.999
376 Using a large language model to create lay summaries of clinical study descriptions
  • Apr 1, 2025
  • Journal of Clinical and Translational Science
  • Rebecca E Kaiser + 7 more

Objectives/Goals: We assessed the feasibility of using a large language model (LLM) to create lay language descriptions of study protocols for recruitment, which has the potential to improve accessibility and transparency of clinical studies and enable participants to make informed decisions. Methods/Study Population: All studies from a clinical research recruitment platform were included, which features human-written lay descriptions and titles for study recruitment. Corresponding protocol summaries in the IRB system were extracted and translated into lay language using a LLM (gpt-35-turbo-0613). A subset was used to develop prompt variations through an iterative process. Prompt strategies evaluated include chain-of-thought and few-shot prompting techniques. LLM-generated and human-written descriptions were compared for readability using Flesch–Kincaid and Simple Measure of Gobbledygook (SMOG) reading grade levels and information completeness using Word Movers’ Distance (WMD). Results/Anticipated Results: A total of 55 study descriptions were included – 10 were used to develop prompts and 45 were used for evaluation. The final LLM instructions included multistep prompts. The LLM was first instructed to produce a two- to three-sentence long description without using scientific jargon and included two pairs of examples. The LLM was then asked to shorten the description and finally to provide an engaging title. LLM-generated and human-written summaries were similar in length (median (IQR) 328 (278.5–360.5) vs. 342 (203–532.5) characters, respectively). LLM-generated summaries had lower Flesch-Kincaid grade level (5.15 vs. 8.28, p Discussion/Significance of Impact: An LLM can be used to generate lay language summaries that are readable at a lower grade level while maintaining semantic similarity. This approach can be used to improve the drafting of summaries for recruitment, thereby improving accessibility to potential participants. Future work includes human evaluation and implementation into practice.

  • Research Article
  • Cite Count Icon 3
  • 10.34010/komputa.v13i1.11197
Peringkasan Teks Otomatis Abstraktif Menggunakan Transformer Pada Teks Bahasa Indonesia
  • Apr 27, 2024
  • Komputa : Jurnal Ilmiah Komputer dan Informatika
  • Andika Bahari + 1 more

Abstractive text summarization is used to generate summaries that are similar to human-written summaries. To achieve that capability, recurrent deep learning architectures were usually applied, such as RNN, LSTM, and GRU. In previous studies about abstractive text summarization in Indonesian, recurrent models were widely used and there were cohesion and grammar errors that appeared in the generated summaries—this could have an impact on performance. Currently, there is Transformer, a relatively new architecture that relies on the attention mechanism entirely. Due to its non-recurrent nature, Transformer overcomes the problem of dependency the on hidden states that occurs in recurrent models and it can retain information on all input sequences. In this study, we use Transformer to evaluate how good it is at abstractive text summarization in Indonesian. The training was conducted using the pre-trained T5 model with IndoSum dataset which contains around 19K news-summary pairs. We achieved evaluation scores of 0.61 ROUGE-1 and 0.51 ROUGE-2.

  • Open Access Icon
  • Research Article
  • Cite Count Icon 1
  • 10.3934/aci.2024001
Linguistic summarisation of multiple entities in RDF graphs
  • Jan 1, 2024
  • Applied Computing and Intelligence
  • Elizaveta Zimina + 5 more

&lt;abstract&gt;&lt;p&gt;Methods for producing summaries from structured data have gained interest due to the huge volume of available data in the Web. Simultaneously, there have been advances in natural language generation from Resource Description Framework (RDF) data. However, no efforts have been made to generate natural language summaries for groups of multiple RDF entities. This paper describes the first algorithm for summarising the information of a set of RDF entities in the form of human-readable text. The paper also proposes an experimental design for the evaluation of the summaries in a human task context. Experiments were carried out comparing machine-made summaries and summaries written by humans, with and without the help of machine-made summaries. We develop criteria for evaluating the content and text quality of summaries of both types, as well as a function measuring the agreement between machine-made and human-written summaries. The experiments indicated that machine-made natural language summaries can substantially help humans in writing their own textual descriptions of entity sets within a limited time.&lt;/p&gt;&lt;/abstract&gt;

  • Research Article
Towards Inpatient Discharge Summary Automation via Large Language Models: A Multidimensional Evaluation with a HIPAA-Compliant Instance of GPT-4o and Clinical Expert Assessment.
  • Jan 1, 2024
  • AMIA ... Annual Symposium proceedings. AMIA Symposium
  • Tyler Osborne + 9 more

Large language models (LLMs) have demonstrated potential to automate clinical documentation tasks that may reduce clinician burden, such as generation of hospital discharge summaries. Prior research used older LLMs and limited data, raising concerns about fabrications and omissions. In this study, we evaluated the automatic generation of inpatient Internal Medicine discharge summaries using a HIPAA-compliant Microsoft Azure instance of OpenAI's GPT-4o. Both human-written and AI-generated discharge summaries were scored by Internal Medicine hospital faculty for quality, readability/conciseness, factuality and completeness, presence of hallucinations/omissions and their impact on safety, and compared with the actual discharge summaries. Our results showed that the AI-generated discharge summaries significantly outperformed actual human written summaries in both quality and readability/conciseness and were comparable to humans in factuality and completeness, with a minimal cost.

  • Research Article
  • 10.61784/wjit3001
SUMMAGAN: ENHANCING WEB NEWS SUMMARIZATION THROUGH GENERATIVE ADVERSARIAL NETWORKS
  • Jan 1, 2024
  • World Journal of Information Technology
  • Matthew Walkman + 1 more

This paper introduces SummaGAN, a novel application of Generative Adversarial Networks (GANs) for text summarization. Unlike traditional summarization methods that rely on extractive techniques, SummaGAN uses adversarial learning to generate coherent and contextually accurate summaries. The model includes a transformer-based generator that creates summaries and a discriminator that evaluates their quality, guiding the generator to produce outputs that closely mimic human-written summaries. A large, diverse dataset of over 100,000 articles from domains such as news, scientific literature, and blogs was used to train and fine-tune the model. Experimental results show that SummaGAN significantly outperforms existing baseline models, including traditional extractive summarizers and advanced abstractive models, across multiple evaluation metrics such as ROUGE, BLEU, METEOR, and the newly introduced Coherence and Consistency Score (CCS). SummaGAN achieved a 15% improvement in ROUGE-1 scores and a 20% enhancement in BLEU scores, indicating better summary relevance and fluency. The CCS metric highlights SummaGAN's superior ability to maintain the logical flow and factual accuracy of the source text. This research demonstrates the potential of GANs to address challenges in text summarization, such as redundancy and loss of meaning, through dynamic adversarial learning. The integration of GANs with transformer architectures presents a robust framework for future NLP advancements. Future research will explore scaling the model for larger datasets, applying it in multilingual contexts, and refining the adversarial training process for improved efficiency and performance.

  • Research Article
  • Cite Count Icon 19
  • 10.1109/tse.2023.3279774
Function Call Graph Context Encoding for Neural Source Code Summarization
  • Sep 1, 2023
  • IEEE Transactions on Software Engineering
  • Aakash Bansal + 4 more

Source code summarization is the task of writing natural language descriptions of source code. The primary use of these descriptions is in documentation for programmers. Automatic generation of these descriptions is a high value research target due to the time cost to programmers of writing these descriptions themselves. In recent years, a confluence of software engineering and artificial intelligence research has made inroads into automatic source code summarization through applications of neural models of that source code. However, an Achilles' heel to a vast majority of approaches is that they tend to rely solely on the context provided by the source code being summarized. But empirical studies in program comprehension are quite clear that the information needed to describe code much more often resides in the context in the form of Function Call Graph surrounding that code. In this paper, we present a technique for encoding this call graph context for neural models of code summarization. We implement our approach as a supplement to existing approaches, and show statistically significant improvement over existing approaches. In a human study with 20 programmers, we show that programmers perceive generated summaries to generally be as accurate, readable, and concise as human-written summaries.

  • 1
  • 2
  • 1
  • 2

Popular topics

  • Latest Artificial Intelligence papers
  • Latest Nursing papers
  • Latest Psychology Research papers
  • Latest Sociology Research papers
  • Latest Business Research papers
  • Latest Marketing Research papers
  • Latest Social Research papers
  • Latest Education Research papers
  • Latest Accounting Research papers
  • Latest Mental Health papers
  • Latest Economics papers
  • Latest Education Research papers
  • Latest Climate Change Research papers
  • Latest Mathematics Research papers

Most cited papers

  • Most cited Artificial Intelligence papers
  • Most cited Nursing papers
  • Most cited Psychology Research papers
  • Most cited Sociology Research papers
  • Most cited Business Research papers
  • Most cited Marketing Research papers
  • Most cited Social Research papers
  • Most cited Education Research papers
  • Most cited Accounting Research papers
  • Most cited Mental Health papers
  • Most cited Economics papers
  • Most cited Education Research papers
  • Most cited Climate Change Research papers
  • Most cited Mathematics Research papers

Latest papers from journals

  • Scientific Reports latest papers
  • PLOS ONE latest papers
  • Journal of Clinical Oncology latest papers
  • Nature Communications latest papers
  • BMC Geriatrics latest papers
  • Science of The Total Environment latest papers
  • Medical Physics latest papers
  • Cureus latest papers
  • Cancer Research latest papers
  • Chemosphere latest papers
  • International Journal of Advanced Research in Science latest papers
  • Communication and Technology latest papers

Latest papers from institutions

  • Latest research from French National Centre for Scientific Research
  • Latest research from Chinese Academy of Sciences
  • Latest research from Harvard University
  • Latest research from University of Toronto
  • Latest research from University of Michigan
  • Latest research from University College London
  • Latest research from Stanford University
  • Latest research from The University of Tokyo
  • Latest research from Johns Hopkins University
  • Latest research from University of Washington
  • Latest research from University of Oxford
  • Latest research from University of Cambridge

Popular Collections

  • Research on Reduced Inequalities
  • Research on No Poverty
  • Research on Gender Equality
  • Research on Peace Justice & Strong Institutions
  • Research on Affordable & Clean Energy
  • Research on Quality Education
  • Research on Clean Water & Sanitation
  • Research on COVID-19
  • Research on Monkeypox
  • Research on Medical Specialties
  • Research on Climate Justice
Discovery logo
FacebookTwitterLinkedinInstagram

Download the FREE App

  • Play store Link
  • App store Link
  • Scan QR code to download FREE App

    Scan to download FREE App

  • Google PlayApp Store
FacebookTwitterTwitterInstagram
  • Universities & Institutions
  • Publishers
  • R Discovery PrimeNew
  • Ask R Discovery
  • Blog
  • Accessibility
  • Topics
  • Journals
  • Open Access Papers
  • Year-wise Publications
  • Recently published papers
  • Pre prints
  • Questions
  • FAQs
  • Contact us
Lead the way for us

Your insights are needed to transform us into a better research content provider for researchers.

Share your feedback here.

FacebookTwitterLinkedinInstagram
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.

Privacy PolicyCookies PolicyTerms of UseCareers