Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Climbing towards the NLU of the universal reading of shei ‘who’

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

Abstract The wh -expressions shei ‘who’ and shenme ‘what’ in Mandarin Chinese not only convey an interrogative meaning but also exhibit existential and universal readings in specific contexts ( Huang 1982 , Cheng 1991 , 1995 , Li 1992 , Tsai 1994 , Lin 1996 , 1998 ). Focusing on the universal interpretation of shei , this paper has three objectives. First, we demonstrate that current state-of-the-art large language models (LLMs) such as ChatGPT, lack reliability in distinguishing these three distinct readings of shei . Second, we develop a specialized natural language processing and understanding (NLP/NLU) system capable of processing and interpreting shei across diverse contexts with greater accuracy, transparency, and consistency. Unlike current LLMs, our system is built upon Wang et al.’s ( 2019a , 2019b ) generative linguistics-based NLP/NLU software tools, Articut and Loki, enabling it to require significantly less training data to interpret the universal reading of shei . Third, we compare our model’s performance with that of ChatGPT, demonstrating its superior accuracy and robustness in interpreting the universal reading of shei .

Similar Papers
  • Conference Article
  • 10.1145/3711875.3729128
CrossLM: A Data-Free Collaborative Fine-Tuning Framework for Large and Small Language Models
  • Jun 23, 2025
  • Yongheng Deng + 5 more

While large language models (LLMs) are endowed with broad knowledge, their task-specific performance is often suboptimal. Fine-tuning LLMs with task-specific data from diverse nodes is necessary, but this data is typically safeguarded and not shared publicly due to privacy concerns. A common solution involves downstream nodes downloading the LLM locally and fine-tuning it with their proprietary data. However, owners often regard pre-trained LLMs as valuable assets and are reluctant to share them. Additionally, the significant computational resources required by LLMs make local fine-tuning impractical for many nodes. To mitigate these problems, this paper proposes CrossLM, a data-free collaborative fine-tuning framework for large and small language models. CrossLM enables resource-constrained nodes to train smaller language models (SLMs) using their private task-specific data. These SLMs are subsequently leveraged to promote the task-specific natural language generation and understanding capabilities of the LLMs. Simultaneously, the SLMs of nodes also benefit from enhancement by the fine-tuned LLMs. In this way, CrossLM avoids sharing private data and proprietary LLMs, and also reduces the resource requirements of nodes. Through extensive experiments across a range of benchmark tasks and popular language models, we demonstrate that CrossLM significantly boosts the task-specific performance of both LLMs and SLMs while preserving the generalization capabilities of LLMs.

  • Research Article
  • Cite Count Icon 2
  • 10.1007/s10143-025-03785-7
Current trends and future prospects of language models and processing systems in spine surgery - a scoping review.
  • Sep 5, 2025
  • Neurosurgical review
  • Vivek Sanker + 9 more

Natural language processing (NLPs) and Large language models (LLM), such as ChatGPT, represent transformative advancements in artificial intelligence (AI). Their implementation into the medical field has a broad potential, and this review discusses the current trends and prospects of NLPs and LLMs in spine surgery, assessing their potential benefits, applications, and limitations. The methodology involved a comprehensive narrative review of existing English literature related to the use of NLPs and LLMs in spine surgery. We searched the databases PubMed, EMBASE, Web of Science and Scopus from inception until 16th June 2025 using keywords evolving around LLM, natural language processing and spine surgery. Original studies, clinical reports, and case series were included, while abstracts or unpublished studies were excluded. From 221 initial records, 37 studies were included: 18 evaluated LLMs and 19 evaluated NLP-based tools. LLMs were commonly used for clinical decision-making (n = 8), patient counseling (n = 7), classification (n = 2), and in research (n = 1). NLPs were applied in classification tasks (n = 12), clinical decision-making (n = 3), patient counseling (n = 1), postoperative opioid monitoring (n = 2), and research registry development (n = 1). ChatGPT-4 achieved up to 92% accuracy in clinical recommendations, outperforming GPT-3.5 in multiple tasks. Comparative analyses have found that newer versions of LLMs, such as ChatGPT-4, outperform previous versions, evident by greater accuracy and to a lesser extent of artificial hallucination. However, limitations persist, including overconfident outputs, adherence gaps to clinical guidelines, and inconsistent patient readability. While this review suggests that NLPs and LLMs can have a significant impact on spine practice, it is important to keep their limitations in mind and implement them with caution. To maximize the benefits of these models in spine surgery, future research should focus on improving model sensitivity and specificity, promoting multi-disciplinary collaborations, and addressing ethical considerations regarding the use of language models in medical practice, including the inherent issue of hallucination of these models.

  • Dissertation
  • 10.32657/10356/184392
Towards trustworthy and reliable language models
  • Jan 1, 2025
  • Ruochen Zhao

This thesis addresses the critical challenge of developing trustworthy and reliable Natural Language Processing (NLP) systems, specifically the newly emerged Large Language Models (LLMs). As LLMs become increasingly prevalent in various domains, the need for transparent, interpretable, and controllable AI systems has never been more pressing. However, the complexity of LLMs, the compositional nature of language, and the potential for hallucinations pose significant obstacles to achieving these goals. To increase user trust of AI systems in real-life deployment, we hope to enhance the trustworthiness and reliability of LLMs without requiring model revisions or compromising performance. Motivated by this overarching goal, we delve into two main goals that enhance trustworthiness, providing user-friendly explanations of the LLM’s decisions and controlling the LLM’s behaviors. Specifically, we raise three main research questions: How can we disentangle the true reasons behind LLM decisions from the complex architecture and vast number of parameters? How can we provide user-friendly explanations for LLM generations? How can we increase LLM controllability with minimal interventions?

  • Research Article
  • 10.14569/ijacsa.2024.0151224
A Multimodal Data Scraping Tool for Collecting Authentic Islamic Text Datasets
  • Jan 1, 2024
  • International Journal of Advanced Computer Science and Applications
  • Abdallah Namoun + 2 more

Making decisions based on accurate knowledge is agreed upon to provide ample opportunities in different walks of life. Machine learning and natural language processing (NLP) systems, such as Large Language Models, may use unrecognized sources of Islamic content to fuel their predictive models, which could often lead to incorrect judgments and rulings. This article presents the development of an automated method with four distinct algorithms for text extraction from static websites, dynamic websites, YouTube videos with transcripts, and for speech-to-text conversion from videos without transcripts, particularly targeting Islamic knowledge text. The tool is tested by collecting a reliable Islamic knowledge dataset from authentic sources in Saudi Arabia. We scraped Islamic content in Arabic from text websites of prominent scholars and YouTube channels administered by five authorized agencies in Saudi Arabia. These agencies include the general authority for the affairs of the grand mosque and the prophet’s mosque and charitable foundations in Saudi Arabia. For websites, text data were scraped using Python tools for static and dynamic web scraping such as Beautiful Soup and Selenium. For YouTube channels, data were scraped from existing transcripts or transcribed using automatic speech recognition tools. The final Islamic content dataset comprises 31225 records from regulated sources. Our Islamic knowledge dataset can be used to develop accurate Islamic question answering, AI chatbots and other NLP systems.

  • Conference Article
  • Cite Count Icon 2
  • 10.3115/980691.980695
A test environment for natural language understanding systems
  • Jan 1, 1998
  • Li Li + 4 more

The Natural Language Understanding Engine Test Environment (ETE) is a GUI software tool that aids in the development and maintenance of large, modular, natural language understanding (NLU) systems. Natural language understanding systems are composed of modules (such as part-of-speech taggers, parsers and semantic analyzers) which are difficult to test individually because of the complexity of their output data structures. Not only are the output data structures of the internal modules complex, but also many thousands of test items (messages or sentences) are required to provide a reasonable sample of the linguistic structures of a single human language, even if the language is restricted to a particular domain. The ETE assists in the management and analysis of the thousands of complex data structures created during natural language processing of a large corpus using relational database technology in a network environment.

  • Video Transcripts
  • 10.48448/4pse-7755
Differentially Private Decoding in Large Language Models
  • Jul 9, 2022
  • Underline Science Inc.
  • Jimit Majmudar

Recent large-scale natural language processing (NLP) systems use a pre-trained Large Language Model (LLM) on massive and diverse corpora as a headstart. In practice, the pre-trained model is adapted to a wide array of tasks via fine-tuning on task-specific datasets. LLMs, while effective, have been shown to memorize instances of training data thereby potentially revealing private information processed during pre-training. The potential leakage might further propagate to the downstream tasks for which LLMs are fine-tuned. On the other hand, privacy-preserving algorithms usually involve retraining from scratch, which is prohibitively expensive for LLMs. In this work, we propose a simple, easy to interpret, and computationally lightweight perturbation mechanism to be applied to an already trained model at the decoding stage. Our perturbation mechanism is model-agnostic and can be used in conjunction with any LLM. We provide theoretical analysis showing that the proposed mechanism is differentially private, and experimental results showing a privacy-utility trade-off.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 20
  • 10.1038/s41598-024-66576-y
Clinical efficacy of pre-trained large language models through the lens of aphasia
  • Jul 6, 2024
  • Scientific Reports
  • Yan Cong + 2 more

The rapid development of large language models (LLMs) motivates us to explore how such state-of-the-art natural language processing systems can inform aphasia research. What kind of language indices can we derive from a pre-trained LLM? How do they differ from or relate to the existing language features in aphasia? To what extent can LLMs serve as an interpretable and effective diagnostic and measurement tool in a clinical context? To investigate these questions, we constructed predictive and correlational models, which utilize mean surprisals from LLMs as predictor variables. Using AphasiaBank archived data, we validated our models’ efficacy in aphasia diagnosis, measurement, and prediction. Our finding is that LLMs-surprisals can effectively detect the presence of aphasia and different natures of the disorder, LLMs in conjunction with the existing language indices improve models’ efficacy in subtyping aphasia, and LLMs-surprisals can capture common agrammatic deficits at both word and sentence level. Overall, LLMs have potential to advance automatic and precise aphasia prediction. A natural language processing pipeline can be greatly benefitted from integrating LLMs, enabling us to refine models of existing language disorders, such as aphasia.

  • Research Article
  • 10.71097/ijsat.v16.i2.4275
Locally Deployed NLP System for Secure Document Summarization and Context-Aware Question Answering Using LLMs and Vector Embeddings.
  • May 1, 2025
  • International Journal on Science and Technology
  • Rama Krishna Pradhana - + 4 more

This paper presents a locally deployed Natural Language Processing (NLP) system designed to perform secure document summarization and context-aware question answering. The proposed framework addresses critical concerns related to data privacy by eliminating dependence on external APIs or cloud-based services. By leveraging advanced Large Language Models (LLMs), vector databases, and modern NLP tools such as LangChain, Hugging Face, Qdrant, and Streamlit, the system enables users to interact intelligently with unstructured documents in a fully offline environment. The architecture facilitates efficient document parsing, semantic embedding, and retrieval-augmented generation (RAG), empowering users to generate concise summaries and retrieve precise information through natural language queries. This solution is particularly well-suited for privacy-sensitive domains such as healthcare, law, finance, and research, where local data control and confidentiality are essential. The framework not only enhances document comprehension but also promotes efficient information retrieval and workflow automation

  • Research Article
  • Cite Count Icon 2
  • 10.3390/bioengineering12111219
The Development and Evaluation of a Retrieval-Augmented Generation Large Language Model Virtual Assistant for Postoperative Instructions
  • Nov 7, 2025
  • Bioengineering
  • Syed Ali Haider + 9 more

Background: During postoperative recovery, patients and their caregivers often lack crucial information, leading to numerous repetitive inquiries that burden healthcare providers. Traditional discharge materials, including paper handouts and patient portals, are often static, overwhelming, or underutilized, leading to patient overwhelm and contributing to unnecessary ER visits and overall healthcare overutilization. Conversational chatbots offer a solution, but Natural Language Processing (NLP) systems are often inflexible and limited in understanding, while powerful Large Language Models (LLMs) are prone to generating “hallucinations”. Objective: To combine the deterministic framework of traditional NLP with the probabilistic capabilities of LLMs, we developed the AI Virtual Assistant (AIVA) Platform. This system utilizes a retrieval-augmented generation (RAG) architecture, integrating Gemini 2.0 Flash with a medically verified knowledge base via Google Vertex AI, to safely deliver dynamic, patient-facing postoperative guidance grounded in validated clinical content. Methods: The AIVA Platform was evaluated through 750 simulated patient interactions derived from 250 unique postoperative queries across 20 high-frequency recovery domains. Three blinded physician reviewers assessed formal system performance, evaluating classification metrics (accuracy, precision, recall, F1-score), relevance (SSI Index), completeness, and consistency (5-point Likert scale). Safety guardrails were tested with 120 out-of-scope queries and 30 emergency escalation scenarios. Additionally, groundedness, fluency, and readability were assessed using automated LLM metrics. Results: The system achieved 98.4% classification accuracy (precision 1.0, recall 0.98, F1-score 0.9899). Physician reviews showed high completeness (4.83/5), consistency (4.49/5), and relevance (SSI Index 2.68/3). Safety guardrails successfully identified 100% of out-of-scope and escalation scenarios. Groundedness evaluations demonstrated strong context precision (0.951), recall (0.910), and faithfulness (0.956), with 95.6% verification agreement. While fluency and semantic alignment were high (BERTScore F1 0.9013, ROUGE-1 0.8377), readability was 11th-grade level (Flesch–Kincaid 46.34). Conclusion: The simulated testing demonstrated strong technical accuracy, safety, and clinical relevance in simulated postoperative care. Its architecture effectively balances flexibility and safety, addressing key limitations of standalone NLP and LLMs. While readability remains a challenge, these findings establish a solid foundation, demonstrating readiness for clinical trials and real-world testing within surgical care pathways.

  • Research Article
  • Cite Count Icon 22
  • 10.1016/j.cmpb.2024.108326
Information extraction from medical case reports using OpenAI InstructGPT
  • Jul 18, 2024
  • Computer Methods and Programs in Biomedicine
  • Veronica Sciannameo + 8 more

Background and objectiveResearchers commonly use automated solutions such as Natural Language Processing (NLP) systems to extract clinical information from large volumes of unstructured data. However, clinical text's poor semantic structure and domain-specific vocabulary can make it challenging to develop a one-size-fits-all solution. Large Language Models (LLMs), such as OpenAI's Generative Pre-Trained Transformer 3 (GPT-3), offer a promising solution for capturing and standardizing unstructured clinical information. This study evaluated the performance of InstructGPT, a family of models derived from LLM GPT-3, to extract relevant patient information from medical case reports and discussed the advantages and disadvantages of LLMs versus dedicated NLP methods. MethodsIn this paper, 208 articles related to case reports of foreign body injuries in children were identified by searching PubMed, Scopus, and Web of Science. A reviewer manually extracted information on sex, age, the object that caused the injury, and the injured body part for each patient to build a gold standard to compare the performance of InstructGPT. ResultsInstructGPT achieved high accuracy in classifying the sex, age, object and body part involved in the injury, with 94%, 82%, 94% and 89%, respectively. When excluding articles for which InstructGPT could not retrieve any information, the accuracy for determining the child's sex and age improved to 97%, and the accuracy for identifying the injured body part improved to 93%. InstructGPT was also able to extract information from non-English language articles. ConclusionsThe study highlights that LLMs have the potential to eliminate the necessity for task-specific training (zero-shot extraction), allowing the retrieval of clinical information from unstructured natural language text, particularly from published scientific literature like case reports, by directly utilizing the PDF file of the article without any pre-processing and without requiring any technical expertise in NLP or Machine Learning. The diverse nature of the corpus, which includes articles written in languages other than English, some of which contain a wide range of clinical details while others lack information, adds to the strength of the study.

  • Conference Article
  • Cite Count Icon 18
  • 10.3115/100964.100984
Plans for a task-oriented evaluation of natural language understanding systems
  • Jan 1, 1989
  • Beth M Sundheim

A plan is presented for evaluating natural language processing (NLP) systems that have focused on the issues of text understanding as exemplified in short texts from military messages. The plan includes definition of bodies of text to use as development and test data, namely the narrative lines from one type of naval message, and definition of a simulated database update task that requires NLP systems to fill a template with information found in the texts. Documentation related to the naval messages and examples of filled templates have been prepared to assist NLP system developers. It is anticipated that developers of a number of different NLP systems will participate in the evaluation and will meet afterwards to present and interpret the results and to critique the test design.

  • Conference Article
  • Cite Count Icon 7
  • 10.1109/tai.1993.633965
CARAMEL: A step towards reflection in natural language understanding systems
  • Nov 8, 1993
  • G Sabah + 1 more

The authors wish to extend to natural language understanding (NLU) systems a paradigm now seen as essential for AI: the use of meta-knowledge and the capability a system may have to observe its own functioning. First, they recall why multi-expert systems seem the best architecture for dealing efficiently with most NL constraints. Then they give general information about reflective reasoning models and propose some extensions to currently admitted ideas in the domain of distributed AI. Lastly, an illustration of these ideas with CARAMEL (in French: Comprehension Automatique de Recits, Apprentissage et Modelisation des Echanges Langagiers), the system developed in the group and able to perform various tasks using NLU.

  • Conference Article
  • Cite Count Icon 2
  • 10.1145/3589335.3641300
Large Language Models for Graph Learning
  • May 13, 2024
  • Yujuan Ding + 3 more

Graphs are widely applied to encode entities with various relations in web applications such as social media and recommender systems. Meanwhile, graph learning-based technologies, such as graph neural networks, are demanding to support the analysis, understanding, and usage of the data in graph structures. Recently, the boom of language foundation models, especially Large Language Models (LLMs), has advanced several main research areas in artificial intelligence, such as natural language processing, graph mining, and recommender systems. The synergy between LLMs and graph learning holds great potential to prompt the research in both areas. For example, LLMs can facilitate existing graph learning models by providing high-quality textual features for entities and edges, or enhancing the graph data with encoded knowledge and information. It may also innovate with novel problem formulations on graph-related tasks. Due to the research significance as well as the potential, the convergent area of LLMs and graph learning has attracted considerable research attention. Therefore, we propose to hold the workshop Large Language Models for Graph Learning at WWW'24, in order to provide a venue to gather researchers in academia and practitioners in the industry to present the recent progress on relevant topics and exchange their critical insights.

  • Research Article
  • 10.1093/jamiaopen/ooaf182
Lightweight open-source large language models versus cTAKES for information extraction from discharge summaries: tobacco smoking status test case
  • Jan 13, 2026
  • JAMIA Open
  • David M Dávila-García + 2 more

ObjectivesTo compare lightweight open-source large language models (LLMs) with cTAKES, a state-of-the-art natural language processing (NLP) system, in an information extraction task from hospitalization discharge summaries.Materials and MethodsTwo readers annotated 250 randomly sampled adult discharge summaries (BJC HealthCare, 2018-2023) for tobacco smoking status as “Smoker,” “Never smoker,” “Unknown.” Six LLMs (Llama-3 [1B-70B], gpt-oss-20B, MedGemma-27B) and cTAKES extracted smoking status from summaries. Performance was benchmarked against consensus annotations using weighted F1-score, macro F1-score, and per-class F1-scores and a noninferiority test.ResultsInter-reader agreement was excellent (κ = 0.91). LLM size (2.3-47.3 GB) and inference time (2.5-14.5 s/note) varied. gpt-oss-20B achieved non-inferior performance vs cTAKES (F1 = 0.99 vs 0.97; P < .021).DiscussionThe high accuracy and efficiency of gpt-oss-20B support its potential as a practical, open-source alternative to traditional NLP for clinical information extraction.ConclusionLightweight LLMs can be applied for use across diverse clinical information extraction tasks without the need for task-specific fine-tuning.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 30
  • 10.1007/s10462-025-11147-4
Adversarial machine learning: a review of methods, tools, and critical industry sectors
  • May 3, 2025
  • Artificial Intelligence Review
  • Sotiris Pelekis + 7 more

The rapid advancement of Artificial Intelligence (AI), particularly Machine Learning (ML) and Deep Learning (DL), has produced high-performance models widely used in various applications, ranging from image recognition and chatbots to autonomous driving and smart grid systems. However, security threats arise from the vulnerabilities of ML models to adversarial attacks and data poisoning, posing risks such as system malfunctions and decision errors. Meanwhile, data privacy concerns arise, especially with personal data being used in model training, which can lead to data breaches. This paper surveys the Adversarial Machine Learning (AML) landscape in modern AI systems, while focusing on the dual aspects of robustness and privacy. Initially, we explore adversarial attacks and defenses using comprehensive taxonomies. Subsequently, we investigate robustness benchmarks alongside open-source AML technologies and software tools that ML system stakeholders can use to develop robust AI systems. Lastly, we delve into the landscape of AML in four industry fields –automotive, digital healthcare, electrical power and energy systems (EPES), and Large Language Model (LLM)-based Natural Language Processing (NLP) systems– analyzing attacks, defenses, and evaluation concepts, thereby offering a holistic view of the modern AI-reliant industry and promoting enhanced ML robustness and privacy preservation in the future.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant