Articles published on Information extraction
Authors
Select Authors
Journals
Select Journals
Duration
Select Duration
24796 Search results
Sort by Recency
- New
- Research Article
- 10.1109/tvcg.2026.3694842
- Jul 1, 2026
- IEEE transactions on visualization and computer graphics
- Min Wu + 6 more
Surface normal estimation is a fundamental task in point cloud processing and plays a crucial role in downstream applications. Existing methods typically extract features from local neighborhoods or patches, followed by surface fitting or direct regression to predict normals. However, the scale ambiguity in determining the optimal neighborhood hinders effective extraction of geometric information, making normal estimation for unstructured point clouds with significant density variations particularly challenging. To address this challenge, we propose APN-Net, an adaptive perception network for point cloud normal estimation. Specifically, we design the Graphical Information Self-perception (GIS) module, which provides an implicit manner for region partitioning and expands the receptive field, enabling automatic extraction of both local geometric details and global structural information, while alleviating the scale ambiguity in determining the optimal neighborhood. Moreover, to capture complex geometric details, we introduce the Adaptive Graph Convolution (AGC) module, which employs adaptive kernels to model relationships among points across different semantic regions, thereby enabling richer feature representation. Extensive experiments on both synthetic and real-world scanned datasets demonstrate that APN-Net achieves superior performance in unoriented normal estimation, particularly for point clouds with significant density variations.
- New
- Research Article
- 10.1016/j.clsr.2026.106313
- Jul 1, 2026
- Computer Law & Security Review
- Huai-Hsuan Huang + 3 more
Information extraction in legal texts: Investigating LLMs’ performance on traffic accident verdicts
- New
- Research Article
- 10.1088/1748-0221/21/07/p07002
- Jul 1, 2026
- Journal of Instrumentation
- Y An + 8 more
In this work, a 16-channel charge-to-digital converter (QDC) readout electronics system is developed for the time-of-flight (TOF) detectors of the High-energy FRagment Separator (HFRS) at the High-Intensity heavy-ion Accelerator Facility (HIAF) in China. The electronics is based on the time-over-threshold (TOT) technique and employs an FPGA-based fast linear discharge method for signal processing. A tapped delay-line time-to-digital converter implemented in the FPGA is used to measure both the leading edge timing and the signal duration above threshold, enabling simultaneous extraction of timing and energy information from particle events. The performance of the QDC readout electronics was evaluated using a pulse generator and a plastic scintillation detector hit by a laser source. Test results demonstrate that the system achieves an integral nonlinearity of 0.88% and an energy resolution better than 1% across an input charge dynamic range of 5–102 pC. Within the region of interest from 10 to 100 pC, a timing resolution better than RMS = 26.8 ps is obtained after time-walk correction. The system maintains sufficiently good timing and energy readout performance at 1 MHz. In joint tests with a plastic scintillation TOF detector, a timing resolution of 31.2 ps RMS after time-walk correction is achieved. To address the time-walk issue in QDC readout electronics applied to plastic scintillation detectors, a reference-free local correction method is proposed. The performance obtained using this method is comparable to that achieved with the reference-signal-based method, demonstrating the feasibility of this approach for future beam applications.
- New
- Research Article
- 10.1016/j.ultramic.2026.114377
- Jul 1, 2026
- Ultramicroscopy
- Umer Masood Chaudry + 6 more
Navigating the unavoidable: Enhancing the EBSD indexability through adaptation to three hurdles in femtosecond laser ablation.
- New
- Research Article
- 10.1016/j.newast.2026.102534
- Jul 1, 2026
- New Astronomy
- Tyler J Kapolka + 5 more
For the Circular Restricted 3-Body Problem (CR3BP), the topologies present within a Poincaré map enable the extraction of useful information regarding periodic, quasi-periodic, and chaotic trajectory behavior. Aside from the prominent topologies that follow distinct concentric patterns around fixed points, indicative of the periodic and quasi-periodic motion that is often the central focus of CR3BP research, there are also many “dusty” regions on the Poincaré map that appear random without an apparent structure and are indicative of chaotic motion. This paper, for the first time in literature, identifies dynamical structures associated with chaotic transport residing in the “dusty” region of a Poincaré map for the Earth-Moon system employing a novel methodology called conditional behavior mapping. Using a Jacobi constant of 3.175 to allow access to the Moon via L 1 but preventing system exit via L 2 , 8 distinct deterministic pathways–a series of dynamical conveyor belts –for chaotic transport to the Moon are mapped traversing these dynamical structures. In addition, this paper identifies sub-structures present within the dynamical structures that create a definable pattern distinguishing whether a chaotic transfer will result in lunar collision or a circumlunar trajectory, and shows the existence of at least one unstable translunar periodic orbit associated with these structures. This paper advances ongoing multi-body astrodynamics research by allowing for the mapping of chaotic transport to the Moon. • Dynamical structures associated with chaotic transport reside in the “dusty” region of a Poincaré map. • Sub-structures exist indicating if chaotic transfers result in lunar collision or circumlunar trajectories. • At least one unstable translunar periodic orbit can be mapped to the dynamical structures for chaotic transport.
- New
- Research Article
- 10.1016/j.jbi.2026.105036
- Jul 1, 2026
- Journal of biomedical informatics
- Ramtin Babaeipour + 2 more
AI-assisted protocol information extraction for improved accuracy and efficiency in clinical trial workflows.
- New
- Research Article
- 10.1016/j.eswa.2026.132010
- Jul 1, 2026
- Expert Systems with Applications
- Runwei Situ + 1 more
Self-improving multi-agent framework for zero-shot multimodal information extraction
- New
- Research Article
- 10.1016/j.patcog.2026.113121
- Jul 1, 2026
- Pattern Recognition
- Fukuan Wang + 4 more
Underwater image enhancement via degradation information extraction and guidance
- New
- Research Article
- 10.1245/s10434-026-19482-8
- Jul 1, 2026
- Annals of surgical oncology
- Anika Shah + 6 more
Nomograms predicting the likelihood of sentinel lymph node (SLN) metastasis in early-stage breast cancer can aid surgical decision-making but are underused due to the burden of variable review and input. This study evaluated whether OpenAI's large language models (LLMs) can extract information required for nomogram use and reproduce the estimated rate of SLN metastasis using the Memorial Sloan Kettering and MD Anderson nomograms. The study analyzed the de-identified radiology and pathology notes for 20 patients. Three prompts were tested: (1) o1 prompted to generate SLN metastasis estimates without nomograms, (2) o1 prompted to use the nomogram from online calculators with and without serial corrections, and (3) GPT-4o prompted to use chain-of-thought reasoning from nomogram variables and the corresponding point values from a pictorial nomogram. Artificial intelligence (AI) estimates of SLN metastasis rates were compared with a physician-expert's manual use of nomograms. OpenAI o1 captured all clinical variables in 65% of cases without serial correction (94-95% of individual variables) and 80% of cases with serial correction (96-97% of individual variables), with tumor size most frequently misidentified. Agreement between LLM-generated and physician-calculated risk prediction was low (exact matches in 0-10% of cases, near agreement in 20-45% of cases), indicating moderate rater reliability (intraclass correlation coefficient [ICC], 0.57-0.62). The optimized GPT-4o prompt demonstrated greater agreement (exact matches in 25% and near agreement in 90% of cases) and reliability (ICC, 0.95). Large language models can reliably extract nomogram inputs from clinical notes but require task-specific prompt engineering for accurate SLN metastasis risk estimation. Currently, automating nomogram-based risk estimators with AI may not justify the significant resources required for optimization.
- New
- Research Article
- 10.1016/j.csl.2026.101950
- Jul 1, 2026
- Computer Speech & Language
- Nikolaos Malamas + 2 more
Question Answering (QA) and Content Retrieval (CR) systems have experienced a boost in performance in recent years leveraging state-of-the-art Transformer models to process user expressions and retrieve and extract information requested. Despite the constant language understanding improvements, very little effort has been put into the design of such systems for personal desktop use, where data are kept locally and are not sent to cloud services and decisions and outputs are transparent and explainable to the user. To that end, we present QuAVA, a conversational desktop content retrieval assistant, designed on four pillars: privacy and security, explainability, low-resource requirements, and multi-source data fusion. QuAVA is a data and privacy-preserving assistant that enables users to access their private data such as files, emails, and message exchanges, conversationally and transparently. The proposed architecture automatically extracts and preprocesses content from various sources and organizes it in a 3-layered hierarchical structure, namely a topic, a subtopic, and a content layer by employing ML algorithms for clustering and labeling. This way, users can navigate and access information via a set of conversation rules embedded in the assistant. We conduct a qualitative comparison analysis of the QuAVA architecture with other well-established QA and CR architectures against the four pillars defined, as well as privacy tests, and conclude that QuAVA is the only -to our knowledge- virtual assistant that successfully satisfies them. • Design and development of a desktop Content Retrieval assistant. • Private and secure Content Retrieval system. • Personal desktop assistant for Content Retrieval for everyday use. • Hierarchical structure for explainable and transparent Virtual Assistants.
- New
- Research Article
- 10.1016/j.disamonth.2026.102113
- Jul 1, 2026
- Disease-a-month : DM
- Praveen Gunasekaran + 7 more
Approved weight loss drugs for obesity with a thorough emphasis on GLP-1 agonist medications: A systematic review.
- New
- Research Article
- 10.1007/s40820-026-02265-x
- Jun 30, 2026
- Nano-micro letters
- Huijae Park + 4 more
Soft electronics are an emerging class of mechanically compliant platforms that enable conformal, skin-interfaced sensing and actuation on curvilinear and dynamic surfaces. These systems combine deformation-tolerant electrical functionality with soft contact mechanics, but their in-use performance is strongly influenced by time-varying interfaces, motion-induced artifacts, and the system burden associated with dense multimodal integration. Advances in soft electronics are now converging with artificial intelligence, which supports reliable information extraction from high-dimensional signals and enables on-device inference that tolerates variability across users and day-to-day conditions. Here, progress in this convergence from materials to intelligent systems is summarized. Material and interface foundations are introduced first, focusing on deformation-tolerant conductors, low-impedance biointerfaces, and breathable substrate strategies that support extended wear. Manufacturing and integration approaches are then discussed, highlighting scalable fabrication, multilayer interconnects, and energy-autonomous wireless operation that enable higher channel counts and multifunctional architectures. Learning-based pipelines are subsequently reviewed with emphasis on artifact suppression, nonideality compensation, multimodal inference, and efficient edge deployment. Finally, emerging directions including neuromorphic computing and in-sensor computing are discussed, together with current challenges and future opportunities toward deployable intelligent soft systems that operate continuously and reliably in everyday settings.
- New
- Research Article
- 10.1038/s41467-026-74657-x
- Jun 30, 2026
- Nature communications
- Shijie Xu + 1 more
Metalloproteins are essential to many cellular processes. They use metal ions as cofactors to catalyze reactions, stabilize protein structures, and mediate electron transfer. Identifying their metal-binding sites remains difficult because of the complexity of protein environments and the promiscuous binding of metal ions, and existing computational methods are limited by accuracy and data scarcity. Here we introduce PRIME, a hybrid deep learning framework that combines evolutionary and structural signals to predict metal-binding sites accurately and efficiently. PRIME employs protein language models and pre-trained structure models to extract information from protein sequences and structures, together with a probe generation algorithm that bridges sequence- and structure-based predictions by scanning candidate sites. PRIME outperforms existing methods across diverse metal ions, from abundant zinc and calcium to challenging potassium and sodium. Ablation analysis shows that pretrained structure models improve accuracy. Case studies on AlphaFold2 models further demonstrate PRIME's potential for high-throughput metalloproteomics.
- New
- Research Article
- 10.1080/02687038.2026.2691169
- Jun 29, 2026
- Aphasiology
- Jessica De Leon + 6 more
ABSTRACT Background The relative scarcity of studies of bilingual speakers with primary progressive aphasia (PPA) leads to knowledge gaps regarding their symptoms and patterns of language impairment. It is unknown if PPA symptoms in bilingual speakers differ from those of monolingual speakers. In addition, it remains unclear if one of a bilingual speaker’s languages is more vulnerable to the onset of symptoms and/or is better preserved. Furthermore, bilingualism factors (e.g. order/age of language acquisition (L1/L2) and language dominance) may explain patterns of impairment both within and across the variants of PPA. Aims To characterize the bilingualism factors, speech and language symptoms, and patterns of language impairment in a retrospective cohort of bilingual speakers with PPA. Methods In a large cohort (N = 69) of bilingual speakers, we performed a chart review to extract information regarding self/caregiver-reported PPA symptoms, first language impacted by PPA at symptom onset, less preserved language at time of evaluation, and bilingual history factors (age/order of acquisition (L1/L2) and language dominance). We explored how emergence and presentation of symptoms associated with bilingualism factors. Results The presenting symptoms were mostly consistent with current PPA diagnostic criteria, although symptoms unique to bilingual speakers were reported. The language reported to first be impacted by PPA, as well as the less preserved language at the time of evaluation, was generally reported to be the less dominant language, regardless of PPA variant. These measures did not associate as closely with age/order of acquisition (L1 versus L2). Conclusion Bilingual speakers with PPA may report additional symptoms that reflect their ability to speak more than one language. The less dominant language was most susceptible to the initial impact of PPA and also tended to the less preserved language at the time of evaluation. Future studies of bilingual speakers with neurodegenerative diseases should systematically consider bilingualism factors, which are likely to contribute to clinical presentation and disease progression.
- New
- Research Article
- 10.1038/s41598-026-59180-9
- Jun 23, 2026
- Scientific reports
- Zhikun Zhao + 5 more
Between the occurrence of a disaster and the initiation of formal investigation and assessment, the overall evolution of the disaster often remain unclear, which hinders the timely progress of investigative work. To address this challenge, this study proposes a framework for constructing disaster event chains from limited post-disaster information. The framework first defines event patterns based on disaster system theory, then employs the unified information extraction model to automatically extract event chain nodes and their relationships from news texts. DBSCAN clustering is further applied to achieve semantic normalization and entity standardization, enabling the reconstruction of the disaster event chain. Finally, a fuzzy comprehensive evaluation is conducted to assess the effectiveness of the proposed framework. The results demonstrate that the framework can provide valuable initial analytical insights and auxiliary support for disaster investigation efforts.
- New
- Research Article
- 10.1093/jamia/ocag101
- Jun 22, 2026
- Journal of the American Medical Informatics Association : JAMIA
- Han Yang + 11 more
We aimed to develop a data model and a natural language processing (NLP) pipeline for representing physical activity (PA) in Electronic Health Records (EHRs), and to evaluate transformer- and Large Language Model (LLM)-based classifiers for sentence-level PA attribute classification. We analyzed PA documentation across three patient cohorts (cancer, COVID, and Alzheimer's disease) using structured and unstructured EHR data. A conceptual schema was developed to represent PA and its linguistic attributes. Five BERT models and three modern LLMs (Llama3-8B, MedAlpaca-13B, and PMC-Llama-13B) were evaluated for classifying PA attributes (binary status, negation, exclusion, and an eleven-class Category) on pre-extracted PA-related sentences. Clinical notes were a richer source of PA information than structured ICD or SDoH data. On binary tasks, the best BERT model reached F1 0.619 (Exclusion); with Supervised Fine-Tuning (SFT), Llama3-8B reached F1 0.689 (Exclusion). On the 11-class Category task, performance was modest (best macro-F1 0.262, ROC-AUC 0.803, by Llama3-8B). In-Context Learning (ICL) was highly variable: while Llama3-8B-ICL achieved the best ROC-AUC on Category, the domain-specific MedAlpaca-13B and PMC-LLaMA-13B essentially failed. These results, together with sparsely represented PA elements (Amount, Frequency, Assessment) and the absence of a downstream evaluation, position this work as an initial proof of feasibility, with supervised domain adaptation still required for reliable clinical PA extraction. We contribute a PA data model, annotation schema, and a working NLP pipeline with a BERT/LLM benchmark for sentence-level PA attribute classification. The pipeline supports future end-to-end PA extraction and downstream applications such as phenotyping, risk prediction, and cohort identification.
- New
- Research Article
- 10.1016/j.compbiomed.2026.111806
- Jun 20, 2026
- Computers in biology and medicine
- Guillermo Villanueva Benito + 5 more
CineScribe: LLM-based detection and mitigation of ambiguity in cine cardiac magnetic resonance reports.
- New
- Research Article
- 10.1093/jamia/ocag084
- Jun 19, 2026
- Journal of the American Medical Informatics Association : JAMIA
- Enshuo Hsu + 6 more
Generative information extraction using large language models (LLMs), particularly through prompting combined with few-shot learning, has become a popular method. In many ways such prompts with examples resemble the annotation guidelines long used for manual labeling of data for information extraction, and indeed studies have demonstrated the direct use of these guidelines as effective prompts. However, constructing annotation guidelines is both labor- and knowledge-intensive. Instead, this paper proposes to leverage LLMs' impressive ability to automatically create such annotation guidelines. Specifically, we propose a zero-shot hierarchical prompt engineering method that harvests the knowledge summarization and text generation capacity of LLMs to synthesize annotation guidelines to improve downstream LLMs while requiring minimal human input. Zero-shot clinical named entity recognition benchmarks, 2012 i2b2 EVENT, 2012 i2b2 TIMEX, 2014 i2b2, and 2018 n2c2 showed improvements of 0.2% to 25.86% for Llama 3.1 and 5.82% to 16.13% for GPT-OSS in strict F1 scores from the no-guideline baseline. The LLM-synthesized guidelines showed equivalent or better performance compared to human-written guidelines by 0.23% to 10.00% in most tasks. LLMs generate high-quality annotation guidelines following a consistent pattern (eg, title, entity types, examples) without human guidance, indicating that a representation of such a concept has been encoded during the pre-training. Nuances in definitions, however, still require adjustment by researchers to align with the project. This study proposes a novel hierarchical prompt engineering method that requires minimal knowledge transfer from a human expert and is applicable to multiple biomedical domains.
- Research Article
- 10.1007/s00198-026-08100-8
- Jun 18, 2026
- Osteoporosis international : a journal established as result of cooperation between the European Foundation for Osteoporosis and the National Osteoporosis Foundation of the USA
- Tzu-Hao Tseng + 13 more
Osteoporosis is a major cause of fragility fractures, yet limited access to DXA leads to underdiagnosis and delayed treatment. Recent advances in artificial intelligence enable extraction of bone structural information from routine radiographs, providing a potential tool for opportunistic osteoporosis screening. Whether AI-derived BMD can approximate DXA and predict real-world fracture risk remains unclear. Adults aged ≥ 20years who underwent both lumbar DXA and radiographic examinations (lumbosacral or kidney-ureter-bladder) within six months between January 2014 and December 2024 were retrospectively analyzed. Lumbar BMD was estimated using DeepXray Spina and compared with DXA using Pearson correlation, intraclass correlation coefficient (ICC), and Bland-Altman analysis. Diagnostic performance for osteoporosis (T-score ≤ - 2.5) and fracture prediction was evaluated using receiver operating characteristic (ROC) analysis, Cohen's κ, and logistic regression. Among 540 participants (73.9% female; mean age 57.0years; mean follow-up 6.4years), AI- and DXA-derived BMD showed strong agreement (r = 0.943; ICC = 0.934). For osteoporosis diagnosis, AI-derived T-scores achieved an AUC of 0.959, κ = 0.74, and 90% accuracy. AI- and DXA-derived BMD showed comparable performance for predicting vertebral (AUC 0.704 vs. 0.678) and hip fractures (0.716 vs. 0.678). For all-site fractures, AI-derived BMD showed a modestly higher AUC than DXA-derived BMD (AUC, 0.699 vs 0.677; P = 0.042). Lower AI-derived BMD was independently associated with higher fracture risk. AI-derived BMD from routine radiographs closely correlates with DXA and demonstrates comparable fracture prediction. This approach may support opportunistic osteoporosis screening without reliance on DXA.
- Research Article
- 10.1038/s41598-026-57300-z
- Jun 18, 2026
- Scientific reports
- Seonghu Jung + 1 more
Orbital angular momentum (OAM) of light has acquired the interest of scientists due to its novel properties. The high-dimensional Hilbert space and entanglement that OAM offers make it a valuable DOF for many quantum optical protocols. However, conventional measurement of OAM state, which exploits coincidence count-based quantum state tomography (QST), suffers from low-brightness and subsequent need for data accumulation times. Such problems are a serious obstacle, especially for high-dimensional OAM qudits. In this work, we suggest stimulated emission tomography (SET) as a solution for bright and efficient measurement of OAM entangled photons. We show that SET can successfully extract information of SPDC photons by measuring the spiral bandwidth and reconstructing the density matrix of SPDC photons. The fidelity and linear entropy of the reconstructed density matrix are [Formula: see text] and [Formula: see text], respectively, while the time required for each projective measurement is 1s. Our results show the potential of SET as an efficient alternative to conventional QST.