Articles published on Computer Vision
Authors
Select Authors
Journals
Select Journals
Duration
Select Duration
69595 Search results
Sort by Recency
- New
- Research Article
- 10.1016/j.jconrel.2026.114994
- Jul 10, 2026
- Journal of controlled release : official journal of the Controlled Release Society
- Bochang Feng + 11 more
PSMA-directed nanoprobe enabling precision imaging and MR-guided chemoradiotherapy in prostate cancer.
- New
- Research Article
- 10.58257/ijprems51116
- Jul 8, 2026
- International Journal of Progressive Research in Engineering Management and Science
Automated microplastic quantification is currently compromised by morphological mimicry: air bubbles, organic biofilms, and sediment create high false-positive rates in standard Convolutional Neural Networks (CNNs).This study introduces AquaEye, a containerized computer vision framework that mitigates artifact misclassification by fusing Bayesian Deep Learning with ISO-standard morphometrics.Unlike deterministic U-Net implementations, we deploy a Monte Carlo Dropout inference pipeline to estimate epistemic uncertainty, enabling the suppression of predictions where model variance exceeds a safety threshold.To enforce physical validity, a post-processing geometric gate rejects candidates based on Circularity (4A/P 2 ) and Solidity, filtering non-polymer structures that bypass the neural filter.The system, deployed via a Dockerized microservices architecture, ensures reproducibil-ity often absent in -lab-bench scripts.Experimental validation confirms that AquaEye statistically decouples true microplastic instances from background noise, offering a robust alternative to manual microscopy for highthroughput environmental monitoring.
- New
- Research Article
1
- 10.1109/tpami.2026.3672629
- Jul 1, 2026
- IEEE transactions on pattern analysis and machine intelligence
- Qiyang Wan + 3 more
Visual recognition models have achieved unprecedented success in various tasks. While researchers aim to understand the underlying mechanisms of these models, the growing demand for deployment in safety-critical areas like autonomous driving and medical diagnostics has accelerated the development of eXplainable AI (XAI). Distinct from generic XAI, visual recognition XAI is positioned at the intersection of vision and language, which represent the two most fundamental human modalities and form the cornerstones of multimodal intelligence. This paper provides a systematic survey of XAI in visual recognition by establishing a multi-dimensional taxonomy from a human-centered perspective based on intent, object, presentation, and methodology. Beyond categorization, we summarize critical evaluation desiderata and metrics, conducting an extensive qualitative assessment across different categories and demonstrating quantitative benchmarks within specific dimensions. Furthermore, we explore the interpretability of Multimodal Large Language Models and practical applications, identifying emerging trends and opportunities. By synthesizing these diverse perspectives, this survey provides an insightful roadmap to inspire future research on the interpretability of visual recognition models.
- New
- Research Article
- 10.1016/j.applanim.2026.106988
- Jul 1, 2026
- Applied Animal Behaviour Science
- Dana L.M Campbell + 11 more
AI-driven analysis of cattle behaviour in lairage: Impact of flooring substrate and implications for welfare and management
- New
- Research Article
- 10.1016/j.patcog.2026.113091
- Jul 1, 2026
- Pattern Recognition
- Abu Sufian + 5 more
• DemoFace provides a demographically balanced pixelated real face image dataset to mitigate biases in face biometric systems. • The dataset’s structured image–text embedding multimodality supports downstream tasks and facilitates the analysis of model biases in CVFMs. • Through our novel Responsible AI methodology, we developed the dataset and established a set of baselines using SOTA CVFMs to assess the performance and fairness of CVFMs in face biometric tasks. • Through our evaluation using both new and adapted metrics, we revealed inherent bias patterns in several SOTA CVFMs, providing valuable insights to guide future research. Bias and fairness are critical challenges in data-driven computer vision (CV), where limited demographic diversity in training data worsens these challenges. Face biometric (face recognition) systems are core tasks of CV that are highly impacted by these challenges, as existing real-face datasets lack comprehensive demographic representation, whereas current synthetic datasets promote stereotypes. CV Foundation Models (CVFMs) are currently at the forefront of CV applications, including face biometrics, which use global features in multimodal data. However, the scarcity of large-scale, demographic multimodal datasets, such as image-text embeddings for model fine-tuning (or training), limits the fairness in state-of-the-art (SOTA) CVFMs for downstream face biometric tasks. To address these issues, we introduce DemoFace, a balanced demographic face dataset comprising 30,240 pixelated real face images of 672 representative individuals evenly distributed across 48 demographic groups, categorized by ethnicity/race, gender, and age. We gathered images using an API set up from multiple copyright-free public forums. The collected images were then manually filtered, anonymized, and annotated by two independent research groups, and then lightly pixelated for privacy preservation. DemoFace’s image-text embedding multimodality enables fine-tuning (or training) of CVFMs for fairness-focused face biometrics tasks and bias pattern evaluation. Through two empirical studies: face authentication as classification and textual description as token generation, we established baseline scores across ethnicity/race, gender, and age groups. Our baselines identified inherent bias patterns through both new and tailored metrics derived from existing ones, emphasizing the need for more equitable AI models. Here is the Repository: Link
- New
- Research Article
- 10.1037/xlm0001533
- Jul 1, 2026
- Journal of experimental psychology. Learning, memory, and cognition
- Roberto G De Almeida + 2 more
How is a word visually recognized? We investigated word recognition by employing a novel psychophysical task in which word segments (true morphemes or nonmorphemic orthographic sequences) were colored blue/red, red/blue, or black, while participants wore anaglyph glasses. The color combinations split the words either at the morpheme boundary (legal split) or at one character to the left or to the right of the split (illegal split). This manipulation divided the retinal information along the fovea, allowing us to project the different word segments to both left and right hemispheres via contralateral and ipsilateral visual pathways. Our goal was to investigate the roles of morphology and semantics in the early moments of visual word recognition, potentially tapping functional properties of the visual word form area. Participants (N = 72) performed a masked lexical decision task (60 ms exposure) to semantically transparent compounds (e.g., football), monomorphemic pseudocompounds (e.g., shamrock), and unsegmentable monomorphemic words (e.g., jeopardy). We found an effect of legality for compounds and pseudocompounds, when presented via the stronger contralateral visual pathways, suggesting an early morpho-orthographic segmentation, but with true compounds yielding faster and more accurate responses than pseudocompounds across all projection combinations. Our results are compatible with a model of visual word recognition involving an initial stage of morpho-orthographic analysis that is insensitive to semantics, followed by later stages of morphological computations and semantic access. Although our evidence relies on a psychophysical task, we suggest that these processing stages are in line with a functional division between visual word form area-1 (orthographic) and visual word form area-2 (lexical) in the early visual word recognition system. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
- New
- Research Article
- 10.1016/j.marenvres.2026.108108
- Jul 1, 2026
- Marine environmental research
- Ziyu Wang + 10 more
Artificial intelligence for marine oil spill management: Recent advances and future directions.
- New
- Research Article
- 10.1107/s2052252526003568
- Jul 1, 2026
- IUCrJ
- Ben A Coulson + 5 more
The `photometric selection' approach as a high throughput, sample tolerant, low cost and highly automatable method of carrying out serial crystallography is presented. Crystalline samples are loaded and distributed onto a simple transparent substrate and an in-line camera identifies crystals using image recognition algorithms from the computer vision project OpenCV. In contrast to established serial techniques, which generally require that crystal samples be refined with narrow size distributions and defined habits, the sample requirements when using photometric selection are shown to be minimal. We demonstrate how broadly effective photometric selection can be by collecting high-quality datasets from three exemplar systems: a small-molecule organometallic, a small-molecule organic and a metal-organic framework system. In contrast to previously established grid-scanning techniques, data collection using photometric selection can be up to six times faster.
- New
- Research Article
- 10.1016/j.gaitpost.2026.110213
- Jul 1, 2026
- Gait & posture
- Raquel Costa + 2 more
Comparison of marker-based and markerless motion capture systems to assess gait kinematics and kinetics in children with cerebral palsy.
- New
- Research Article
- 10.1177/00220345261434568
- Jul 1, 2026
- Journal of dental research
- Z Chen + 7 more
Artificial intelligence (AI) holds transformative potential for advancing oral health surveillance by streamlining data collection, integration, and dissemination. This review critically synthesizes AI applications in oral health surveillance, highlighting its roles in 1) mapping population-level trends and oral health inequities using machine learning on epidemiological data; 2) enabling remote screening of oral diseases/conditions, including caries, oral hygiene, gingivitis, oral cancer, and malocclusion from intraoral images via computer vision models; and 3) integrating multimodal data through emerging large language models (LLMs) to enhance precision public health. We clarify the comparative strengths of distinct AI modeling for processing the primary data types in surveillance: structured clinical records, unstructured images, and integrated multimodal data. Traditional machine learning methods have been effectively applied to map population-level oral health disparities and identify risk factors but are constrained to structured data. Computer vision methods excel in individual-level diagnostics using intraoral photographs. To translate such capability into scalable surveillance, it is recommended to establish standardized imaging protocols for nonclinical settings, develop scalable models for fine-grained feature extraction, and implement reliable evaluation. These steps are essential to address pervasive challenges, including inconsistent image quality, domain shift, prevalence imbalance, and cost-effectiveness constraints. The future of AI-driven oral health surveillance lies in developing dental-adapted multimodal LLMs (MLLMs). Such MLLMs are uniquely capable of synthesizing disparate data streams, from structured clinical data to heterogeneous imaging modalities (e.g., intraoral photographs and radiographs), or even biomolecular data. This integration capacity facilitates a paradigm shift, moving current applications in dental consultation and clinical decision support toward a novel, tiered system for population-level monitoring. Such a system would provide actionable insights for public health policymaking via spatiotemporal analysis and causal inference. Next-generation AI-driven oral health surveillance systems can only succeed when built on a strong foundation of rigorous ethical principles and safeguards.
- New
- Research Article
1
- 10.1016/j.psj.2026.106887
- Jul 1, 2026
- Poultry science
- Bidur Paneru + 6 more
Computer vision models for precision poultry farming: A narrative review of behavioral and welfare monitoring studies.
- New
- Research Article
- 10.1109/tpami.2026.3673525
- Jul 1, 2026
- IEEE transactions on pattern analysis and machine intelligence
- Banglei Guan + 2 more
In recent years, affine correspondences (ACs) have emerged as widely adopted alternative to point correspondences (PCs) in geometric problems in computer vision. An AC is composed of a PC across two different views plus an affine transformation between the small patches around this PC. Prior studies have shown that a single affine correspondence (AC) generally yields three independent constraints for estimating relative pose. This work addresses relative pose estimation in multi-perspective camera systems, a relevant problem given their prevalence in modern technologies such as autonomous vehicles and augmented reality. More specifically, we introduce the first comprehensive suite of minimal solvers for 6DoF relative pose estimation across multiple cameras using only two ACs, which is notably valuable for robust model fitting scenarios. We analyze all possible configurations of two ACs in two views, and present minimal solvers covering all identified minimal cases. We make use of the hidden variable technique to eliminate the translation parameters, and represent rotation using either Cayley parameters or quaternions. We furthermore introduce novel constraints on the generalized relative pose problem that are beneficial in deriving more compact solvers with fewer solutions. Comprehensive experiments on synthetic and real-world data show that the proposed affine correspondence-based solvers are highly effective and computationally efficient.
- New
- Research Article
- 10.1109/tvcg.2026.3662816
- Jul 1, 2026
- IEEE transactions on visualization and computer graphics
- Songle Chen + 4 more
Estimating the 6-DoF posture of parts in assembly-based modeling is a critical task in the fields of computer graphics, computer vision and robotics. A typical scenario involves enabling a machine agent to automatically assemble IKEA furniture using the provided parts. This paper presents HiFormer, a novel Hierarchical Transformer with Box-packed Positional Encoding, designed for highly automatic 3D part assembly. Our method addresses three important issues commonly encountered in 3D part assembly: 1) How to mitigate the overfitting problem associated with Transformer-based feature learning for 3D point clouds? 2) How to effectively model the relationships between the intragroup and intergroup parts? 3) How to compute positional encoding and integrate it into the Transformer for parts with diverse geometric forms in the coarse-to-fine assembly process? These challenges are tackled through three key contributions: 1) a multi-task 3D Swin Transformer with a two-stage training strategy for feature extraction, 2) a novel hierarchical Transformer for capturing part relationships at flattening, intragroup, and intergroup levels, and 3) an innovative box-packed positional encoding that enhances the Transformer by incorporating query, key, and value information derived from relative box positions. On the PartNet benchmark, our method outperforms the state-of-the-art PWH-MP model on three representative categories-Chair, Table, and Lamp-, achieving average improvements of 2.84% in Part Accuracy (PA) and 3.72% in Connection Accuracy (CA) for diversity modeling (with noise), and 3.55% in PA and 3.21% in CA for deterministic modeling (without noise).
- New
- Research Article
- 10.1016/j.robot.2026.105442
- Jul 1, 2026
- Robotics and Autonomous Systems
- Hao Chen + 3 more
Generalizable task-oriented object grasping through LLM-guided ontology and similarity-based planning
- New
- Research Article
- 10.1016/j.compag.2026.111787
- Jul 1, 2026
- Computers and Electronics in Agriculture
- Lepeng Song + 8 more
Design and implementation of an orchard variable spray system based on visual recognition and fuzzy adaptive decoupling
- New
- Research Article
- 10.1016/j.envres.2026.124520
- Jul 1, 2026
- Environmental research
- Sina Borzooei + 6 more
Morphology-informed deep learning for risk assessment of filamentous bulking in a full-scale industrial wastewater treatment plant.
- New
- Research Article
- 10.1037/xge0001931
- Jul 1, 2026
- Journal of experimental psychology. General
- Conor J R Smithson + 3 more
The domain-general ability to recognize objects at the subordinate level (o) is most often tested in the visual modality. However, previous work has evidenced a strong relationship between visual and auditory o abilities, suggesting that o may be multimodal. We measured this relationship using a structural equation modeling approach. We found a strong relationship (r = .79) between visual and auditory o. There were similarly strong relationships between spatial ability and both visual (r = .73) and auditory (r = .63) o. When the shared influence of general ability was controlled for, the relationship between visual and auditory o remained substantial (r = .61), while relationships between spatial ability and visual (r = .31) and auditory o (r = .05) were nonsignificant. These results suggest that individual differences in object recognition are largely multimodal and that the relationship between o and spatial ability is not clearly independent of general ability. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
- New
- Research Article
- 10.1016/j.cortex.2026.04.008
- Jul 1, 2026
- Cortex; a journal devoted to the study of the nervous system and behavior
- Joonwoo Kim + 2 more
Rapid orthographic and delayed phonological processing: ERP and oscillatory evidence from masked priming in Korean.
- New
- Research Article
1
- 10.1016/j.ijmedinf.2026.106417
- Jul 1, 2026
- International journal of medical informatics
- Abdur Rasool + 3 more
Challenges in translating AI-driven ASD/ADHD diagnosis: A methodological systematic review.
- New
- Research Article
- 10.1111/1541-4337.70551
- Jul 1, 2026
- Comprehensive reviews in food science and food safety
- Yao Zheng + 4 more
Seafood provides high-quality protein and essential nutrients but is highly susceptible to rapid postharvest deterioration. Conventional quality evaluation methods are often destructive, labor-intensive, and difficult to implement in real-time industrial and consumer settings. In recent years, deep learning-assisted computer vision (DL-CV) has emerged as a promising technical route for nondestructive and rapid seafood quality assessment in industrial processing lines and retail inspection systems. This review synthesizes recent advances in DL-CV for seafood quality evaluation from both technical and application-oriented perspectives. The technical framework is first clarified by comparing visible-light imaging with other imaging modalities and by discussing the transition from traditional machine learning to end-to-end deep learning-based feature extraction. Recent developments in representative deep learning architectures, such as convolutional neural networks and vision transformers, are then summarized alongside emerging trends toward lightweight design, architectural enhancement, and interpretability. Current application scenarios are further reviewed systematically, with particular emphasis on freshness evaluation, including key image regions and two dominant labeling strategies: storage time-based labeling (STBL) and traditional indicator-based labeling (TIBL). Additional applications such as species identification, weight estimation, defect detection, and quantitative determination of compositional attributes are also discussed. Overall, DL-CV demonstrates strong potential for accurate, nondestructive seafood quality prediction, especially when leveraging visible-light imaging systems that are easily deployable in practical environments. Future research should focus on finer quality differentiation, multidimensional and integrated quality evaluation, and scalable deployment in both industrial processing and consumer-oriented applications.