Articles published on Matrix multiplication
Authors
Select Authors
Journals
Select Journals
Duration
Select Duration
7582 Search results
Sort by Recency
- New
- Research Article
- 10.1016/j.engappai.2026.114709
- Jul 1, 2026
- Engineering Applications of Artificial Intelligence
- Haiyang Pan + 4 more
Multiple restriction cross-border matrix machine for multiple objective fault diagnosis
- New
- Research Article
- 10.1016/j.mejo.2026.107181
- Jul 1, 2026
- Microelectronics Journal
- Jianzi Jin + 10 more
A full-stack, energy-efficient 8-Mb NOR-flash compute-in-memory SoC with on-chip acceleration and end-to-end toolchain for edge AI
- New
- Research Article
- 10.1109/tvcg.2026.3683714
- Jul 1, 2026
- IEEE transactions on visualization and computer graphics
- Fang-Chi Chang + 1 more
Rendering large-scale, unbounded scenes on AR/VR-class devices is constrained by the computation, bandwidth, and storage cost of 3D Gaussian Splatting (3DGS). We propose a low-power, low-cost 3DGS hardware accelerator that renders full-HD images in real time, together with a hardware-friendly compression pipeline that combines iterative Gaussian pruning and fine-tuning, progressive spherical harmonics (SH) degree reduction, and vector quantization of all SH coefficients and colors. The scheme achieves a $51.6\times$51.6× model-size reduction with a 0.743 dB PSNR loss. The accelerator uses a frame-level pipeline that integrates point-based culling and projection with tile-based sorting and rasterization, skips zero-Jacobian matrix multiplications (reducing processing elements by 63% and computation by 53%), and adopts comparison-free tile-based sorting with deterministic latency. Implemented in a TSMC 28-nm process at 800MHz, the design occupies $\text{0.66}\;\text{mm}^{2}$0.66mm2 with 1.1438 M gates and 120 kB SRAM, consumes 0.219 W, and delivers 1219 Mpixels/J at 267.5 Mpixels/s, enabling 1080p at 129 FPS. Overall, it is $5.98\times$5.98× smaller in area, $5.94\times$5.94× higher throughput, and delivers $7.5\times$7.5× higher energy efficiency than prior 3DGS accelerators.
- New
- Research Article
- 10.1088/1361-6528/ae83c0
- Jun 29, 2026
- Nanotechnology
- Min-Woo Kwon + 2 more
The rapid growth of artificial intelligence computing has intensified the demand for energyefficient hardware accelerators capable of large-scale matrix-vector multiplication. Resistive random-access memory has attracted significant interest for such applications due to its analog weight storage capability and compatibility with crosspoint array architectures. However, the sneak-path current remains a critical challenge that limits the scalability and reliability of highdensity RRAM arrays. In this work, a palladium (Pd)-MoS 2 based self-rectifying RRAM device was fabricated and experimentally characterized, and a physics-informed compact model is developed to quantitatively evaluate its sneak-pass current suppression capability at the array level. The device exhibited asymmetric bipolar resistive switching originating from Schottkybarrier-controlled carrier injection at the metal/MoS 2 interface. To accurately capture this intrinsic rectifying behavior, a modulated thermionic-emission formulation based on the Richardson-Dushman equation was incorporated into the conventional Lehtonen-Laiho framework. This formulation preserved the essential Schottky barrier physics while ensuring numerical stability in circuit-level simulations. The proposed compact model reproduced the measured current-voltage characteristics, including a rectification ratio of approximately 60, a memory window on the order of 10 3 , and stable bipolar switching behavior. Furthermore, by systematically varying key physical parameters, such as metal work function, MoS 2 electron affinity, and Fermi-level pinning factor, the model enabled predictive estimation of Schottky barrier height and corresponding rectification characteristics for various metal/MoS 2 combinations.
- New
- Research Article
- 10.1021/acs.analchem.6c02208
- Jun 29, 2026
- Analytical chemistry
- Zhenshengnan Li + 7 more
Although significant progress has been achieved in diagnosing papillary thyroid carcinoma (PTC), it remains a challenge to diagnose thyroid nodules (TNs) with ambiguous ultrasound features. Matrix metalloproteinases (MMPs), key mediators of tumor invasion and metastasis, are highly promising biomarkers for PTC. In this manuscript, a machine learning (ML)-assisted peptide microarray sensing platform (MLPM) is developed for accurate diagnosis of PTC through the combination of the high-throughput ability of peptide microarray and the powerful predictive performance of the gradient boosting (GB) algorithm. The peptide microarray enables the simultaneous detection of seven MMP activities with low limits of detection (LODs) (down to the pg mL-1 level) and wide dynamic ranges (more than 3 orders of magnitude) across diverse sample types, including cells, tissues, and plasma. The GB model with MMP activity profiles exhibits exceptional diagnostic performance, achieving area under the receiver operating characteristic curve (AUC) values of 0.99 for discriminating benign from malignant nodules and 0.97 for predicting lymph node metastasis status in PTC. Crucially, in an independent blind validation, the model maintains high accuracy of 96.0% and 94.1% for these two critical clinical tasks, confirming its robustness and generalizability.
- New
- Research Article
- 10.1021/acsnano.6c05866
- Jun 20, 2026
- ACS nano
- Ik Joon Seo + 6 more
Here, we demonstrate an interface-controlled, self-rectifying resistive switching memory integrated in a 4K (64 × 64) crossbar array (CA). A simple Ru/HfAlOx/TiN stack composed of fabrication-friendly materials enables both nonvolatile resistive switching and polarity-dependent rectification. The interface-controlled operation suppresses the stochastic variability typically observed in conventional resistive switching memories and removes the need for electroforming, which is advantageous for mass production. In the 4K crossbar array, we achieve 100% functional yield without operational failure and experimentally verify analogue vector-matrix multiplication (VMM). The same interface-controlled switching yields analogue-like conductance updates with linear and symmetric modulation, and we present inference simulation results based on the measured characteristics. Finally, current-conduction fitting and drive-level capacitance profiling (DLCP), together with atomistic numerical simulations, elucidate a switching mechanism governed by the motion of internal mobile charges at the interfaces.
- New
- Research Article
- 10.1088/1361-6560/ae7957
- Jun 18, 2026
- Physics in Medicine & Biology
- Rohan Nadkarni + 6 more
Objective.Our goal was to develop a simulation platform for photon-counting CT (PCCT) imaging in mouse models of head and neck squamous cell carcinoma (HNSCC).Approach.High-resolution vasculature from an energy-integrating detector micro-CT scan of a barium-enhanced mouse was transferred to the mouse whole-body (MOBY) digital phantom using affine warps. To generate tumors with contrast agent distributions derived from real data, we trained a denoising diffusion probabilistic model (DDPM) on material-decomposed iodine- and barium-enhanced mouse tumors from a prior PCCT study. DDPM synthesized tumors were fused with the vascularized MOBY to create mouse HNSCC phantoms. We improved the accuracy of images from MATLAB PCCT simulation through an adjustment that utilizes matrix multiplication and a multi-layer perceptron (MLP) trained on matched real and simulated material phantoms. We passed MOBY HNSCC phantoms into the adjusted simulation, decomposed PCCT images into water, iodine, calcium, and barium maps, and compared these outputs to the true HNSCC phantoms using quantitative metrics from iodine- and barium-enhanced regions of the tumor.Main results.DDPM synthesized tumors had similar mean iodine and barium concentrations to real tumors. In a test set phantom, our matrix multiplication and MLP adjustment substantially reduced the root mean square error of attenuation measurements in reconstructed images from PCCT simulation. In this phantom, material decomposition of the adjusted image using a real sensitivity matrix produced similar material concentrations and cross-contamination patterns to those from real PCCT imaging. Material maps from adjusted simulations of MOBY HNSCC phantoms suggest that default PCCT settings slightly overestimated iodine content, while barium content was slightly underestimated in high barium tumors and overestimated in low barium tumors.Significance.This work established a PCCT simulation pipeline with ground truth digital mouse HNSCC phantoms, enabling evaluation of PCCT performance within a calibrated imaging configuration while minimizing radiation exposure to live mice.
- New
- Research Article
- 10.1038/s44172-026-00707-3
- Jun 17, 2026
- Communications engineering
- Amin Shafiee + 3 more
The physical scaling of photonic matrix-vector multiplication hardware for deep neural network acceleration is fundamentally limited by accumulated optical losses, crosstalk noise, and the prohibitive footprint of conventional devices such as Mach-Zehnder interferometers. Here we present LightPro, a fully programmable linear photonic processor designed to optimize scalability, power efficiency, and area footprint. At its core, our architecture integrates a neural architecture search and pruning framework with tunable phase-change material directional couplers. By thermally modulating the phase-change material state, we dynamically adjust coupling coefficients to achieve precise splitting ratios, facilitating highly optimized topologies for matrix-vector multiplication operations. The underlying phase-change material-based devices are evaluated using numerical multiphysics simulations and compact models, which are validated against reported experimental data from prior work. System-level evaluations demonstrate that the neural architecture search-optimized LightPro architectures achieve up to an 85% footprint reduction and a greater than 50% decrease in power consumption. Network scaling evaluations using handwritten digit and Gaussian datasets yield an inference accuracy degradation of less than 5%. Experimental prototyping on a commercial photonic processor validates the computational accuracy of LightPro, establishing a scalable and efficient pathway for next-generation photonic artificial intelligence accelerators.
- New
- Research Article
- 10.1109/tpami.2026.3704741
- Jun 17, 2026
- IEEE transactions on pattern analysis and machine intelligence
- Sihao Lin + 7 more
Semantic mismatch remains a key challenge in conventional knowledge distillation, where representational features are typically regressed from the teacher to the student in a one-to-one spatial matching fashion. In this paper, we address semantic mismatch by examining architectural differences between teacher and student networks. Specifically, due to the variations in network width and depth, the teacher network has a larger receptive field than the student, enabling it to integrate a broader spatial context. In contrast, the student model captures more localized features. This disparity exacerbates semantic misalignment. To alleviate this issue, we propose a novel one-to-all spatial matching knowledge distillation approach, wherein each pixel of the teacher's feature is distilled to all spatial locations of the student's feature map, weighted by a similarity map produced by a Target-aware Transformer (TaT). To further enhance TaT, we reduce its quadratic computational complexity and prevent incorrect spatial alignment, such as distilling background regions from the teacher to foreground regions in the student, and vice versa. In addition, we introduce the "looking broader" strategy, which rearranges the distilled representations of the student and teacher to align their receptive fields. This strategy is motivated by the observation that while individual pixels in student features typically have smaller receptive fields, aggregating multiple pixels can effectively bridge this gap. Therefore, we propose integrating feature pixels from multiple spatial positions using an efficient matrix multiplication. We validate our method through extensive experiments and demonstrate its superior performance and broad generalization capability across various backbone networks and vision tasks, including image classification, semantic segmentation, and object detection.
- Research Article
- 10.1016/j.saa.2026.127602
- Jun 5, 2026
- Spectrochimica acta. Part A, Molecular and biomolecular spectroscopy
- Ye He + 7 more
Multiple carbon quantum dots-enhanced excitation-emission matrix fluorescence strategy for accurate vintage discrimination of Anhua dark tea.
- Research Article
- 10.1364/oe.600141
- Jun 1, 2026
- Optics express
- Masato Shotoku + 4 more
Holographic displays are highly anticipated as a three-dimensional (3D) display technology that satisfies the essential requirements of human visual perception. However, the generation of computer-generated holograms (CGHs) is hindered by the significant challenges of high computational complexity and long processing times. In this study, we propose a method to reformulate the point-cloud method into a matrix multiplication framework. By representing the redundancy of object coordinates in the separable convolution method and the computational domain constraints of the wavefront recording plane method as sparse matrices, we eliminate redundant calculations and achieve efficient hologram generation. In addition, we further accelerate matrix-based hologram calculations by leveraging the matrix computation capabilities of modern graphics processing units.
- Research Article
- 10.1002/cbin.70165
- Jun 1, 2026
- Cell biology international
- Abir Salek + 10 more
Berberine Chloride Induces Apoptosis and Inhibits Adhesion, Migration, and Invasion in MDA-MB-231 and 4T1 Breast Cancer Cells: An Integrative In Vitro and In Silico Study.
- Research Article
- 10.1016/j.rineng.2026.110194
- Jun 1, 2026
- Results in Engineering
- David Zamora-Arranz + 3 more
Leveraging programmable logic controllers for machine learning applications in industrial setups
- Research Article
- 10.1021/acs.jctc.6c00175
- May 26, 2026
- Journal of chemical theory and computation
- Ryan Stocks + 3 more
We present the first implementation of double hybrid density functional theory (DHDFT) with all major computational steps accelerated on GPUs. To efficiently utilize GPU hardware, we employ the resolution-of-identity (RI) approximation, transforming electron repulsion integral (ERI) computations into large dense matrix multiplications and enabling basis functions up to g angular momentum. We demonstrate revDSD-PBEP86-D4(noFC)/def2-QZVPP calculations on the entire COMPAS-3x data set of ∼39,000 peri-condensed polybenzenoid hydrocarbon isomer geometries (up to 68 atoms) using just 900 node-hours on the Perlmutter supercomputer. For medium-sized organic molecules (up to ∼3k basis functions), the PT2 component adds minimal cost relative to the initial SCF step. This demonstrates that efficient GPU acceleration reduces the practical computational requirements of DHDFT comparable to conventional hybrid DFT. We additionally benchmark a range of LDA, GGA, and MGGA functionals against the revDSD-PBEP86-D4(noFC) isomerization energies. Without dispersion corrections, the SVWN5 LDA functional (MAD 4.47 kJ/mol) outperforms all tested GGAs and MGGAs. With dispersion corrections, only two MGGAs, led by M06-L-D4 (MAD 3.82 kJ/mol), are able to surpass the SVWN5 results.
- Research Article
- 10.1002/csr.70654
- May 26, 2026
- Corporate Social Responsibility and Environmental Management
- Sajid Ullah + 1 more
ABSTRACT Green innovation is a powerful tool for enhancing sustainability in organizations by focusing on environmental, social, and economic perspectives. Despite the benefits of green innovation, its implementation remains abysmal in developing countries. While previous research has identified various policies, it has often overlooked their combined impact hindering the successful adoption of green innovation. To address this lacuna, the current study empirically identified policies to promote green innovation adoption in an emerging economy by using the Pakistani manufacturing sector as a case. A unique approach integrating the fuzzy Delphi method (FDM), interpretive structural modeling (ISM), and cross‐impact matrix multiplication applied to classification (MICMAC) was developed to analyze the policies. First, green innovation policies were identified through an extensive literature review; they were then filtered using the Delphi method. Second, the ISM approach was used to meticulously adjudicate interactions between the identified policies. Finally, MICMAC was used to determine the driving and dependence power of policies. The study's findings indicate that “rules and regulations” and “green taxes and emission trading scheme” are the most significant policies for green innovation adoption while “big vision” and “corporate social responsibility” are the least important. The practitioners and manufacturing industry managers can promote green innovation initiatives by incorporating rules and regulations and an emission trading scheme.
- Research Article
- 10.1111/ajco.70121
- May 23, 2026
- Asia-Pacific journal of clinical oncology
- Gamze Sonmez + 3 more
Traditional approaches such as two-dimensional (2D) tumor models and animal studies often fall short in accurately replicating human cancer biology. In response to these limitations, three-dimensional (3D) bioprinting has emerged as a powerful platform for generating complex, physiologically relevant tumor models that better recapitulate the tumor microenvironment (TME). These bioprinted constructs enable precise spatial organization of multiple cell types and extracellular matrix components, allowing more faithful in vitro representation of tumor architecture and cellular interactions. Increasingly, patient-derived bioprinted models have demonstrated promise in personalized drug screening and therapeutic optimization, with applications reported in cancers such as glioblastoma, breast cancer, and hepatocellular carcinoma, highlighting their potential to improve predictive accuracy in preclinical testing and precision oncology. Despite these advances, significant challenges remain, including optimization of bioinks, reproducibility, scalability, and standardization of bioprinting workflows for broader clinical adoption. This review provides an integrated, clinician-oriented overview of 3D bioprinting in oncology, discussing its role in modeling tumor heterogeneity, angiogenesis, and metastasis, while critically evaluating current limitations and translational barriers. Emerging strategies-including smart bioinks, microfluidic integration, and four-dimensional (4D) bioprinting-are also examined as potential solutions to enhance functional complexity and clinical relevance, ultimately supporting the advancement of drug development pipelines and personalized cancer therapy.
- Research Article
- 10.1371/journal.pone.0349146
- May 20, 2026
- PLOS One
- Mateusz Gruzewski + 1 more
In this article, we present an efficient and concise OpenMP implementation of the Nussinov RNA folding algorithm, a well-known representative of non-serial polyadic dynamic programming (NPDP). Our goal is to develop an optimized implementation that can serve as a template for related dynamic programming applications. The proposed code is derived from a detailed analysis of manual implementations, emphasizing the separation of problematic and non-problematic instances and structuring computations in a way analogous to matrix multiplication. This design enables the semi-automatic extraction of data locality using tools based on Presburger arithmetic—techniques widely employed in classical loop transformations and advanced source-to-source compilers grounded in the polyhedral model. In the experimental evaluation, we assess the performance of our implementation on modern massively parallel AMD and Intel processors with 64, 128, and 192 threads. Our approach leverages cache-aware tiling, parallelism, and explicit vectorization to maximize computational efficiency, achieving performance that surpasses both automatically generated compiler-based solutions and manually tuned implementations on the evaluated platforms. Specifically, our implementation achieves execution times up to two orders of magnitude faster than polyhedral code, while also outperforming unvectorized manual approaches—being at least 30 × faster than array transposition–based methods and at least 5 × faster than the tiled sparsified Four Russians variant. Additionally, our results indicate that CPU implementations do not exhibit significantly worse performance compared to their corresponding GPU counterparts. These results demonstrate the importance of leveraging Advanced Vector Extensions (AVX) to fully exploit the capabilities of modern multi-core processors, particularly those in the AMD Epyc family.
- Research Article
- 10.1364/oe.593081
- May 18, 2026
- Optics express
- Chengli Chai + 4 more
Photonic processors offer a promising approach to accelerating matrix multiplications, which dominate the workload in AI computations. Photonic matrix multiplication processors based on time-division multiplexing (TDM) have attracted significant interest due to their superior scalability, which is critical for realizing large-scale matrices with a limited number of on-chip optical components. To further enhance parallelism, wavelength-division multiplexing (WDM) can be incorporated, with each wavelength carrying different vector information. While several prior studies have demonstrated such photonic matrix processors using both TDM and WDM, the (de)multiplexing of multiple wavelengths is typically performed off-chip. In this work, we propose and demonstrate a hybrid WDM-TDM photonic matrix-matrix multiplication processor, which employs microring resonator (MRR) arrays as on-chip wavelength (de)multiplexers, achieving a higher level of integration compared with approaches that use off-chip (de)multiplexers. A 3-wavelength, 12-channel circuit is fabricated on a Si-on-insulator (SOI) platform and applied to handwritten digit recognition, achieving a classification accuracy of 93.2% on 2000 images. Furthermore, the scalability of wavelength multiplexing and computational performance are analyzed, providing design guidelines for scaling the proposed architecture.
- Research Article
- 10.1073/pnas.2532022123
- May 15, 2026
- Proceedings of the National Academy of Sciences
- Ziao Wang + 14 more
Modern deep learning relies nearly exclusively on dedicated electronic hardware accelerators. Photonic approaches, with low consumption and high operation speed, are increasingly considered for inference but, to date, remain mostly limited to relatively basic tasks. Simultaneously, the problem of training deep and complex neural networks, overwhelmingly performed through backpropagation, remains a significant limitation to the size and, consequently, the performance of current architectures and a major compute and energy bottleneck. Here, we experimentally implement a versatile and scalable training algorithm, called direct feedback alignment, on a hybrid electronic-photonic platform. An optical processing unit performs large-scale random matrix multiplications, which is the central operation of this algorithm. We perform optical training of modern deep learning architectures, including Transformers, with more than 1B parameters, and obtain good performances on language, vision, and diffusion-based generative tasks. We study the scaling of the training time and demonstrate a potential advantage of our hybrid opto-electronic approach for ultra-deep and wide neural networks, thus opening a promising route to sustain the exponential growth of modern artificial intelligence beyond traditional von Neumann approaches.
- Research Article
- 10.1038/s41563-026-02609-3
- May 13, 2026
- Nature materials
- Haoran Qi + 21 more
Information units are progressively approaching the fundamental physical limits of integration density, including in terms of extremely small sizes, multistates and probabilistic traversal. However, simultaneously encompassing all of these characteristics in a unit remains elusive. Here, via real-time in situ electrical monitoring, we clearly observed stochastic alterations of multiple conductance states in Sc2C2@C88. The true random bit sequence generated exhibited an autocorrelation function whose confidence interval fell within ±0.02, demonstrating high-quality randomness. The alterations of multiple conductance states are controllable, that is, whose probability distributions could traverse from 0 to 1, enabling us to factorize 551 into its prime factors. Furthermore, we proposed a matrix-chain multiplication scheme and experimentally verified the multiplication of two 4 × 4 state-transition matrices with a small maximum error of <0.05. Combined with theoretical calculations, the stochastic but controllable multistates are probably attributed to the rich energy landscape, which could be stepwise changed by the electric field. Our findings reveal extremely small multilevel probabilistic bit for matrix multiplication, which pave the way for ultra-compact intelligent electronic devices.