Articles published on Parameter learning
Authors
Select Authors
Journals
Select Journals
Duration
Select Duration
4837 Search results
Sort by Recency
- New
- Research Article
- 10.1016/j.neunet.2026.108711
- Jul 1, 2026
- Neural networks : the official journal of the International Neural Network Society
- Mingyue Kong + 4 more
Trainable-parameter-free structural-diversity message passing for graph neural networks.
- New
- Research Article
- 10.1016/j.jbi.2026.105043
- Jul 1, 2026
- Journal of biomedical informatics
- Muhammad Aslanimoghanloo + 2 more
Generative modeling of clinical time series via latent stochastic differential equations.
- New
- Research Article
- 10.1016/j.aap.2026.108503
- Jul 1, 2026
- Accident; analysis and prevention
- Rui Shen + 3 more
The nonlinear impact of road safety policy implementation on the severity of road traffic crashes: A fusion of deep learning and Bayesian random parameter methods.
- New
- Research Article
- 10.1016/j.seppur.2026.137628
- Jul 1, 2026
- Separation and Purification Technology
- Luis Lillo Otarola + 4 more
Predicting filtration clogging behavior using machine learning and physicochemical parameters in final wine filtration processes
- New
- Research Article
- 10.1016/j.media.2026.104124
- Jul 1, 2026
- Medical image analysis
- Yiding Wang + 2 more
Continuous-time causal distribution learning with identifiability for brain dynamic effective connectivity inference.
- New
- Research Article
- 10.3390/app16136401
- Jun 26, 2026
- Applied Sciences
- Minghao Ma + 4 more
Wearable electrocardiogram (ECG) monitoring enables continuous cardiovascular assessment, yet signals acquired in ambulatory environments are inevitably corrupted by baseline wander, electrode motion artifacts, and muscle interference, which obscure diagnostically critical waveform features. Existing deep learning denoisers rely on heuristic attention mechanisms and time-domain-only losses, lacking principled control over what information the network retains or discards. To address this limitation, we propose EFIB-Net, an information bottleneck-guided multi-resolution network for robust ECG denoising. The framework introduces two complementary components: an efficient frequency-guided attention module that derives temporal attention weights directly from the energy distribution of parallel multi-resolution convolutional branches, requiring only four learnable parameters while providing physically interpretable feature selection that naturally highlights QRS complexes, and a variational information bottleneck constraint at the encoder–decoder bottleneck that forces the latent representation to retain only reconstruction-relevant information and discard noise, guided by a spectral–temporal composite loss. To the best of our knowledge, we are among the first to explicitly introduce the information bottleneck principle into deep-learning-based ECG signal denoising. Experiments on the MIT-BIH Arrhythmia Database show that EFIB-Net outperforms ten traditional and deep learning baselines across four standard metrics—signal-to-noise ratio (SNR), root mean square error, percentage root-mean-square difference, and correlation coefficient; at an input SNR of −5 dB it reaches 8.12 dB output SNR, surpassing the strongest attention-based competitor by 1.77 dB (p<0.01) while using only 0.45 M parameters and 10.8 ms inference latency per segment; downstream evaluation further demonstrates that the denoised signals achieve 99.18% R-peak detection sensitivity and 91.26% heartbeat classification F1-score, both within approximately one percentage point of the clean-signal upper bound, making it practical for real-time cardiac monitoring on resource-constrained wearable devices. Zero-shot cross-database evaluation on the QT Database further confirms generalizability, with only 0.54 dB degradation without retraining.
- New
- Research Article
- 10.1108/ecam-07-2025-1114
- Jun 26, 2026
- Engineering, Construction and Architectural Management
- M.K.C.S Wijewickrama + 3 more
Purpose This study aims to assess the applicability of the Bayesian belief network (BBN)-based conceptual information processing model for quality assurance (QA) within real-world reverse logistics supply chains (RLSCs) of demolition waste (DW) and to observe how it supports decision-making related to information processing for QA. Design/methodology/approach A multiple-case study strategy was adopted, focusing on two RLSCs in South Australia (SA). Focus group discussions were conducted within case studies to obtain expert-elicited probabilistic inferences for parametric learning in BBN-based modelling. A series of analyses was conducted using GeNIe software, including sensitivity analysis, root cause analysis and scenario analysis based on macro-, meso- and micro-level epistemic uncertainties. Findings The study confirmed that the BBN-based model is a useful decision-support tool for internal stakeholders in RLSCs. QA was most sensitive to micro-level workflow uncertainties and health and safety concerns, especially those related to the waste processor. Notably, RLSCs with small and medium-scale organisations were more vulnerable to epistemic uncertainties at all levels. Macro-level uncertainties emerged as root causes that propagate through the system and affect QA outcomes. Originality/value This is the first empirical application of a developed BBN-based information processing model specifically designed for QA in RLSCs of DW. This study contributes by demonstrating how this model can be operationalised as a decision-support tool, providing empirical insights into how epistemic uncertainties propagate through real-world supply chains.
- Research Article
- 10.21203/rs.3.rs-9916271/v1
- Jun 19, 2026
- Research square
- Dean Buonomano + 3 more
In computational neuroscience and deep learning models, synapses are viewed as simple computational elements with a single learnable parameter: synaptic strength. Here we propose that synapses are more sophisticated elements with multiple learnable parameters. Specifically, that the temporal profile of short-term synaptic plasticity can be learned to optimize the processing of temporal information. We confirm a prediction of this hypothesis by showing that in mouse and human neocortex, synaptic dynamics is not preferentially shaped by the identity of the pre- or postsynaptic neuron. Using computational approaches, we demonstrate that learned synaptic dynamics allows for interval-selectivity and counting at the synaptic level, and significantly enhances the performance of feedforward networks on complex spatiotemporal tasks. Our results provide a framework to understand the mismatch between the computational simplicity of synapses in models and their biochemical complexity, as well as address a fundamental gap between how the brain and artificial neural networks process temporal information.
- Research Article
- 10.1111/desc.70238
- Jun 17, 2026
- Developmental Science
- Karen E Smith + 4 more
ABSTRACTPeople learn most effectively when they can flexibly modify strategies to accommodate environmental changes. Here, we explore how chronic early life stress influences the ways individuals weight and prioritize new information when making decisions. To do so, we examined the choices of 11–16‐year‐old children in a reward learning task. Children varied on their level of stress exposure. We used a Bayesian computational learning model to assess whether differences in children's decision making were driven by difference in their expectations about the task environment. Children with high stress exposure switched their responses rather than maintaining choices that had previously been effective for them when making decisions about rewards. This switching was driven by children with high stress exposure expecting environments to be more volatile. These findings offer insight into how early life experiences can shape the subsequent parameters of human learning.SummaryHow experience shapes an individual's ability to flexibly adjust learning strategies when making decisions is not well understood.Using behavioral tasks combined with a novel computational model, we find exposure to high stress environments biases children towards learning strategies optimal to unpredictable environments.Stress in childhood shapes the development of learning parameters in ways that have implications for decision making.
- Research Article
- 10.3390/biomimetics11060418
- Jun 13, 2026
- Biomimetics (Basel, Switzerland)
- Hakan Işıker + 6 more
The power flow problem is one of the most challenging tasks in power systems, affecting both generation cost and energy quality. Optimal power flow (OPF) further complicates this task by requiring the optimal adjustment of system variables and parameters. This paper adapts the Modified Effective Butterfly Optimizer (MEBO) to solve multi-objective optimal power flow (MOOPF) problems with the contribution of optimized weighting using multiple Pareto archives. MEBO is an advanced optimization algorithm that utilizes population reduction and parameter learning to guide subsequent searches for unconstrained problems. The proposed technique has been tested on IEEE 30 and 57 bus test systems, and the results have been compared with existing methods reported in the literature. In the paper, four single-objective functions, namely generator cost, active power loss, fuel emission, and voltage deviation, are used to construct four multi-objective (MO) problems: cost-loss, cost-voltage, cost-emission, and emission-loss. For the cost-emission case, the proposed MEBO achieved compromised solutions of 791.1951 $/h fuel cost with 0.10873 ton/h emission and 801.8172 $/h fuel cost with 0.10044 ton/h emission under different Pareto-based optimization metrics. In the emission-loss case, the algorithm obtained 0.20539 ton/h emission with 3.1403 MW/h power loss, demonstrating the effectiveness of the proposed approach in balancing conflicting objectives. The Pareto curves of MEBO in achieving MO problems are presented, along with the suggested compromised solutions acquired from the literature. In the literature, this is the first application of MEBO for solving MOOPF problems. The results demonstrate that MEBO performs better than most other alternatives; this shows potential for further improvements with respect to the MOOPF problem.
- Research Article
- 10.1109/tpami.2026.3702798
- Jun 11, 2026
- IEEE transactions on pattern analysis and machine intelligence
- Jiang-Xin Shi + 2 more
The fine-tuning paradigm has emerged as a prominent approach for addressing long-tail learning tasks in the era of foundation models. However, the impact of fine-tuning strategies on long-tail learning performance remains unexplored. In this work, we disclose that existing paradigms exhibit a profound misuse of fine-tuning methods, leaving significant room for improvement in both efficiency and accuracy. Specifically, we reveal that heavy fine-tuning (fine-tuning a large proportion of model parameters) can lead to non-negligible performance deterioration on tail classes, whereas lightweight fine-tuning demonstrates superior effectiveness. Through comprehensive theoretical and empirical validation, we identify this phenomenon as stemming from inconsistent class conditional distributions induced by heavy fine-tuning. Building on this insight, we propose LIFT+, an innovative lightweight fine-tuning framework to optimize consistent class conditions. Furthermore, LIFT+ incorporates semantic-aware initialization, minimalist data augmentation, and test-time ensembling to enhance adaptation and generalization of foundation models. Our framework provides an efficient and accurate pipeline that facilitates fast convergence and model compactness. Extensive experiments demonstrate that LIFT+ significantly reduces both training epochs (from $\sim$100 to $\leq$15) and learned parameters (less than 1%), while surpassing state-of-the-art approaches by a considerable margin.
- Research Article
- 10.1038/s41598-026-56218-w
- Jun 11, 2026
- Scientific reports
- Siming Meng
To achieve precise diagnostic outcomes for gearbox systems, this study proposes an integrated methodology combining Orthogonalized Variational Mode Decomposition (OVMD) with Kernel Learning Incremental Relevance Vector Machine (KLIRVM). Capitalizing on the inherent efficacy of Bessel basis functions in detecting abrupt transient fluctuations, OVMD achieves pristine isolation of non-stationary characteristics stemming from damaged rolling elements. A novel KLIRVM approach is introduced and formulated for gearbox fault classification, wherein adaptive kernel parameter learning is seamlessly integrated with incremental updates-enabling each basis function to dynamically adjust its position and width for accurate characterization of locally varying signal features in streaming data. Ablation studies are conducted to validate the superiority of the proposed OVMD-KLIRVM framework. Empirical findings demonstrate that this approach outperforms competing methodologies in diagnostic accuracy.
- Research Article
- 10.64898/2026.06.10.731384
- Jun 10, 2026
- bioRxiv
- Shuming Liu + 4 more
Coarse-grained protein force fields enable simulations of biomolecular systems at length and time scales that are difficult to access with atomistic models, but achieving transferability across folded, intrinsically disordered, and multidomain proteins remains challenging. A central difficulty is that one-bead-per-residue models must represent chemically specific residue interactions while also absorbing solvent-mediated and many-body effects into a simplified energy function. Here, we present MOFF2, a transferable coarse-grained protein force field that combines residue-pair-specific interactions with a density-dependent many-body potential. MOFF2 is optimized using a two-stage strategy: bottom-up parameter learning from heterogeneous reference ensembles followed by refinement against experimental conformational observables. The resulting model provides balanced performance across ordered proteins, intrinsically disordered proteins, and multidomain proteins, and predicts condensate saturation-concentration trends for A1-LCD variant systems. Analysis of the learned parameters reveals chemically interpretable interaction patterns and density-dependent effects that explain the model’s improved transferability. These results demonstrate that combining a generalized coarse-grained energy function with data-driven optimization can produce a practical and interpretable force field for protein conformational and condensate simulations.
- Research Article
- 10.22219/kinetik.v11i3.2700
- Jun 7, 2026
- Kinetik: Game Technology, Information System, Computer Network, Computing, Electronics, and Control
- Eddy Nurraharjo + 3 more
Precision agriculture demands artificial intelligence solutions that are both accurate and deployable on resource-constrained hardware, yet conventional machine learning models require excessive memory while traditional ANFIS architectures suffer from training instability. This study developed a memory-efficient and gradient-stable lightweight Adaptive Neuro-Fuzzy Inference System (ANFIS) for real-time humidity prediction on microcontroller-class devices. The proposed architecture strategically reduced the rule base from 27 to only 4 interpretable fuzzy rules and limited membership functions to two per input, achieving an 85.2% reduction in learnable parameters. A gradient-stable training mechanism was introduced, combining physics-informed parameter initialization with adaptive gradient clipping to prevent gradient explosion. The model was trained and validated using 31,474 real-world greenhouse samples collected over 218 days, with 80% allocated for training and 20% for temporal testing. Experimental results demonstrated that the gradient-stable architecture successfully converged from a catastrophic R² of -64.08 to 0.9148, with a root mean square error of 1.32% and mean absolute error of 1.05%. The model required only 0.211 KB of memory, representing a 99.9% reduction compared to baseline Random Forest models, while achieving inference time of 8.2 milliseconds on Arduino UNO. The system was successfully deployed on three independent hardware modules, maintaining consistent performance with average RMSE of 1.99% over 168 hours of continuous operation. This study concludes that strategic simplification and stability-aware training enable interpretable neuro-fuzzy systems to operate effectively on ultra-low-resource devices, bridging the gap between predictive accuracy and hardware feasibility in embedded agricultural IoT applications.
- Research Article
- 10.1080/01630563.2026.2679010
- Jun 6, 2026
- Numerical Functional Analysis and Optimization
- Tatiana A Bubba + 2 more
In this paper, we revisit a supervised learning approach based on unrolling first introduced in [1] and called Ψ DONet, by providing a deeper microlocal interpretation for its theoretical analysis, and extending its study to the case of sparse-angle tomography. Furthermore, we refine the implementation of the original Ψ DONet considering special filters whose structure is specifically inspired by the streak artifact singularities characterizing tomographic reconstructions from incomplete data. This allows to considerably lower the number of (learnable) parameters while preserving (or even slightly improving) the same quality for the reconstructions from limited-angle data and providing a proof-of-concept for the case of sparse-angle tomographic data.
- Research Article
- 10.1080/00207179.2026.2679228
- Jun 2, 2026
- International Journal of Control
- Daniel Frank + 3 more
Neural networks can learn the behaviour of nonlinear dynamical systems and achieve high prediction accuracy on test data that is drawn from the same distribution as the training data. However, these models fail to generalise during inference when excited with unseen input trajectories. We tackle this generalisation problem in the context of nonlinear system identification with a system-theoretical approach that ensures input–output stability. We enhance a linear approximation with a recurrent neural network (RNN) that models the residual behaviour to capture complex dynamics and to increase the model's generalisation capabilities. We impose constraints on the learnable parameters to ensure dissipativity, an intrinsic property of most physical systems. This leads to improved generalisation on previously unseen inputs. We evaluate our approach using in-distribution (ID) and out-of-distribution (OOD) data from three different use cases and compare it against non-regularised approaches.
- Research Article
- 10.1109/tvcg.2026.3699816
- Jun 2, 2026
- IEEE transactions on visualization and computer graphics
- Yuhao Kang + 6 more
Corner cases, such as severe weather and abnormal lighting, present significant challenges in autonomous driving. The main obstacles involve large-scale data collection and costly annotations. Leveraging generative models to expand corner-case data based on existing annotations offers a promising solution. Unlike monocular videos, multi-view videos introduce an additional "view" dimension, increasing the consistency requirements and making precise control of annotations more challenging. Existing methods decouple multi-view videos along the temporal and view-spatial axes, using separate attention mechanisms, which causes motion discrepancies and limits consistency. Additionally, current approaches employ an independent adapter or ControlNet to encode different 3D annotations, leading to high computational costs and suboptimal alignment between annotations and video latents. These issues arise from neglecting the temporal-spatial relationship and insufficient alignment between 3D annotations and video latents. To address these challenges, we propose DriveGen, which uses 4D position embeddings to encode the positional information of multi-view videos. DriveGen also designs Dual-Scale Full Attention to ensure both global and local spatiotemporal consistency. Furthermore, our Shared Video-Condition Encoding (SVCE) Mechanism converts 3D annotations into 2D masks and encodes both video and annotation sequences using a 3D VAE, requiring only 0.37M learnable parameters to achieve pixel-level alignment and improving generation quality. Numerous experiments have proven that DriveGen has reached the state-of-the-art, capable of generating high-quality controlled autonomous driving videos.
- Research Article
- 10.1002/bimj.70134
- Jun 1, 2026
- Biometrical journal. Biometrische Zeitschrift
- Christoph Wiederkehr + 2 more
We evaluate the performance of targeted maximum likelihood estimation (TMLE) for estimating the average treatment effect in missing data scenarios under varying levels of positivity violations. We employ model- and design-based simulations, with the latter using undersmoothed highly adaptive lasso on the "WASH Benefits Bangladesh" data set to mimic real-world complexities. Five missingness-directed acyclic graphs are considered, capturing common missing data mechanisms in epidemiological research, particularly in one-point exposure studies. These mechanisms include also not-at-random missingness in the exposure, outcome, and confounders. We compare eight missing data methods in conjunction with TMLE as the analysis method, distinguishing between non-multiple imputation (non-MI) and multiple imputation (MI) approaches. The MI approaches use both parametric and machine learning models. Results show that non-MI methods, particularly complete cases with TMLE incorporating an outcome-missingness model, exhibit lower bias compared to all other evaluated missing data methods and greater robustness against positivity violations across. In comparison MI with classification and regression trees (CART) achieve lower root mean squared error, while often maintaining nominal coverage rates. Our findings highlight the trade-offs between bias and coverage, and we recommend using complete cases with TMLE incorporating an outcome-missingness model for bias reduction and MI CART when accurate confidence intervals are thepriority.
- Research Article
- 10.1016/j.mlwa.2026.100857
- Jun 1, 2026
- Machine Learning with Applications
- Gavin Lee Goodship + 2 more
Neural ordinary differential equations provide a principled continuous depth formulation, yet their practical use is often limited by slow, numerically fragile, and computationally expensive training caused by the weaknesses of standard ODEs solvers. We introduce extended stability Runge Kutta methods, which are explicit fixed step solvers designed to remain stable at much larger step sizes than classical schemes such as RK4 or adaptive methods such as Dormand Prince. By enlarging the stability region of the real axis, these solvers cross the integration horizon with far fewer steps, providing deterministic computation, greatly reduced function evaluation cost, and smoother optimization dynamics. Across CIFAR 10, CIFAR 100, and Tiny ImageNet, extended stability solvers match the final accuracy of higher-order methods, usually within one to two percent, while producing significantly more stable gradients. On CIFAR 10, they reduce gradient clipping rates to between zero and twenty-five percent depending on integration horizon, compared with sixty to one hundred percent for standard solvers, and they maintain reduced spectral amplification in moderately stiff regimes, while preserving coherent gradient flow in more extreme stiffness settings. The 15-stage variant achieves a fourteen-fold speedup in wall time relative to Dormand Prince 5 and a two-fold speedup relative to ResNet 20, while requiring eighteen times and twelve times fewer function evaluations, with comparable floating point operations to RK4. These gains require no additional learned parameters, no regularization, no architectural changes, and no adaptive tolerance tuning. They also allow successful optimization in regimes where classical explicit solvers become unstable or diverge. Overall, the results show that the geometry of the stability region, rather than formal order or adaptivity, is the key factor that governs gradient flow conditioning under a fixed integration horizon. Higher-order methods can reach similar accuracy but suffer from unstable and oscillatory gradients. Extended stability methods maintain coherent gradient flow while providing substantial practical speedups, making stability region geometry an effective and simple tool for accelerating and stabilizing the training of neural ordinary differential equation models. • Introduce extended-stability Runge–Kutta (ESRK) solvers to Neural ODE. • Demonstrate ESRK acts as a numerical preconditioner for gradient flow. • ESRK-15 achieves stable training with h = 30 at fixed compute cost. • ESRK-21 further smooths spectral dynamics in high-gain regimes ( Δ T = 20 –30). • Stability-region size, rather than solver order alone, governs gradient-flow conditioning in Neural ODEs training. • Provides FLOPs and wall-time scaling showing ≤ 1.4 × overhead vs RK4. • Provided detailed Resnet-20 comparison benchmarks across three datasets, along with geometric properties analysis.
- Research Article
- 10.1016/j.neucom.2026.133499
- Jun 1, 2026
- Neurocomputing
- Xianglin Wu + 2 more
SigMA: Path signatures and multi-head attention for learning parameters in fBm-driven SDEs