Placebo zones in discontinuity‐based designs: Estimation, inference, and implementation
This paper introduces a model-selection algorithm for regression discontinuity designs that evaluates candidate models within a "placebo zone" of the running variable, demonstrating asymptotic optimality and favorable performance in simulations, alongside a new randomization inference procedure implemented via Stata commands.
Abstract We propose a new model‐selection algorithm for regression discontinuity design and related estimators. The performance of candidate models is assessed within a “placebo zone” of the running variable. Candidate models can differ by bandwidth and other choice parameters. We outline (restrictive) sufficient conditions under which the approach is asymptotically optimal, and then show the approach also performs favorably under more general conditions in Monte Carlo simulations, including simulations calibrated to well‐known real‐world applications. We also propose a new randomization inference procedure which draws on the placebo estimates. Our Stata commands implement the procedure and compare its performance to other approaches.
- Research Article
3
- 10.1007/s40815-018-0509-0
- Jun 19, 2018
- International Journal of Fuzzy Systems
The sufficient conditions for satisfying the monotonicity property of the Takagi–Sugeno–Kang (TSK) fuzzy inference system (FIS) have shown to be useful in many different applications. However, the related sufficient and necessary conditions are still unknown. As such, even when the sufficient conditions are violated, the TSK FIS model may still able to satisfy the monotonicity property. Therefore, the monotonicity test is used as an approximated method to determine the validness of the monotonicity property. To the best of our knowledge, the use of the monotonicity test in FIS is new. In this paper, we focus on single-input zero-order TSK FIS with Gaussian fuzzy membership functions. An algorithm to test the monotonicity property, either accepting or rejecting an TSK FIS model of being monotone, is devised and analyzed. The relationship between TSK FIS and its capability of satisfying the monotonicity property, along with the sufficient conditions and the outcome of the monotonicity test, is established through a Monte Carlo simulation. The Monte Carlo simulation is a useful necessity test for the sufficient conditions. We define a necessity measure of the sufficient conditions (NMSC) as the probability that a randomly generated monotone TSK FIS model (evaluated based on the monotonicity test algorithm) satisfies the sufficient conditions. We empirically show that the NMSC score reduces with increasing number of fuzzy rules. In addition, an application of the TSK FIS model to failure modes and effects analysis is demonstrated. As compared with the sufficient conditions, a better FIS-based model with a lower error measure can be obtained using the monotonicity test. The outcome indicates the effectiveness of the monotonicity test for designing low-dimensional TSK FIS models with large numbers of fuzzy rules.
- Dissertation
1
- 10.17077/etd.j6r5my5q
- Aug 15, 2014
The selection of a best-subset regression model from a candidate family is a common problem that arises in many analyses. In best-subset model selection, we consider all possible subsets of regressor variables; thus, numerous candidate models may need to be fit and compared. One of the main challenges of best-subset selection arises from the size of the candidate model family: specifically, the probability of selecting an inappropriate model generally increases as the size of the family grows. For this reason, it is usually difficult to select an optimal model when best-subset selection is attempted based on a moderate to large number of regressor variables. Model selection criteria are often constructed to estimate discrepancy measures used to assess the disparity between each fitted candidate model and the generating model. The Akaike information criterion (AIC) and the corrected AIC (AICc) are designed to estimate the expected Kullback-Leibler (K-L) discrepancy. For best-subset selection, both AIC and AICc are negatively biased, and the use of either criterion will lead to overfitted models. To correct for this bias, we introduce a criterion AICi, which has a penalty term evaluated from Monte Carlo simulation. A multistage model selection procedure AICaps, which utilizes AICi, is proposed for best-subset selection. In the framework of linear regression models, the Gauss discrepancy is another frequently applied measure of proximity between a fitted candidate model and the generating model. Mallows’ conceptual predictive statistic (Cp) and the modified Cp (MCp) are designed to estimate the expected Gauss discrepancy. For best-subset selection, Cp and MCp exhibit negative estimation bias. To correct for this bias, we propose a criterion CPSi that again employs a penalty term evaluated from Monte Carlo simulation. We further devise a multistage procedure, CPSaps, which selectively utilizes CPSi.
- Research Article
8
- 10.3390/s20185332
- Sep 17, 2020
- Sensors (Basel, Switzerland)
Surrogate Modeling (SM) is often used to reduce the computational burden of time-consuming system simulations. However, continuous advances in Artificial Intelligence (AI) and the spread of embedded sensors have led to the creation of Digital Twins (DT), Design Mining (DM), and Soft Sensors (SS). These methodologies represent a new challenge for the generation of surrogate models since they require the implementation of elaborated artificial intelligence algorithms and minimize the number of physical experiments measured. To reduce the assessment of a physical system, several existing adaptive sequential sampling methodologies have been developed; however, they are limited in most part to the Kriging models and Kriging-model-based Monte Carlo Simulation. In this paper, we integrate a distinct adaptive sampling methodology to an automated machine learning methodology (AutoML) to help in the process of model selection while minimizing the system evaluation and maximizing the system performance for surrogate models based on artificial intelligence algorithms. In each iteration, this framework uses a grid search algorithm to determine the best candidate models and perform a leave-one-out cross-validation to calculate the performance of each sampled point. A Voronoi diagram is applied to partition the sampling region into some local cells, and the Voronoi vertexes are considered as new candidate points. The performance of the sample points is used to estimate the accuracy of the model for a set of candidate points to select those that will improve more the model’s accuracy. Then, the number of candidate models is reduced. Finally, the performance of the framework is tested using two examples to demonstrate the applicability of the proposed method.
- Research Article
100
- 10.1111/faf.12154
- Feb 23, 2016
- Fish and Fisheries
Multimodel frameworks are common in contemporary elasmobranch growth literature. These techniques offer a proposed improvement over individual growth functions by incorporating additional candidate models with alternative characteristics. Sigmoid functions (e.g. Gompertz and logistic) are a popular alternative to the commonly used von Bertalanffy growth function (VBGF) as they are hypothesized to better suit certain taxa based on body shape (such as batoids) or reproductive mode (such as egg‐layers). However, this hypothesis has never been tested. This study examined 74 elasmobranch multimodel growth studies by comparing the growth curves of their respective candidate models. Hypotheses regarding model performances were rejected as the VBGF was equally likely to fit best for all taxa and reproductive modes. Subsequently, no individual model was suited to be used a priori. Differences between candidate model fits were greatest at age zero with Gompertz and logistic functions providing estimates that were 15% and 23% larger on average than the VBGF, respectively. However, length‐at‐age estimates of the different models became negligible at older ages. Differences between candidate models were mostly small (≤5%), and the multimodel framework only marginally affected length‐at‐age estimates. However, there were cases where some candidate models provided inappropriate fits that contrasted considerably to the best fitting model. In some of these instances, a single‐model framework could have yielded biologically unrealistic growth estimates. Therefore, no study could pre‐empt whether or not it required a multimodel framework. A framework was subsequently recommended to maximize the accuracy of model fits for elasmobranch length‐at‐age estimates using multimodel approaches.
- Book Chapter
1
- 10.1007/bfb0006144
- Jan 1, 1982
An efficient algorithm is presented for generating D-optimal designs. The usual sequential D-optimal design algorithm embodies the principle of the greedy algorithm of combinatorial optimization. It is shown that a sufficient condition for applying the accelerated greedy algorithm of M. Minoux to the design problem is satisfied. The actual implementation of the accelerated sequential design algorithm is based on a more general sufficient condition. This allows the evaluation of quadratic forms to replace determinant evaluations. A heap type data structure provides additional efficiency. While the standard sequential design algorithm requires a number of basis function evaluations proportional to the number of iterations, the accelerated design algorithm computation is proportional to a much smaller sum of coefficients.
- Research Article
72
- 10.1103/physreve.86.011403
- Jul 9, 2012
- Physical Review E
We report on the diffusion of purely repulsive and freely rotating colloidal rods in the isotropic, nematic, and smectic liquid crystal phases to probe the agreement between Brownian and Monte Carlo dynamics under the most general conditions. By properly rescaling the Monte Carlo time step, being related to any elementary move via the corresponding self-diffusion coefficient, with the acceptance rate of simultaneous trial displacements and rotations, we demonstrate the existence of a unique Monte Carlo time scale that allows for a direct comparison between Monte Carlo and Brownian dynamics simulations. To estimate the validity of our theoretical approach, we compare the mean square displacement of rods, their orientational autocorrelation function, and the self-intermediate scattering function, as obtained from Brownian dynamics and Monte Carlo simulations. The agreement between the results of these two approaches, even under the condition of heterogeneous dynamics generally observed in liquid crystalline phases, is excellent.
- Research Article
20
- 10.1016/j.jhydrol.2023.130152
- Sep 21, 2023
- Journal of Hydrology
Adaptive selection and optimal combination scheme of candidate models for real-time integrated prediction of urban flood
- Research Article
83
- 10.1016/j.fishres.2004.08.013
- Oct 18, 2004
- Fisheries Research
Beyond ‘lognormal versus gamma’: discrimination among error distributions for generalized linear models
- Research Article
- 10.1080/00207728708963968
- Jan 1, 1987
- International Journal of Systems Science
The multiple-model adaptive filter (MMAF) method is applied to the estimation of error states of inertial navigation systems (INS). Monte Carlo simulations are performed to evaluate the sensitivity of several MMAFs to uncertainties in flight condition, where a Doppler radar receiver or Omega receiver is considered as the reference information source. It is shown that the MMAF method is useful not only for a case where the actual system model is included within the candidate models, but also for a case where the actual system model is not included within the candidate models.
- Research Article
5
- 10.1118/1.4735880
- Jun 1, 2012
- Medical Physics
Purpose: Realistic calculations of x‐ray projection images play an important role in many projects related to CBCT, such as the design of scanners and reconstruction algorithms. However, to yield a desired level of realism it usually requires a tremendously long computational time, which hinders the research process. The purpose of this project is to develop a realistic x‐ray projection image simulation package, gDRR, on a computer graphics processing unit (GPU) to achieve both high accuracy and efficiency. Methods: Primary signals in a projection is computed by GPU‐based ray‐tracing algorithms, with many features considered, e.g. source energy spectrum, fluence map due to a bowtie filter, and detector response. Scatter signals are obtained by Monte Carlo (MC) simulations on GPU under the aforementioned realistic setup followed by a denoising step. Noise signals are calculated by taking the difference between the MC simulated primary and the ray‐tracing primary signals, and the difference between the MC simulated scatter signals and the denoised scatter signals. The noise level is calibrated according to the mAs level in a scan. Results: The primary signals agree well with results from MC simulations. As for the scatter signals, our MC results for real patient cases are in good agreement with those from EGSnrc with less than 2% relative error using 10 million source photons. Various realistic artifacts can be observed in CBCT images reconstructed from the simulated projections, e.g. beam hardening, scatter, and noise. The computation time per projection is ∼20ms for primary signal per energy channel, ∼4sec for scatter signal, and 1∼2sec for all other steps. Conclusions: We have developed a complete simulation package on GPU to compute x‐ray projections in CBCT. Primary, scatter, and noise signals can be calculated with a high level of realism at a high efficiency. This work is supported in part by NIH (1R01CA154747‐01), Varian Medical Systems through a Master Research Agreement, and the Thrasher Research Fund.
- Research Article
- 10.20998/2079-0023.2023.01.13
- Jul 15, 2023
- Bulletin of National Technical University "KhPI". Series: System Analysis, Control and Information Technologies
The subject of research in the article is machine learning algorithms used for requirement elicitation technique selection. The goal of the work is to build effective parsimonious machine learning models to predict the using particular elicitation techniques in IT projects that allow using as few predictor variables as possible without a significant deterioration in the prediction quality. The following tasks are solved in the article: design an algorithm to build parsimonious machine learning candidate models for requirement elicitation technique selection based on gathered information on practitioners' experience, assess parsimonious machine learning model accuracy, and design an algorithm for the best candidate model selection. The following methods are used: algorithm theory, statistics theory, sampling techniques, data modeling theory, and science experiments. The following results were obtained: 1) parsimonious machine learning candidate models were built for the requirement elicitation technique selection. They included less number of features that helps in the future to avoid overfitting problems associated with the best-fit models; 2) according to the proposed algorithm for best candidate selection – a single parsimonious model with satisfied performance was chosen. Conclusion: An algorithm is proposed to build parsimonious candidate models for requirement elicitation technique selection that avoids the overfitting problem. The algorithm for the best candidate model selection identifies when a parsimonious model's performance is degraded and decides on the suitable model's selection. Both proposed algorithms were successfully tested with four datasets and can be proposed for their extensions to others.
- Research Article
25
- 10.1016/j.fishres.2010.11.008
- Nov 18, 2010
- Fisheries Research
Decreasing uncertainty in catch rate analyses using Delta-AdaBoost: An alternative approach in catch and bycatch analyses with high percentage of zeros
- Research Article
- 10.5334/pb.586
- Jan 1, 1976
- Psychologica Belgica
One of the purposes of the investigations reported here was to test Berlyne’s 1960 theory of arousal, i.e., in relation to learning. Among the fundamental hypotheses of this theory there is one implying that reinforcement depends on arousal regulation. This theory was translated into three mathematical models incorporating different assumptions about the relationship between arousal regulation and the effects of learning. Furthermore, on the basis of the assumption that reinforcement is not motivationally regulated but should be treated descriptively, two models were developed, matching all of the assumptions of two of the three motivational models except that of motivational regulation of reinforcement Because of the complexity of the models, they were investigated by means of inferential computer simulation. Parameters were estimated by Monte Carlo methods. The performance of the models was compared with the performance of a group of 50 human subjects in an instrumental learning task with a 100:0 CRF schedule. The performance of the motivational models proved to deviate at a high level of significance from the performance of the subjects (Tab. 2 & 3, Fig. 4). Moreover, the distribution of the performance of the motivational models is bimodal. This can be explained by the fact that most of the replications strive toward the opposite asymptote, closely approximating a criterion of several consecutive incorrect responses Further explorations showed that when negative reinforcement is disconnected from arousal regulation, performance is ameliorated (Fig. 5). This improvement is. however, rather small, so it must be concluded that the motivational models, and consequently the theory from which they are derived, are not useful for generating predictions about learning performance. The non-motivational models, on the other hand, show fairly good agreement with the performance of the subjects. Even a model with variable “learning operators” shows a reasonably good fit for some perseveration statistics as well as the general learning curve statistics (Tab. 2, 3 and 4) It is concluded that models with variable learning operators deserve further exploration, and that Berlyne’s arousal theory, as far as it is concerned with learning, seems to be defective in its most fundamental assumption.
- Research Article
13
- 10.1167/15.6.9
- May 7, 2015
- Journal of Vision
The purpose of this article is to provide mathematical insights into the results of some Monte Carlo simulations published by Tolhurst and colleagues (Clatworthy, Chirimuuta, Lauritzen, & Tolhurst, 2003; Chirimuuta & Tolhurst, 2005a). In these simulations, the contrast of a visual stimulus was encoded by a model spiking neuron or a set of such neurons. The mean spike count of each neuron was given by a sigmoidal function of contrast, the Naka-Rushton function. The actual number of spikes generated on each trial was determined by a doubly stochastic Poisson process. The spike counts were decoded using a Bayesian decoder to give an estimate of the stimulus contrast. Tolhurst and colleagues used the estimated contrast values to assess the model's performance in a number of ways, and they uncovered several relationships between properties of the neurons and characteristics of performance. Although this work made a substantial contribution to our understanding of the links between physiology and perceptual performance, the Monte Carlo simulations provided little insight into why the obtained patterns of results arose or how general they are. We overcame these problems by deriving equations that predict the model's performance. We derived an approximation of the model's decoding precision using Fisher information. We also analyzed the model's contrast detection performance and discovered a previously unknown theoretical connection between the Naka-Rushton contrast-response function and the Weibull psychometric function. Our equations give many insights into the theoretical relationships between physiology and perceptual performance reported by Tolhurst and colleagues, explaining how they arise and how they generalize across the neuronal parameter space.
- Research Article
14
- 10.1145/3379484
- May 27, 2020
- Proceedings of the ACM on Measurement and Analysis of Computing Systems
We study online optimization in a setting where an online learner seeks to optimize a per-round hitting cost, which may be non-convex, while incurring a movement cost when changing actions between rounds. We ask:under what general conditions is it possible for an online learner to leverage predictions of future cost functions in order to achieve near-optimal costs? Prior work has provided near-optimal online algorithms for specific combinations of assumptions about hitting and switching costs, but no general results are known. In this work, we give two general sufficient conditions that specify a relationship between the hitting and movement costs which guarantees that a new algorithm, Synchronized Fixed Horizon Control (SFHC), achieves a 1+O(1/w) competitive ratio, where w is the number of predictions available to the learner. Our conditions do not require the cost functions to be convex, and we also derive competitive ratio results for non-convex hitting and movement costs. Our results provide the first constant, dimension-free competitive ratio for online non-convex optimization with movement costs. We also give an example of a natural problem, Convex Body Chasing (CBC), where the sufficient conditions are not satisfied and prove that no online algorithm can have a competitive ratio that converges to 1.