Incorporating Measurement Errors in Fixed Person Parameter Calibration
This study introduces a fixed person parameter calibration method that incorporates measurement error using a Bayesian iterative approach based on a mixed-effect structural equation model. Simulation results show it consistently outperforms existing FPC and FIC methods across various sample sizes, item compositions, and ability distributions, highlighting the importance of accounting for measurement error in small-sample item calibration.
Abstract In this study, we propose a new fixed person parameter calibration (FPC) strategy that incorporates measurement error in examinee ability estimates. Specifically, the proposed FPC method is an application of the mixed‐effect structural equation model of Junker et al. (2012) to the small‐sample item calibration context and relies on a Bayesian iterative sampling procedure for parameter estimation. We evaluated the proposed FPC method using simulated data sets that varied in terms of sample size, item composition, and examinee ability distribution. The parameter recovery performance of the proposed method was compared to those from alternative small‐sample calibration methods two other FPC methods and the state‐of‐the‐art fixed item parameter calibration (FIC) method. The results from the simulation study showed that the proposed method consistently outperformed the compared FPC and FIC methods. The encouraging performance of the proposed method demonstrates the impact of properly accounting for measurement error and provides a justification for its use as a competent small‐sample item calibration method.
- Research Article
3
- 10.3352/jeehp.2024.21.23
- Sep 12, 2024
- Journal of Educational Evaluation for Health Professions
Computerized adaptive testing (CAT) has become a widely adopted test design for high-stakes licensing and certification exams, particularly in the health professions in the United States, due to its ability to tailor test difficulty in real time, reducing testing time while providing precise ability estimates. A key component of CAT is item response theory (IRT), which facilitates the dynamic selection of items based on examinees' ability levels during a test. Accurate estimation of item and ability parameters is essential for successful CAT implementation, necessitating convenient and reliable software to ensure precise parameter estimation. This paper introduces the irtQ R package (http://CRAN.R-project.org/), which simplifies IRT-based analysis and item calibration under unidimensional IRT models. While it does not directly simulate CAT, it provides essential tools to support CAT development, including parameter estimation using marginal maximum likelihood estimation via the expectation-maximization algorithm, pretest item calibration through fixed item parameter calibration and fixed ability parameter calibration methods, and examinee ability estimation. The package also enables users to compute item and test characteristic curves and information functions necessary for evaluating the psychometric properties of a test. This paper illustrates the key features of the irtQ package through examples using simulated datasets, demonstrating its utility in IRT applications such as test data analysis and ability scoring. By providing a user-friendly environment for IRT analysis, irtQ significantly enhances the capacity for efficient adaptive testing research and operations. Finally, the paper highlights additional core functionalities of irtQ, emphasizing its broader applicability to the development and operation of IRT-based assessments.
- Research Article
2
- 10.1002/ets2.12376
- Feb 4, 2024
- ETS Research Report Series
The multistage testing (MST) design has been gaining attention and popularity in educational assessments. For testing programs that have small test‐taker samples, it is challenging to calibrate new items to replenish the item pool. In the current research, we used the item pools from an operational MST program to illustrate how research studies can be built upon literature and program‐specific data to help to fill the gaps between research and practice and to make sound psychometric decisions to address the small‐sample issues. The studies included choice of item calibration methods, data collection designs to increase sample sizes, and item response theory models in producing the score conversion tables. Our results showed that, with small samples, the fixed parameter calibration (FIPC) method performed consistently the best for calibrating new items, compared to the traditional separate‐calibration with scaling method and a new approach of a calibration method based on the minimum discriminant information adjustment. In addition, the concurrent FIPC calibration with data from multiple administrations also improved parameter estimation of new items. However, because of the program‐specific settings, a simpler model may not improve current practice when the sample size was small and when the initial item pools were well‐calibrated using a two‐parameter logistic model with a large field trial data.
- Research Article
45
- 10.1111/j.1365-2435.2009.01596.x
- Nov 9, 2009
- Functional Ecology
Summary 1. Phylogenetic signal – the similarity in trait values among phylogenetically related species – is pervasive for most types of traits in most organisms. Traits can often be categorized a priori into groups based on the level of biological organization, functional relations, developmental origins, or genetic underpinnings. Traits within such groups are often expected to be correlated and hence show similar levels of phylogenetic signal. 2. We developed multivariate statistical methods to test for phylogenetic signal in groups of traits while also incorporating estimates of trait measurement error (including within‐species variation) that can obscure phylogenetic signal. Simultaneously, these methods produce estimates of correlations between traits that are corrected for phylogenetic relationships among species. 3. We applied these methods to data for 13 morphological and physiological traits gathered in a common‐garden study of nine species of Manglietia (Magnoliaceae). The 13 traits fell into four groups: three traits involved photosynthesis [maximum net photosynthesis (Amax), light saturation point (LSP), light compensation point]; three described leaf morphology (thickness of leaves, palisade tissue, sponge tissue); four related to plant growth (basal stem diameter, crown volume, leaf area, relative growth rate); and three measured thermal tolerance [critical temperature (Tch), peak temperature (Tmax), temperature of half‐inactivation (T50)]. We also constructed a molecular phylogeny for these species from 219 AFLP markers via maximum likelihood estimation under the assumption of sequential binary changes in DNA sequences. 4. Of the 13 traits, only two photosynthesis traits (Amax and LSP) exhibited statistically detectable phylogenetic signal (P < 0·05) when analysed separately, whether using previously published univariate tests or our new univariate tests that incorporate measurement error. In contrast, multivariate analyses of the four trait groups, estimating simultaneously the phylogenetic signal for all traits and the correlations between traits, revealed a statistically significant phylogenetic signal for two of the four groups (photosynthesis and plant growth), comprising seven traits in total. 5. Our results demonstrate that even when the number of species in a comparative study is small, resulting in low power for univariate tests, phylogenetic signal can nonetheless be detected with multivariate tests that incorporate measurement error. Furthermore, our simulations show that the joint estimation of phylogenetic signal and trait correlations can lead to better (less biased and more precise) estimates of both.
- Research Article
- 10.1016/j.engappai.2026.114213
- May 1, 2026
- Engineering Applications of Artificial Intelligence
A novel multisource-intelligent calibration method for discrete element model parameters and application in macro-mesoscopic strength analysis of slope soil
- Research Article
501
- 10.1080/10635150701313830
- Apr 1, 2007
- Systematic Biology
Most phylogenetically based statistical methods for the analysis of quantitative or continuously varying phenotypic traits assume that variation within species is absent or at least negligible, which is unrealistic for many traits. Within-species variation has several components. Differences among populations of the same species may represent either phylogenetic divergence or direct effects of environmental factors that differ among populations (phenotypic plasticity). Within-population variation also contributes to within-species variation and includes sampling variation, instrument-related error, low repeatability caused by fluctuations in behavioral or physiological state, variation related to age, sex, season, or time of day, and individual variation within such categories. Here we develop techniques for analyzing phylogenetically correlated data to include within-species variation, or "measurement error" as it is often termed in the statistical literature. We derive methods for (i) univariate analyses, including measurement of "phylogenetic signal," (ii) correlation and principal components analysis for multiple traits, (iii) multiple regression, and (iv) inference of "functional relations," such as reduced major axis (RMA) regression. The methods are capable of incorporating measurement error that differs for each data point (mean value for a species or population), but they can be modified for special cases in which less is known about measurement error (e.g., when one is willing to assume something about the ratio of measurement error in two traits). We show that failure to incorporate measurement error can lead to both biased and imprecise (unduly uncertain) parameter estimates. Even previous methods that are thought to account for measurement error, such as conventional RMA regression, can be improved by explicitly incorporating measurement error and phylogenetic correlation. We illustrate these methods with examples and simulations and provide Matlab programs.
- Dissertation
1
- 10.17077/etd.fb5dwk9t
- Sep 27, 2017
<p>For unidimensional item response theory (UIRT) models, three linking methods, which are the separate, concurrent, and fixed parameter calibration methods, have been developed and widely used in applications such as vertical scaling, differential item functioning, computerized adaptive testing (CAT), and equating. By contrast, even though a few studies have compared the separate and concurrent calibration methods for full multidimensional IRT (MIRT) models or applied the concurrent calibration method to vertical scaling using the bifactor model, no study has yet provided technical descriptions of the concurrent and fixed parameter calibration methods for any MIRT models. Thus, the purpose of this dissertation was to extend the concurrent and fixed parameter calibration methods for UIRT models to the two-tier item factor analysis model. In addition, the relative performance of the separate, concurrent, and fixed parameter calibration methods was compared in terms of the recovery of item parameters and accuracy of IRT observed score equating using both real and simulated datasets.</p> <p>The separate, concurrent, and fixed parameter calibration methods well recovered the item parameters, with the concurrent calibration method performing slightly better than the other two linking methods. Despite the comparable performance of the three linking methods in terms of the recovery of item parameters, however, some discrepancy was observed between the IRT observed score equating results obtained with the three linking methods. In general, the concurrent calibration method provided equating results with the smallest equating error, whereas the separate calibration method provided equating results with the largest equating error due to the largest standard error of equating. The performance of the fixed parameter calibration method depended on the proportion of common items. When the proportion was , the fixed parameter calibration method provided more biased equating results than the concurrent calibration method because of the underestimated specific slope parameters. However, when the proportion of common items was 40%, the fixed parameter calibration method worked as well as the concurrent calibration method.</p>
- Research Article
2
- 10.1140/epjc/s10052-023-12284-2
- Nov 29, 2023
- The European Physical Journal C
This paper investigates the self-similar solutions of the Einstein-axion-dilaton configuration from type IIB string theory and the global SL(2,R) symmetry. We consider the Continuous Self Similarity (CSS), where the scale transformation is controlled by an SL(2, R) boost or hyperbolic translation. The solutions stay invariant under the combination of space-time dilation with internal SL(2,R) transformations. We develop a new formalism based on Sequential Monte Carlo (SMC) and artificial neural networks (NNs) to estimate the self-similar solutions to the equations of motion in the hyperbolic class in four dimensions. Due to the complex and highly nonlinear patterns, researchers typically have to use various constraints and numerical approximation methods to estimate the equations of motion; thus, they have to overlook the measurement errors in parameter estimation. Through a Bayesian framework, we incorporate measurement errors into our models to find the solutions to the hyperbolic equations of motion. It is well known that the hyperbolic class suffers from multiple solutions where the critical collapse functions have overlap domains for these solutions. To deal with this complexity, for the first time in literature on the axion-dilaton system, we propose the SMC approach to obtain the multi-modal posterior distributions. Through a probabilistic perspective, we confirm the deterministic α\\documentclass[12pt]{minimal} \\usepackage{amsmath} \\usepackage{wasysym} \\usepackage{amsfonts} \\usepackage{amssymb} \\usepackage{amsbsy} \\usepackage{mathrsfs} \\usepackage{upgreek} \\setlength{\\oddsidemargin}{-69pt} \\begin{document}$$\\alpha $$\\end{document} and β\\documentclass[12pt]{minimal} \\usepackage{amsmath} \\usepackage{wasysym} \\usepackage{amsfonts} \\usepackage{amssymb} \\usepackage{amsbsy} \\usepackage{mathrsfs} \\usepackage{upgreek} \\setlength{\\oddsidemargin}{-69pt} \\begin{document}$$\\beta $$\\end{document} solutions available in the literature and determine all possible solutions that may occur due to measurement errors. We finally proposed the penalized Leave-One-Out Cross-validation (LOOCV) to combine the Bayesian NN-based estimates optimally. The approach enables us to determine the optimum weights while dealing with the co-linearity issue in the NN-based estimates and better predict the critical functions corresponding to multiple solutions of the equations of motion.
- Research Article
26
- 10.1007/s40572-017-0160-1
- Oct 5, 2017
- Current Environmental Health Reports
Outdoor air pollution exposures used in epidemiological studies are commonly predicted from spatiotemporal models incorporating limited measurements, temporal factors, geographic information system variables, and/or satellite data. Measurement error in these exposure estimates leads to imprecise estimation of health effects and their standard errors. We reviewed methods for measurement error correction that have been applied in epidemiological studies that use model-derived air pollution data. We identified seven cohort studies and one panel study that have employed measurement error correction methods. These methods included regression calibration, risk set regression calibration, regression calibration with instrumental variables, the simulation extrapolation approach (SIMEX), and methods under the non-parametric or parameter bootstrap. Corrections resulted in small increases in the absolute magnitude of the health effect estimate and its standard error under most scenarios. Limited application of measurement error correction methods in air pollution studies may be attributed to the absence of exposure validation data and the methodological complexity of the proposed methods. Future epidemiological studies should consider in their design phase the requirements for the measurement error correction method to be later applied, while methodological advances are needed under the multi-pollutants setting.
- Conference Article
3
- 10.1117/12.864145
- Aug 19, 2010
- Proceedings of SPIE, the International Society for Optical Engineering/Proceedings of SPIE
Our laboratory has investigated the efficacy of a suite of color calibration and monitor profiling packages which employ a variety of color measurement sensors. Each of the methods computes gamma correction tables for the red, green and blue color channels of a monitor that attempt to: a) match a desired luminance range and tone reproduction curve; and b) maintain a target neutral point across the range of grey values. All of the methods examined here produce International Color Consortium (ICC) profiles that describe the color rendering capabilities of the monitor after calibration. Color profiles incorporate a transfer matrix that establishes the relationship between RGB driving levels and the International Commission on Illumination (CIE) XYZ (tristimulus) values of the resulting on-screen color; the matrix is developed by displaying color patches of known RGB values on the monitor and measuring the tristimulus values with a sensor. The number and chromatic distribution of color patches varies across methods and is usually not under user control. In this work we examine the effect of employing differing calibration and profiling methods on rendition of color images. A series of color patches encoded in sRGB color space were presented on the monitor using color-management software that utilized the ICC profile produced by each method. The patches were displayed on the calibrated monitor and measured with a Minolta CS200 colorimeter. Differences in intended and achieved luminance and chromaticity were computed using the CIE DE2000 color-difference metric, in which a value of ΔE = 1 is generally considered to be approximately one just noticeable difference (JND) in color. We observed between one and 17 JND's for individual colors, depending on calibration method and target. As an extension of this fundamental work<sup>1</sup>, we further improved our calibration method by defining concrete calibration parameters for the display, using the NEC wide gamut puck, and making sure that those calibration parameters did conform, with the help of a state of the art Spectroradiometer, PR670. As a result of this addition of the PR670, and also an in-house developed method of profiling and characterization, it appears that there was much improvement in ΔE, the color difference.
- Research Article
7
- 10.1093/forestry/cpad063
- Dec 12, 2023
- Forestry: An International Journal of Forest Research
Height–diameter (H–D) models are fundamental tools for predicting the relationship between tree H–D at breast height, for numerous applications in forestry. Increasingly, studies develop H–D models that can be calibrated to achieve a high level of precision with only a few observations. Different calibration methods and strategies are employed and compared in these studies, often disregarding the data used to develop the models and the H–D function used. In this study, we examined the transferability of optimal calibration strategies across studies, conducting a literature review and an empirical study. We compared the performance of six H–D functions and different calibration methods when using the same calibration strategies and dataset. Based on our literature review, we found that the most commonly employed calibration strategy is random-effects calibration and that the most common variable used to develop generalized H–D models is dominant height. We observed that different calibration methods can lead to varying results due to their different emphases on various aspects of the data and their individual limitations. Moreover, when the same dataset is used for calibration, different H–D functions may exhibit various performances. However, we found high percentages of agreement for the Curtis, Schumacher, and Wykoff H–D functions across all three calibration methods and low agreement between all functions and the Power H–D function. These observations underscore the need to consider all relevant factors, including the H–D function used, when selecting an H–D function and calibration strategy to ensure optimal transferability of the model. Our study provides insights that can improve the accuracy of H–D models, which are essential for predicting forest growth and structure in the context of changing environmental conditions.
- Research Article
5
- 10.3390/math8020162
- Jan 23, 2020
- Mathematics
The Burr type XII (BurrXII) distribution is very flexible for modeling and has earned much attention in the past few decades. In this study, the maximum likelihood estimation method and two Bayesian estimation procedures are investigated based on constant-stress accelerated life test (ALT) samples, which are obtained from the doubly truncated three-parameter BurrXII distribution. Because computational difficulty occurs for maximum likelihood estimation method, two Bayesian procedures are suggested to estimate model parameters and lifetime quantiles under the normal use condition. A Markov Chain Monte Carlo approach using the Metropolis–Hastings algorithm via Gibbs sampling is built to obtain Bayes estimators of the model parameters and to construct credible intervals. The proposed Bayesian estimation procedures are simple for practical use, and the obtained Bayes estimates are reliable for evaluating the reliability of lifetime products based on ALT samples. Monte Carlo simulations were conducted to evaluate the performance of these two Bayesian estimation procedures. Simulation results show that the second Bayesian estimation procedure outperforms the first Bayesian estimation procedure in terms of bias and mean squared error when users do not have sufficient knowledge to set up hyperparameters in the prior distributions. Finally, a numerical example about oil-well pumps is used for illustration.
- Research Article
70
- 10.1002/2013jc009705
- Jul 1, 2014
- Journal of Geophysical Research: Oceans
The choice of parameter values is crucial in the course of sea ice model development, since parameters largely affect the modeled mean sea ice state. Manual tuning of parameters will soon become impractical, as sea ice models will likely include more parameters to calibrate, leading to an exponential increase of the number of possible combinations to test. Objective and automatic methods for parameter calibration are thus progressively called on to replace the traditional heuristic, “trial‐and‐error” recipes. Here a method for calibration of parameters based on the ensemble Kalman filter is implemented, tested and validated in the ocean‐sea ice model NEMO‐LIM3. Three dynamic parameters are calibrated: the ice strength parameter P*, the ocean‐sea ice drag parameter Cw, and the atmosphere‐sea ice drag parameter Ca. In twin, perfect‐model experiments, the default parameter values are retrieved within 1 year of simulation. Using 2007–2012 real sea ice drift data, the calibration of the ice strength parameter P* and the oceanic drag parameter Cw improves clearly the Arctic sea ice drift properties. It is found that the estimation of the atmospheric drag Ca is not necessary if P* and Cw are already estimated. The large reduction in the sea ice speed bias with calibrated parameters comes with a slight overestimation of the winter sea ice areal export through Fram Strait and a slight improvement in the sea ice thickness distribution. Overall, the estimation of parameters with the ensemble Kalman filter represents an encouraging alternative to manual tuning for ocean‐sea ice models.
- Conference Article
1
- 10.1109/ichceswidr54323.2021.9656297
- Nov 6, 2021
The parameters of the hydrological model are important parts of the model, and parameter calibration plays a key role in the output of the model. To analyze the influence of different calibration methods on the simulation results of the VIC hydrological model, the parameter calibration methods are coupled with the VIC model combining the characteristics of the Qinhuai River Basin. Three automatic optimization methods, including the genetic algorithm, simulated annealing algorithm and SCE-UA(shuffle complex evolution algorithm) algorithm, are used to automatically calibrate the 6 parameters of the VIC model. Three evaluation methods of the Nash efficiency coefficient (NSE), mean square error (MSE) and percentage deviation (PBIAS) are used to optimize the parameters. The model’s simulated runoff results are compared and analyzed. The results show that the SCE-UA algorithm with the Nash efficiency coefficient as the objective function can make the VIC model parameter calibration achieve ideal results. Different objective functions have different effects on the peak runoff during the model simulation verification period. By optimizing the parameters with the Nash efficiency coefficient and mean square error as objective functions, the simulated peak runoff is more similar to the measured data. The results can provide a reference for other models parameters automatic calibration, and are of great practical significance for rapid emergency response of hydrological simulation and forecasting.
- Research Article
16
- 10.1016/j.powtec.2022.117860
- Sep 1, 2022
- Powder Technology
A strain energy-based elastic parameter calibration method for lattice/bonded particle modelling of solid materials
- Research Article
34
- 10.1007/bf02296133
- Sep 1, 1994
- Psychometrika
Hierarchical Bayes procedures for the two-parameter logistic item response model were compared for estimating item and ability parameters. Simulated data sets were analyzed via two joint and two marginal Bayesian estimation procedures. The marginal Bayesian estimation procedures yielded consistently smaller root mean square differences than the joint Bayesian estimation procedures for item and ability estimates. As the sample size and test length increased, the four Bayes procedures yielded essentially the same result.