Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

3PL with Ability‐Expression Gap: Modeling the Discrepancy between Latent and Expressed Ability

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

Abstract This article proposes the Ability‐Attenuated 3PL (AA‐3PL), an extension of the three‐parameter logistic model designed to represent systematic discrepancies between latent ability and expressed performance (the ability‐expression gap). The model retains the conventional interpretation of 3PL item parameters while adding a parsimonious mechanism to capture structured performance attenuation that can arise from non‐ability influences in longitudinal and applied testing contexts. Evidence is established through a Bayesian simulation program spanning multiple study modules, including baseline multistage conditions, shape‐misspecification stress tests, identification‐constraint checks, a growth‐related extension, a one‐stage boundary case, and prior‐sensitivity analyses. Across attenuation conditions, AA‐3PL consistently reduces difficulty misattribution and improves overall model performance, while showing negligible differences from 3PL when no gap is present; robustness depends on aligning the attenuation shape with the underlying data pattern. An empirical two‐wave WJ‐PV illustration provides convergent support: AA‐family models better capture follow‐up systematic deviations than 3PL, with the shift‐augmented extension offering the most coherent improvements when cross‐wave shifts are present.

Similar Papers
  • Dissertation
  • Cite Count Icon 1
  • 10.17077/etd.005860
Neural network methods for application in educational measurement
  • Aug 1, 2021
  • Geoffrey Converse

In educational measurement, Item Response Theory (IRT) provides a means of quantifying student knowledge. Specifically, IRT models the probability of a student answering a particular item correctly as a function of the student’s continuous-valued latent abilities Θ (e.g. add, subtract, multiply, divide) and parameters associated with the item (e.g. difficulty). Given a group of students’ binary responses (correct/incorrect) to an assessment, parameter estimation techniques are used to infer the students’ abilities Θ to better evaluate student performance. But as the number of students, items, and dimension of Θ increases, traditional parameter estimation methods which rely on numerical integration or MCMC techniques become infeasible. In this thesis, a novel modification to a variational autoencoder (VAE), an unsupervised learning method, is presented which incorporates multiple pieces of domain knowledge into the VAE architecture. Specifically, an expert-annotated Q-matrix detailing the association between items and abilities is used to constrain the trainable weights in the VAE decoder. This directly links the generative posterior distribution of a VAE to IRT models, and allows interpretation of trainable weights as item parameters and a hidden neural layer as student ability estimates. The use of a neural network in this application allows for efficient estimation of parameters in high-dimensional datasets. An additional alteration to the VAE encoder allows for modeling of correlated, rather than independent, latent abilities. This architecture is also of interest to areas outside of education where the probability distribution of the VAE latent code is known. The proposed method, titled ML2P-VAE, achieves accuracy similar to traditional parameter estimation methods on smaller datasets and provides quality estimates on educational assessments with a large number of students, items, and latent abilities, where traditional methods struggle. Modifications used in ML2P-VAE are extended to integrate IRT into deep knowledge tracing models, which use time-dependent neural networks such as long short-term memory networks and Transformers. In an online environment where many items are available, the task of knowledge tracing is to track a student’s learning dynamically as they progress through an assessment. This task has received more attention recently due to the emergence of online learning and AI tutoring systems, where following student’s knowledge acquisition can help facilitate automated curriculum in real time. The inclusion of IRT in the knowledge tracing framework presents a trade-off between explainability and accuracy, while remaining competitive with state-of-the-art methods.

  • Research Article
  • Cite Count Icon 27
  • 10.1007/bf02294342
Statistical Inference Based on Latent Ability Estimates
  • Jun 1, 1996
  • Psychometrika
  • Herbert Hoijtink + 1 more

The quality of approximations to first and second order moments (e.g., statistics like means, variances, regression coefficients) based on latent ability estimates is being discussed. The ability estimates are obtained using either the Rasch, or the two-parameter logistic model. Straightforward use of such statistics to make inferences with respect to true latent ability is not recommended, unless we account for the fact that the basic quantities are estimates. In this paper true score theory is used to account for the latter; the counterpart of observed/true score being estimated/true latent ability. It is shown that statistics based on the true score theory are virtually unbiased if the number of items presented to each examinee is larger than fifteen. Three types of estimators are compared: maximum likelihood, weighted maximum likelihood, and Bayes modal. Furthermore, the (dis)advantages of the true score method and direct modeling of latent ability is discussed.

  • Research Article
  • Cite Count Icon 89
  • 10.1097/aog.0000000000000094
A Model for Predicting the Risk of De Novo Stress Urinary Incontinence in Women Undergoing Pelvic Organ Prolapse Surgery
  • Feb 1, 2014
  • Obstetrics & Gynecology
  • J Eric Jelovsek + 14 more

To construct and validate a prediction model for estimating the risk of de novo stress urinary incontinence (SUI) after vaginal pelvic organ prolapse (POP) surgery and compare it with predictions using preoperative urinary stress testing and expert surgeons' predictions. Using the data set (n=457) from the Outcomes Following Vaginal Prolapse Repair and Midurethral Sling trial, a model using 12 clinical preoperative predictors of de novo SUI was constructed. De novo SUI was determined by Pelvic Floor Distress Inventory responses through 12 months postoperatively. After fitting the multivariable logistic regression model using the best predictors, the model was internally validated with 1,000 bootstrap samples to obtain bias-corrected accuracy using a concordance index. The model's predictions were also externally validated by comparing findings against actual outcomes using Colpopexy and Urinary Reduction Efforts trial patients (n=316). The final model's performance was compared with experts using a test data set of 32 randomly chosen Outcomes Following Vaginal Prolapse Repair and Midurethral Sling trial patients through comparison of the model's area under the curve against: 1) 22 experts' predictions; and 2) preoperative prolapse reduction stress testing. A model containing seven predictors discriminated between de novo SUI status (concordance index 0.73, 95% confidence interval [CI] 0.65-0.80) in Outcomes Following Vaginal Prolapse Repair and Midurethral Sling participants and outperformed expert clinicians (area under the curve 0.72 compared with 0.62, P<.001) and preoperative urinary stress testing (area under the curve 0.72 compared with 0.54, P<.001). The concordance index for Colpopexy and Urinary Reduction Efforts trial participants was 0.62 (95% CI 0.56-0.69). This individualized prediction model for de novo SUI after vaginal POP surgery is valid and outperforms preoperative stress testing, prediction by experts, and preoperative reduction cough stress testing. An online calculator is provided for clinical use. III.

  • Research Article
  • 10.1007/s00432-023-05105-2
Invasiveness identification in pure ground-glass nodules: exploring the generalizability of radiomics based on external validation and stress testing.
  • Jul 15, 2023
  • Journal of cancer research and clinical oncology
  • Ziqi Xiong + 6 more

This study aimed to apply external validation and stress tests to evaluate the generalizability of radiomics models built using various machine-learning methods for identifying the invasiveness of lung adenocarcinomas manifesting as pure ground-glass nodules (pGGNs). This retrospective study enrolled 495 patients (514 pGGNs) confirmed as lung adenocarcinomas by postoperative pathology from three centers. All nodules were included in the primary cohort (randomly divided into training and test cohorts), two external validation cohorts, and two stress test cohorts. Six machine-learning radiomics models were constructed in the training cohort using the optimal features. Performance of radiomics models and clinical models were compared in primary cohort and external validation cohorts. The stress tests included stratified performance evaluation and shifted performance evaluation and contrastive evaluation under three single-condition modification settings. The predictive performance was validated by area under curve (AUC) of receiver operating characteristic (ROC). Of the six radiomics models, the best logistic regression (LR) model was able to maintain high differential diagnostic capability (AUC: 0.849 ± 0.049) and good stability (relative standard deviation, 5.719%), but it showed poorer performance (AUC= 0.835) than the clinical model (AUC= 0.862) in the external validation cohort E1. The stress tests suggested LR model had no significant difference in performance between subgroups after stratification and had good consistency in the predictions before and after the three transformations (Kappa = 0.960, 0.840, and 0.933, respectively; p < 0.05, all). The rigorous testing procedure facilitates the selection of high-performance radiomics models with good clinical generalizability.

  • Research Article
  • Cite Count Icon 119
  • 10.1016/j.annemergmed.2004.09.012
Artificial Neural Network Models for Prediction of Acute Coronary Syndromes Using Clinical Data From the Time of Presentation
  • Apr 27, 2005
  • Annals of Emergency Medicine
  • Robert F Harrison + 1 more

Artificial Neural Network Models for Prediction of Acute Coronary Syndromes Using Clinical Data From the Time of Presentation

  • Research Article
  • Cite Count Icon 39
  • 10.1177/001316447403400206
Approximations to Item Parameters of Mental Test Models and Their Uses1
  • Jul 1, 1974
  • Educational and Psychological Measurement
  • Vern W Urry

Equations were derived to enable the graphic approximation of the item parameters of the stochastic mental test models, i.e., the generalized normal ogive and logistic models. The item parameters for the models are discriminatory power ( ai), difficulty ( bi), and lower asymptote of the item characteristic curve ( ci) where the item characteristic curve (ICC) is the regression of the binary item on latent ability. In brief, c i can be approximated through visual inspection of the left-hand (lower) asymptote of the proportion passing the item plotted against the total test score minus the particular item. Thereafter, a graph appropriate to the approximate ci can be consulted to convert an ordinary item-total test point-biserial correlation and proportion passing the item into approximations of item discriminatory power ( ai) and item difficulty ( bi). Suggested uses for the approximations were to provide a basis for screening items for tailored testing, to enable a determination as to the appropriateness of a set of items for tailored testing, and to provide starting values for parameter estimation in maximum likelihood procedures. The conditions and assumptions necessary for an effective application of the method were delineated. Recent empirical results which bear on the properties of the approximations were examined. An investigation was suggested to evaluate a further possible use of the approximations that of their direct applicability in tailored testing procedures. The generated graphs may also be helpful pedagogically in appreciating the relationships between conventional and mental test model parameters.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 1
  • 10.1117/1.jmi.11.2.024013
Simulation of acquisition shifts in T2 weighted fluid-attenuated inversion recovery magnetic resonance images to stress test artificial intelligence segmentation networks
  • Mar 1, 2024
  • Journal of Medical Imaging
  • Christiane Posselt + 6 more

.PurposeTo provide a simulation framework for routine neuroimaging test data, which allows for “stress testing” of deep segmentation networks against acquisition shifts that commonly occur in clinical practice for T2 weighted (T2w) fluid-attenuated inversion recovery magnetic resonance imaging protocols.ApproachThe approach simulates “acquisition shift derivatives” of MR images based on MR signal equations. Experiments comprise the validation of the simulated images by real MR scans and example stress tests on state-of-the-art multiple sclerosis lesion segmentation networks to explore a generic model function to describe the F1 score in dependence of the contrast-affecting sequence parameters echo time (TE) and inversion time (TI).ResultsThe differences between real and simulated images range up to 19% in gray and white matter for extreme parameter settings. For the segmentation networks under test, the F1 score dependency on TE and TI can be well described by quadratic model functions (). The coefficients of the model functions indicate that changes of TE have more influence on the model performance than TI.ConclusionsWe show that these deviations are in the range of values as may be caused by erroneous or individual differences in relaxation times as described by literature. The coefficients of the F1 model function allow for a quantitative comparison of the influences of TE and TI. Limitations arise mainly from tissues with a low baseline signal (like cerebrospinal fluid) and when the protocol contains contrast-affecting measures that cannot be modeled due to missing information in the DICOM header.

  • Research Article
  • Cite Count Icon 53
  • 10.1016/j.jbankfin.2016.02.004
Extreme risk modeling: An EVT–pair-copulas approach for financial stress tests
  • Mar 8, 2016
  • Journal of Banking &amp; Finance
  • Lyes Koliai

Extreme risk modeling: An EVT–pair-copulas approach for financial stress tests

  • Research Article
  • Cite Count Icon 2
  • 10.56038/ejrnd.v4i1.422
Deep Learning Approaches for Stream Flow and Peak Flow Prediction: A Comparative Study
  • Mar 31, 2024
  • The European Journal of Research and Development
  • Levent Latifoğlu + 1 more

Stream flow prediction is crucial for effective water resource management, flood prevention, and environmental planning. This study investigates the performance of various deep neural network architectures, including LSTM, biLSTM, GRU, and biGRU models, in stream flow and peak stream flow predictions. Traditional methods for stream flow forecasting have relied on hydrological models and statistical techniques, but recent advancements in machine learning and deep learning have shown promising results in improving prediction accuracy. The study compares the performance of the models using comprehensive evaluations with 1-6 input steps for both general stream flow and peak stream flow predictions. Additionally, a detailed analysis is conducted specifically for the biLSTM model, which demonstrated high performance results. The biLSTM model is evaluated for 1-4 ahead forecasting, providing insights into its specific strengths and capabilities in capturing the dynamics of stream flow. Results show that the biLSTM model outperforms other models in terms of prediction accuracy, especially for peak stream flow forecasting. Scatter plots illustrating the forecasting performances of the models further demonstrate the effectiveness of the biLSTM model in capturing temporal dependencies and nonlinear patterns in stream flow data. This study contributes to the literature by evaluating and comparing the performance of deep neural network models for general and peak stream flow prediction, highlighting the effectiveness of the biLSTM model in improving the accuracy and reliability of stream flow forecasts.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 1
  • 10.3390/en13133325
Application of IRT Models to Selection of Bidding Paths in Financial Transmission Rights Auction: U.S. New England
  • Jun 30, 2020
  • Energies
  • Peter Jang + 2 more

This paper explores a way to apply Item Response Theory (IRT), one of the popular statistical methodologies in measurement and psychometrics, to evaluate Financial Transmission Rights (FTR) paths in the U.S. electricity market. FTR is an energy derivative product to hedge congestion cost risks inherent in constrained transmission lines. In New England, with about 1200 pricing locations, the theoretical combinations of FTR paths amount to 1.4 million in prevailing flows alone. With capital constraints, it is imperative that FTR market participants build the capability to evaluate FTR paths to bid on. IRT provides a framework of how well tests work, and how individual items work on tests, estimating respondents’ latent abilities, and individual item parameters. IRT is utilized to analyze historical electricity data of 2019 for a daily congestion cost of eight customer load zones and one hub in the U.S., New England, for the evaluation of FTR paths. In the analysis, an item represents an FTR path, while item difficulty, item discrimination, and a latent trait variable for the path correspond to the path profitability, risk level, and daily congestion ability, respectively. This paper explores the experimental procedures by which IRT, a psychometric tool, may also be applicable in complex energy markets, providing a consistent and standardized analytical framework to address the issues of selection and prioritization among multiple opportunities. FTR path evaluation is conducted in three steps to determine bid priority paths in FTR auctions: parameter significance tests, ranking on path profitability and risk level, and weighting scores of individual rankings on the two criteria.

  • Research Article
  • Cite Count Icon 2
  • 10.1080/15366367.2020.1742054
Estimation of Mixture Rasch Models from Skewed Latent Ability Distributions
  • Oct 1, 2020
  • Measurement: Interdisciplinary Research and Perspectives
  • Tugba Karadavut + 2 more

Mixture Rasch (MixRasch) models conventionally assume normal distributions for latent ability. Previous research has shown that the assumption of normality is often unmet in educational and psychological measurement. When normality is assumed, asymmetry in the actual latent ability distribution has been shown to result in extraction of spurious latent classes in MixRasch models. In this study, the assumption of a skew-t distribution for the latent ability was examined for its effect on reducing extraction of spurious latent classes. A simulation study was conducted with eight different latent ability distributions with varying levels of skewness and kurtosis, two sample sizes (600 and 2,000), and two test lengths (10-item and 30-item). Results showed that the 30-item test but not the 10-item test were robust to extraction of spurious latent classes independent of the sample size or the shape of the ability distribution. Use of a skew-t prior, on the other hand, reduced spurious latent class extraction for the 10-item test, particularly for the sample size of 2,000 for high levels of skewness compared to use of a normal prior. Thus, a skew-t prior was useful for reducing spurious latent class extraction for short tests, although it did not improve item parameter estimation compared to a normal prior.

  • Research Article
  • Cite Count Icon 5
  • 10.1186/s13054-025-05591-5
Evaluating the o1 reasoning large language model for cognitive bias: a vignette study
  • Aug 21, 2025
  • Critical Care
  • Or Degany + 3 more

BackgroundCognitive biases, systematic deviations from logical judgment, are well documented in clinical decision-making, particularly in clinical settings characterized by high decision load, limited time, and diagnostic uncertainty-such as critical care. Prior work demonstrated that large language models, particularly GPT-4, reproduce many of these biases, sometimes to a greater extent than human clinicians.MethodsWe tested whether the o1 model (o1-2024–12-17), a newly released AI system with enhanced reasoning capabilities, is susceptible to cognitive biases that commonly affect medical decision-making. Following the methodology established by Wang and Redelmeier [15], we used ten pairs of clinical scenarios, each designed to test a specific cognitive bias known to influence clinicians. Each scenario had two versions, differed by subtle modifications designed to trigger the bias (such as presenting mortality rates versus survival rates). The o1 model generated 90 independent clinical recommendations for each scenario version, totalling 1,800 responses. We measured cognitive bias as systematic differences in recommendation rates between the paired scenarios, which should not occur with unbiased reasoning. The o1 model's performance was compared against previously published results from both the GPT-4 model and historical human clinician studies.ResultsThe o1 model showed no measurable cognitive bias in seven of the ten vignettes. In two vignettes, the o1 model showed significant bias, but its absolute magnitude was lower than values previously reported for GPT-4 and human clinicians. In a single vignette, Occam’s razor, the o1 model exhibited consistent bias. Therefore, although overall bias appears less frequent overall with the reasoning model than with GPT-4, it was worse in one vignette. The model was more prone to bias in vignettes that included a gap-closing cue, seemingly resolving the clinical uncertainty. Across eight vignette versions, intra‑scenario agreement exceeded 94%, indicating lower decision variability than previously described with GPT‑4 and human clinicians.ConclusionReasoning models may reduce cognitive bias and random variation in judgment (i.e., “noise”). However, our findings caution that reasoning models are still not entirely immune to cognitive bias. These findings suggest that reasoning models may impart some benefits as decision-support tools in medicine, but they also imply a need to explore further the circumstances in which these tools may fail.Graphical Supplementary InformationThe online version contains supplementary material available at 10.1186/s13054-025-05591-5.

  • Research Article
  • Cite Count Icon 14
  • 10.1002/j.2333-8504.2011.tb02276.x
THE SENSITIVITY OF PARAMETER ESTIMATES TO THE LATENT ABILITY DISTRIBUTION
  • Dec 1, 2011
  • ETS Research Report Series
  • Xueli Xu + 1 more

ABSTRACTEstimation of item response model parameters and ability distribution parameters has been, and will remain, an important topic in the educational testing field. Much research has been dedicated to addressing this task. Some studies have focused on item parameter estimation when the latent ability was assumed to follow a normal distribution, whereas others have utilized nonparametric or semiparametric techniques to substitute the normal ability assumption. However, both approaches have their limitations. A normal ability assumption is not flexible enough to reflect possible deviations from symmetry, whereas the nonparametric and semiparametric techniques used to capture possible nonnormal features of the latent ability have difficulty in reaching satisfactory estimates for certain quantities of the ability distribution such as quantiles. Hence a continuous generalized skew normal (GSN) distribution was applied in this study to better capture the possible underlying asymmetric ability distribution. In addition, simultaneous estimation of both the item parameters and the distributional parameters was employed. The performance of the GSN was compared with the normal ability assumption in terms of item parameter and distributional parameter recoveries, based on a series of simulation studies. The results showed that (a) under the Rasch model, both the item parameter estimates and the distributional parameter estimates are robust to the misspecification of the ability distribution, and (b) under the two‐parameter logistic model, although the distributional parameter estimates are fairly robust, the item parameter estimates are slightly more sensitive to the misspecification of the ability distribution, especially when the underlying ability distribution is highly skewed.

  • Research Article
  • Cite Count Icon 1229
  • 10.1007/bf02291411
Estimating Item Parameters and Latent Ability when Responses are Scored in Two or More Nominal Categories
  • Mar 1, 1972
  • Psychometrika
  • R Darrell Bock

A multivariate logistic latent trait model for items scored in two or more nominal categories is proposed. Statistical methods based on the model provide 1) estimation of two item parameters for each response alternative of each multiple choice item and 2) recovery of information from “wrong” responses when estimating latent ability. An application to a large sample of data for twenty vocabulary items shows excellent fit of the model according to a chi-square criterion. Item and test information curves are compared for estimation of ability assuming multiple category and dichotomous scoring of these items. Multiple scoring proves substantially more precise for subjects of less than median ability, and about equally precise for subjects above the median.

  • Research Article
  • Cite Count Icon 1
  • 10.1002/ets2.12325
Robustness of Weighted Differential Item Functioning (DIF) Analysis: The Case of Mantel–Haenszel DIF Statistics
  • Aug 8, 2021
  • ETS Research Report Series
  • Ru Lu + 2 more

Two families of analysis methods can be used for differential item functioning (DIF) analysis. One family is DIF analysis based on observed scores, such as the Mantel–Haenszel (MH) and the standardized proportion‐correct metric for DIF procedures; the other is analysis based on latent ability, in which the statistic is a measure of departure from measurement invariance (DMI) for two studied groups. Previous research has shown, that DIF and DMI do not necessarily agree with each other. In practice, many operational testing programs use the MH DIF procedure to flag potential DIF items. Recently, weighted DIF statistics has been proposed, where weighted sum scores are used as the matching variable and the weights are the item discrimination parameters. It has been shown theoretically and analytically that, given the item parameters, weighted DIF statistics can close the gap between DIF and DMI. The current study investigates the robustness of using weighted DIF statistics empirically through simulations when item parameters have to be estimated from data.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant