Signposts on the Path From Nominal to Ordinal Scales: Moving From a Discrete to a Continuous View.
This study redefines the nominal-ordinal classification of polytomous response data as a continuum, introducing six indices to measure category ordering. Empirical evaluation shows two parametric indices are robust and informative, enhancing measurement accuracy and model selection across diverse datasets.
Polytomous item response data are typically classified as either nominal or ordinal, but this binary distinction may oversimplify their true structure. In this paper, we reframe the nominal-ordinal distinction as a continuum and introduce six empirical indices to quantify the degree of category ordering in item response data. Through extensive simulations with various item response theory (IRT) models and applications to 245 empirical datasets, we evaluate the indices' sensitivity, computational efficiency, and interpretability across diverse measurement contexts. Our findings show that two parametric indices-Mean Difference between Slope Parameters (Index 5) and Arctangent of Paired Category Ratios (Index 6)-are particularly robust and informative, even with low-frequency categories. These indices offer a practical tool for assessing whether and how item categories align with ordinal assumptions, supporting more accurate measurement and model selection. We conclude that treating ordering as a continuum, rather than a binary property, provides deeper insights for psychometric practice and strengthens the connection between empirical response patterns and their theoretical representations.
- Dissertation
- 10.32469/10355/91690
- May 1, 2022
[EMBARGOED UNTIL 6/1/2023] Traditional item response theory (IRT) models assume a symmetric error distribution and rely on symmetric (logit or probit) link functions to model the response probabilities. However, this assumption does not always hold in item response data. To explore the possible benefits of alternate link functions, I developed and investigated a set of one-parameter IRT models for unidimensional tests with dichotomous items by specifying the complementary log-log (CLL), negative log-log model (NLL), and cauchit links. In a series of simulations studies, I demonstrated that these parametrically parsimonious alternate-link IRT models are comparable to traditional IRT models with regard to (1) data distribution shape, (b) inflection point shift, (c) structural similarities, and (d) response behaviors. Importantly, these four properties provide a rationale for how certain parametrically simple alternate-link models can outperform more parametrically complex traditional IRT models. Specifically, I demonstrate that the CLL model accounts for the effect of guessing because it has an item characteristic curve with a higher inflection point than that of the traditional Rasch or two-parameter logistic (2PL) models. Similarly, the NLL model addresses the effects of slipping via an inflection point that is lower than in the Rasch or 2PL models. Importantly, unlike traditional IRT models that focus on guessing and slipping as response behaviors associated with the extreme levels of the latent trait, I show that the CLL and NLL models assume such behaviors affect responses across the entire range of the trait. I also present the cauchit model, which treats both guessing and slipping effects as outliers. The simulation results reveal that these one-parameter alternate-IRT models are robust to small sample sizes (e.g., N = 100), and facilitate item-weighted scoring. I then provide further evidence for these claims by applying the CLL, NLL, and cauchit models to empirical data. Finally, I propose some extensions of alternate-link IRT modeling beyond the particular (unidimensional dichotomous) context that formed the basis of my simulations and empirical analysis. I conclude by discussing the implications of this work for applied measurement and psychometric research.
- Book Chapter
10
- 10.1007/978-1-4613-0169-1_22
- Jan 1, 2001
(1964) Theory of Data is a classification system of behavioral data. In his system, item response data are individual-comparison or individual-stimulus differences data. In this contribution, item response data are further classified using the following facets: Intended Behavior Type, Task Type, Intermediating Variable Type, Construct Level Type, Construct Dimension Type, Recording Type, Scaling Type, Person × Stimulus Interaction Type, Construct Relation Type, and Step Task Type. These facets cannot be completely crossed because some of the facets are nested within certain (combinations of) cells of crossed facets. The facets constitute a theory of item responses (TIR). The TIR and item response theory (IRT) complete each other: the TIR structures the item response data, while IRT explains these data. The TIR can be used to facilitate the choice of an appropriate IRT model.
- Front Matter
34
- 10.1016/s1551-7144(09)00212-2
- Jan 1, 2010
- Contemporary Clinical Trials
Classical and modern measurement theories, patient reports, and clinical outcomes
- Research Article
20
- 10.3102/1076998607306451
- Dec 1, 2008
- Journal of Educational and Behavioral Statistics
The randomized response technique ensures that individual item responses, denoted as true item responses, are randomized before observing them and so-called randomized item responses are observed. A relationship is specified between randomized item response data and true item response data. True item response data are modeled with a (non)linear mixed effects and/or item response theory model. Although the individual true item responses are masked through randomizing the responses, the model extension enables the computation of individual true item response probabilities and estimates of individuals’ sensitive behavior/attitude and their relationships with background variables taking into account any clustering of respondents. Results are presented from a College Alcohol Problem Scale (CAPS) where students were interviewed via direct questioning or via a randomized response technique. A Markov Chain Monte Carlo algorithm is given for estimating simultaneously all model parameters given hierarchical structured binary or polytomous randomized item response data and background variables.
- Dissertation
- 10.17077/etd.9yfctk57
- Feb 5, 2015
<p>The use of testlets in a test can cause multidimensionality and local item dependence (LID), which can result in inaccurate estimation of item parameters, and in turn compromise the quality of item response theory (IRT) true and observed score equating of testlet-based tests. Both unidimensional and multidimensional IRT models have been developed to control local item dependence caused by testlets. The purposes of the current study were to (1) investigate how different levels of LID can affect IRT true and observed score equating of testlet-based tests when the traditional three parameter logistic (3PL) IRT model was used for calibration, and (2) compare the performance of four different IRT models, including the 3PL IRT model, graded response model (GRM), testlet response theory model (TRT), and bifactor model, in IRT true and observed score equating of testlet-based tests with various levels of local item dependence.</p><p>Both real and simulated data analyses were conducted in this study. Two testlet-based tests (i.e., Test A and Test B) that differed in subjects, test length, and testlet length were used in the real data analysis. For simulated data analysis, two main factors were investigated in this study: (1) testlet length (5 or 10), and (2) LID level within testlets that was defined by testlet effect variance (0, 0.25, 0.5625, 0.75, 1, and 1.5). For the unidimensional IRT models (i.e., 3PL IRT model and GRM), unidimensional IRT true score and observed score equating procedures, explained in Kolen and Brennan (2004), were used. For the two investigated multidimensional IRT models (i.e., 3PL TRT model and bifactor model), the unidimensional approximation of multidimensional item response theory (MIRT) true score equating procedure and the unidimensional approximation of MIRT observed score equating procedure (Brossman & Lee, 2013) were applied. The traditional equipercentile equating method was used as the baseline for comparison in both real data and simulated data analyses.</p><p>It was found in the study that both testlet length and the LID level affected the performance of the investigated models on IRT true and observed score equating of testlet-based tests. When the traditional 3PL IRT model was used for tests with long testlets, higher levels of local item dependence led to IRT equating results that deviated further away from those obtained from the baseline method. However, the effect of local item dependence on IRT equating results was not prominent for tests with short testlets.</p><p>Moreover, for tests consisting of long testlets (e.g., a testlet length of 10 or more) and having a very low level of local item dependence (e.g., a LID level of 0.25 or lower), and for tests consisting of short testlets (e.g., a testlet length around 5), all four investigated IRT models worked well in IRT true and observed score equating. For tests with long testlets and a relatively high level of local item dependence (e.g., a LID level of 0.5625 or higher), the GRM, bifactor, and TRT models outperformed the traditional 3PL IRT model in IRT true and observed equating of testlet-based tests.</p><p>The study suggested that the selection of models for IRT true and observed score equating of testlet-based tests should be considered with respect to the features of the testlet-based tests and the groups of examinees from which the data is collected. It is hoped that this study encourages researchers to identify differences among existing models for IRT true and observed score equating of testlet-based tests with various features, and to develop new models that are appropriate for modeling testlet-based tests to obtain accurate IRT number correct score equating results.</p>
- Research Article
46
- 10.1177/00131644211045351
- Sep 13, 2021
- Educational and Psychological Measurement
Disengaged item responses pose a threat to the validity of the results provided by large-scale assessments. Several procedures for identifying disengaged responses on the basis of observed response times have been suggested, and item response theory (IRT) models for response engagement have been proposed. We outline that response time-based procedures for classifying response engagement and IRT models for response engagement are based on common ideas, and we propose the distinction between independent and dependent latent class IRT models. In all IRT models considered, response engagement is represented by an item-level latent class variable, but the models assume that response times either reflect or predict engagement. We summarize existing IRT models that belong to each group and extend them to increase their flexibility. Furthermore, we propose a flexible multilevel mixture IRT framework in which all IRT models can be estimated by means of marginal maximum likelihood. The framework is based on the widespread Mplus software, thereby making the procedure accessible to a broad audience. The procedures are illustrated on the basis of publicly available large-scale data. Our results show that the different IRT models for response engagement provided slightly different adjustments of item parameters of individuals’ proficiency estimates relative to a conventional IRT model.
- Research Article
75
- 10.1080/10705511.2011.581993
- Jun 30, 2011
- Structural Equation Modeling: A Multidisciplinary Journal
Linear factor analysis (FA) models can be reliably tested using test statistics based on residual covariances. We show that the same statistics can be used to reliably test the fit of item response theory (IRT) models for ordinal data (under some conditions). Hence, the fit of an FA model and of an IRT model to the same data set can now be compared. When applied to a binary data set, our experience suggests that IRT and FA models yield similar fits. However, when the data are polytomous ordinal, IRT models yield a better fit because they involve a higher number of parameters. But when fit is assessed using the root mean square error of approximation (RMSEA), similar fits are obtained again. We explain why. These test statistics have little power to distinguish between FA and IRT models; they are unable to detect that linear FA is misspecified when applied to ordinal data generated under an IRT model.
- Research Article
8
- 10.1177/0146621616679394
- Nov 28, 2016
- Applied Psychological Measurement
In psychometric practice, the parameter estimates of a standard item-response theory (IRT) model can become biased when item-response data, of persons' individual responses to test items, contain outliers relative to the model. Also, the manual removal of outliers can be a time-consuming and difficult task. Besides, removing outliers leads to data information loss in parameter estimation. To address these concerns, a Bayesian IRT model that includes person and latent item-response outlier parameters, in addition to person ability and item parameters, is proposed and illustrated, and is defined by item characteristic curves (ICCs) that are each specified by a robust, Student's t-distribution function. The outlier parameters and the robust ICCs enable the model to automatically identify item-response outliers, and to make estimates of the person ability and item parameters more robust to outliers. Hence, under this IRT model, it is unnecessary to remove outliers from the data analysis. Our IRT model is illustrated through the analysis of two data sets, involving dichotomous- and polytomous-response items, respectively.
- Book Chapter
- 10.4324/9781315871493-17
- Aug 20, 2015
This chapter seeks to highlight some of the unique types of item response theory (IRT) models that have emerged in support of computer-based testing (CBT). The multicategory scoring of many item types used in CBT are statistically improving measurement innovative item types has made polytomous IRT models of significant value in CBT. The polytomous IRT models are useful in evaluating the extent to which the innovative ciency. It is important to acknowledge other variants of testlet-based administration that can impact IRT modelling. The IRT models of increased relevance in CBT consist of multidimensional IRT models. In MIRT, item scores are modelled as a function of multiple person abilities. The diversity of IRT and IRT-related models needed for CBT has led to new thinking about how IRT models function within a broader assessment framework. The computer is offering much exciting future work within field of psychometrics for those who like to think creatively about the use of models in assessment contexts.
- Research Article
18
- 10.1177/0146621612440305
- Apr 25, 2012
- Applied Psychological Measurement
When tests consist of multiple-choice and constructed-response items, researchers are confronted with the question of which item response theory (IRT) model combination will appropriately represent the data collected from these mixed-format tests. This simulation study examined the performance of six model selection criteria, including the likelihood ratio test, Akaike’s information criterion (AIC), corrected AIC, Bayesian information criterion, Hannon and Quinn’s information criterion, and consistent AIC, with respect to correct model selection among a set of three competing mixed-format IRT models (i.e., one-parameter logistic/partial credit [1PL/PC], two-parameter logistic/generalized partial credit [2PL/GPC], and three-parameter logistic/generalized partial credit [3PL/GPC]). The criteria were able to correctly select less parameterized IRT models, including the PC, 1PL, and 1PL/PC models. In contrast, the criteria were less able to correctly select more parameterized IRT models, including the GPC, 3PL, and 3PL/GPC models. Implications of the findings and recommendations are discussed.
- Research Article
10
- 10.1080/00273171.2016.1178567
- Jun 20, 2016
- Multivariate Behavioral Research
ABSTRACTWhen categorical ordinal item response data are collected over multiple timepoints from a repeated measures design, an item response theory (IRT) modeling approach whose unit of analysis is an item response is suitable. This study proposes a few longitudinal IRT models and illustrates how a popular compensatory multidimensional IRT model can be utilized to formulate such longitudinal IRT models, which permits an investigation of ability growth at both individual and population levels. The equivalence of an existing multidimensional IRT model and those longitudinal IRT models is also elaborated so that one can make use of an existing multidimensional IRT model to implement the longitudinal IRT models.
- Research Article
13
- 10.1093/swr/34.2.94
- Jun 1, 2010
- Social Work Research
The need to develop measures that tap into constructs of interest to social work, refine existing measures, and ensure that measures function adequately across diverse populations of interest is critical. Item response theory (IRT) is a modern measurement approach that is increasingly seen as an essential tool in a number of allied professions. IRT-based measurement uses a model-based approach that has several analytical and explanatory advantages over classical test theory. In particular, IRT-based techniques facilitate the process of specific item selection, allow for increased measurement precision with fewer items, and provide greater capacity for understanding and accounting for measurement bias across diverse populations. A survey of the top (as rated by impact factor) 20 social work journals revealed that few measurement articles in the social work literature use IRT or other modern measurement approaches. The benefit of incorporating more IRT-based approaches for developing, refining, and ensuring the application of measures to diverse populations is discussed. KEY WORDS: bias; classical test theory; item response theory; measurement; social work ********** The state of measurement within the social work literature is integrally related to knowledge base development and, ultimately, the extent to which research is able to meaningfully inform practice (Holden, Nizza, & Weissman, 1995). Scholarship highlights at least three measurement-related research domains within the field of social work. The first concerns the development of valid and reliable measures that capture the diverse set of phenomena relevant to social work, particularly those phenomena that may not be adequately represented by existent standardized instruments. The second is the assessment and validation of such measures. In particular, high-quality intervention research hinges on the validity and reliability of measures used to assess outcomes (Rosen, Proctor, & Staudt, 1999).Third, a growing body of literature challenges the extent to which well-validated measures adequately account and adjust for within- and across-population sources of diversity (see Ramirez, Ford, Stewart, & Teresi, 2005; Snowden, 2003), and such concerns are highly salient to social work's commitment to diversity-sensitive and -responsive research and practice. During the 1980s and 1990s, social work researchers outlined the relative benefits of item response theory (IRT) over classical test theory (CTT) measurement models, calling explicitly for IRT-based models' increased utilization to address measurement problems in social work research (DeRoos & Allen-Meares, 1993, 1998; Nugent & Hankins, 1989,1992). Indeed, IRT models have largely subsumed CTT approaches within a wide range of allied fields and disciplines (for example, medicine, psychology, nursing, public health, education) (see Dunn, Resnicow, & Klesges, 2006; Embretson & Reise, 2000; Fries, Bruce, & Cella, 2005; Lord, 1980; Ware, Bjorner, & Kosinski, 2000). Given early interest among social work researchers and the recent proliferation of IRT methods within other applied social sciences, our overall objective in the present study was to assess the extent to which these methods are represented within social work research. This review thus realizes three overlapping aims. First, it provides a description and comparison of IRT and CTT models and outlines the potential contributions of IRT methods to social work scholarship; it also briefly discusses IRT more generally as a latent variable model and its overlap with confirmatory factor analytic (CFA) and multi-level modeling methods. Second, it presents the results of a structured review assessing the penetration of IRT-based methods into the field of social work as reflected in key social work research journals. Third, using these results as a launching point, we highlight particular lines of inquiry within social work research where the application of IRT methods would likely yield substantial innovation. …
- Research Article
3
- 10.5750/ijpcm.v6i4.614
- Feb 2, 2017
- International Journal of Person Centered Medicine
Background: More robust and rigorous psychometric models, such as Item Response Theory (IRT) models, have been advocated for applications measuring health sciences outcomes. However, there are challenges to the use of IRT models with health assessments. In particular, item responses from measuring health-related outcomes are typically determined by multiple traits or dimensions. This multidimensionality can be caused by various factors including designed multidimensional structure to the instrument, heterogeneity in item content, and from other sources such as differential item functioning in subpopulations and individual differences in response styles to survey items and rating scales. Objectives: This paper discusses different extensions to IRT models that can be used to account for different types of multidimensionality as well as the use of Bayesian methods with person-centered medicine research.Methods: Use of the SAS PROC MCMC platform for implementing Bayesian analyses is illustrated to estimate and analyze IRT applications to health-related assessments. Results: PROC MCMC involves a straightforward translation of the response probability model along with specifications of the model parameters and prior distributions for the model parameters. Conclusions: Bayesian analysis of multidimensional IRT models is more accessible to researchers and scale developers in measuring health sciences outcomes for person-centered medicine research.
- Research Article
28
- 10.3389/fpsyg.2017.00484
- Apr 4, 2017
- Frontiers in Psychology
In item response theory (IRT) models, assessing model-data fit is an essential step in IRT calibration. While no general agreement has ever been reached on the best methods or approaches to use for detecting misfit, perhaps the more important comment based upon the research findings is that rarely does the research evaluate IRT misfit by focusing on the practical consequences of misfit. The study investigated the practical consequences of IRT model misfit in examining the equating performance and the classification of examinees into performance categories in a simulation study that mimics a typical large-scale statewide assessment program with mixed-format test data. The simulation study was implemented by varying three factors, including choice of IRT model, amount of growth/change of examinees’ abilities between two adjacent administration years, and choice of IRT scaling methods. Findings indicated that the extent of significant consequences of model misfit varied over the choice of model and IRT scaling methods. In comparison with mean/sigma (MS) and Stocking and Lord characteristic curve (SL) methods, separate calibration with linking and fixed common item parameter (FCIP) procedure was more sensitive to model misfit and more robust against various amounts of ability shifts between two adjacent administrations regardless of model fit. SL was generally the least sensitive to model misfit in recovering equating conversion and MS was the least robust against ability shifts in recovering the equating conversion when a substantial degree of misfit was present. The key messages from the study are that practical ways are available to study model fit, and, model fit or misfit can have consequences that should be considered when choosing an IRT model. Not only does the study address the consequences of IRT model misfit, but also it is our hope to help researchers and practitioners find practical ways to study model fit and to investigate the validity of particular IRT models for achieving a specified purpose, to assure that the successful use of the IRT models are realized, and to improve the applications of IRT models with educational and psychological test data.
- Dissertation
- 10.17077/etd.005181
- Dec 1, 2019
Nowadays it is not uncommon that tests, especially high-stakes assessments, are administered with time constraints. When a test is constructed to assess examinees’ abilities in academic knowledge, but the imposed time limits affect examinees’ test performance, speededness effects become a concern. Under such circumstances, inaccurate psychometric results and inferences might be drawn if unidimensional item response theory (IRT) models are applied in testing practice. Speededness detection methods were proposed to identify speeded responses/examinees. Thus, the purpose of the study was to comprehensively investigate how the performance of various detection methods combined with various calibration treatments compared in reducing speededness effects under the 2PL and 3PL IRT models with: (1) simulated test data under various speededness conditions, and (2) real test data. Both simulated and real data analyses were conducted in this study. Two simulation studies were conducted. For the first simulation study, two main factors were investigated: (1) degree of speededness (three levels: None, 10%, and 25%), and (2) IRT calibration model (two models: 2PL, and 3PL). The performance of various combinations of detection methods and calibration treatments were evaluated by assessing Pearson correlation, item parameter recovery, and model-data fit statistics. Data generated in the second simulation study were based on the estimated person and item parameter values obtained from IRT model calibration of the real data used in this study. Thus, the second simulation study served as a link between the pure simulation study and the real data study, because such a generation process enabled the simulated dataset to carry some characteristics of the real data, while true parameter values were known. The real data came from a large pool of a high-stakes standardized assessment items. In the current study, it was found that treating the identified speeded responses as “not-presented” could always lead to more accurate psychometric results compared to the other calibration treatments across various speededness levels under both the 2PL and 3PL IRT models. When the speededness level was large, “removing speeded examinees” could usually yield comparable results compared to “not-presented” treatments across different detection methods, and is a feasible and easily manipulated option in practice. In addition, it was found that detection methods using the item response time (RT) distribution as a speededness indicator (i.e., the INSPECT and VITP methods in the current study) generally showed better performance than the other detection methods in dealing with speededness effects. Moreover, in this study, it was found that the inclusion of the c-parameter could deal with rapid guessing strategy well. Thus, when the speededness level was not large, and mainly caused by rapid guessing behavior, “no treatment” under the 3PL IRT model yielded accurate psychometric results. The findings of the current study provide several feasible options for practitioners when speededness is a concern and unidimentional IRT models are used in the calibration or scoring process. It is hoped that this study will inspire researchers and practitioners to develop new detection methods, or ways of dealing with speededness effects under unidimensional IRT models.