Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Controlling the False Discovery Rate in DIF Detection With e-Values: Evidence From Multidimensional and Testlet Simulations.

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

This study presents the first application of e-value-based false discovery rate (FDR) control to Differential Item Functioning (DIF) detection, addressing long-standing limitations of p-value-based approaches when model assumptions are violated-for example, under multidimensionality, local item dependence, or extreme sample sizes. Two comprehensive simulation studies were conducted to evaluate e-BH (the e-value analogue of BH) procedures, using K-fold and Multisplit likelihood-ratio e-values, under (a) multidimensional contamination and (b) testlet-based local dependence. Across both scenarios, e-BH consistently provided stronger and more stable control of Type I error, FDR, and family-wise error rate (FWER) than classical procedures such as Benjamini-Hochberg (BH) and Holm. Even under severe model misspecification, e-BH maintained substantially lower false-positive rates while remaining relatively competitive in terms of Type II error. A key finding concerns sample size: classical p-value methods exhibited inflation of Type I error as N increased, whereas e-BH preserved stable error control due to its model-agnostic calibration. An empirical application using Progress in International Reading Literacy Study (PIRLS) data further demonstrated that e-BH produces a more defensible and operationally sustainable set of DIF flags than traditional approaches. Together, these results establish e-values as a powerful and robust evidential tool for DIF detection in modern assessment contexts.

Similar Papers
  • Research Article
  • Cite Count Icon 2
  • 10.21831/reid.v10i1.65284
A psychometric evaluation of an item bank for an English reading comprehension tool using Rasch analysis
  • Jun 30, 2024
  • REID (Research and Evaluation in Education)
  • Louis Wai Keung Yim + 2 more

This study reports the psychometric evaluation of an item bank for an Assessment for Learning (AfL) tool to assess primary school students’ reading comprehension skills. A pool of 46 primary 1 to 6 reading passages and their accompanying 522 multiple choice and short answer items were developed based on the Progress in International Reading Literacy Study (PIRLS) assessment framework. They were field-tested at 27 schools in Singapore involving 9834 students aged between 7 and 13. Four main comprehension processes outlined in PIRLS were assessed: focusing on and retrieving explicitly stated information, making straightforward inferences, interpreting and integrating ideas and information, and evaluating and critiquing content and textual elements. Rasch analysis was employed to examine students’ item response patterns for (1) model and item fit; (2) differential item functioning (DIF) about gender and test platform used; (3) local item dependence (LID) within and amongst reading passages; and (4) distractor issues about options within the multiple-choice-type items. Results showed that the data adequately fit the unidimensional Rasch model across all test levels with good internal consistency. Psychometric issues found amongst items were primarily related to ill-functioning distractors and local dependence on items. Problematic items identified were reviewed and subsequently amended by a panel of assessment professionals for future recalibration. This psychometrically and theoretically sound item bank is envisaged to be valuable to developing comprehensive classroom AfL tools that provide information for the English reading comprehension instructional design in the Singaporean context.

  • Research Article
  • Cite Count Icon 6
  • 10.1080/15305058.2015.1039644
Recursive Partitioning to Identify Potential Causes of Differential Item Functioning in Cross-National Data
  • Aug 28, 2015
  • International Journal of Testing
  • W Holmes Finch + 2 more

Differential item functioning (DIF) assessment is key in score validation. When DIF is present scores may not accurately reflect the construct of interest for some groups of examinees, leading to incorrect conclusions from the scores. Given rising immigration, and the increased reliance of educational policymakers on cross-national assessments such as Programme for International Student Assessment, Trends in International Mathematics and Science Study, and Progress in International Reading Literacy Study (PIRLS), DIF with regard to native language is of particular interest in this context. However, given differences in language and cultures, assuming similar cross-national DIF may lead to mistaken assumptions about the impact of immigration status, and native language on test performance. The purpose of this study was to use model-based recursive partitioning (MBRP) to investigate uniform DIF in PIRLS items across European nations. Results demonstrated that DIF based on mother's language was present for several items on a PIRLS assessment, but that the patterns of DIF were not the same across all nations.

  • Research Article
  • Cite Count Icon 8
  • 10.1007/s12564-009-9039-7
Examining type I error and power for detection of differential item and testlet functioning
  • Jun 10, 2009
  • Asia Pacific Education Review
  • Young-Sun Lee + 2 more

In this study, the effectiveness of detection of differential item functioning (DIF) and testlet DIF using SIBTEST and Poly-SIBTEST were examined in tests composed of testlets. An example using data from a reading comprehension test showed that results from SIBTEST and Poly-SIBTEST were not completely consistent in the detection of DIF and testlet DIF. Results from a simulation study indicated that SIBTEST appeared to maintain type I error control for most conditions, except in some instances in which the magnitude of simulated DIF tended to increase. This same pattern was present for the Poly-SIBTEST results, although Poly-SIBTEST demonstrated markedly less control of type I errors. Type I error control with Poly-SIBTEST was lower for those conditions for which the ability was unmatched to test difficulty. The power results for SIBTEST were not adversely affected, when the size and percent of simulated DIF increased. Although Poly-SIBTEST failed to control type I errors in over 85% of the conditions simulated, in those conditions for which type I error control was maintained, Poly-SIBTEST demonstrated higher power than SIBTEST.

  • Research Article
  • Cite Count Icon 2
  • 10.17159/2520-9868/i87a07
Investigating the differential item functioning of a PIRLS Literacy 2016 text across three languages
  • Jul 25, 2022
  • Journal of Education
  • Karen Roux + 2 more

This study forms part of a larger study (Roux, 2020), which looked at the equivalence of a literary text across English, Afrikaans, and isiZulu from the Progress in International Reading Literacy Study (PIRLS). PIRLS is a large-scale reading comprehension assessment that assesses Grade 4 students' reading literacy achievement. PIRLS Literacy 2016 results for South African Grade 4 students indicated poor performance in reading comprehension, with approximately eight out of 10 Grade 4 students who could not read for meaning. Descriptive statistics led to the Rasch analysis, which was conducted using the South African PIRLS Literacy 2016 data. Even though the Rasch analysis indicated differential item functioning across the three languages for this specific passage, there was no universal discrimination against one particular language. By conducting differential item functioning, it was possible to determine whether the selected text had metric equivalence, in other words, whether the test questions were of similar difficulty across languages.

  • Research Article
  • Cite Count Icon 11
  • 10.1080/13803611.2011.630560
How do different versions of a test instrument function in a single language? A DIF analysis of the PIRLS 2006 German assessments
  • Nov 30, 2011
  • Educational Research and Evaluation
  • Tobias C Stubbe

The challenge inherent in cross-national research of providing instruments in different languages measuring the same construct is well known. But even instruments in a single language may be biased towards certain countries or regions due to local linguistic specificities. Consequently, it may be appropriate to use different versions of an instrument in a single language to allow for regional differences. Using data from the Progress in International Reading Literacy Study (PIRLS) 2006, this article examines the consequences of differing German translations of the reading assessment on item difficulty in Austria, Germany, Luxembourg, and in the German-speaking Community of Belgium. A differential item functioning (DIF) analysis indicates that a substantial number of items have significantly different item difficulties. Especially items with differing translations function differently. A closer look at these items shows that there are plausible reasons why one version is easier and another version more difficult.

  • Research Article
  • Cite Count Icon 3
  • 10.1177/00131644211028995
DIF Detection With Zero-Inflation Under the Factor Mixture Modeling Framework.
  • Jul 26, 2021
  • Educational and psychological measurement
  • Sooyong Lee + 2 more

Response data containing an excessive number of zeros are referred to as zero-inflated data. When differential item functioning (DIF) detection is of interest, zero-inflation can attenuate DIF effects in the total sample and lead to underdetection of DIF items. The current study presents a DIF detection procedure for response data with excess zeros due to the existence of unobserved heterogeneous subgroups. The suggested procedure utilizes the factor mixture modeling (FMM) with MIMIC (multiple-indicator multiple-cause) to address the compromised DIF detection power via the estimation of latent classes. A Monte Carlo simulation was conducted to evaluate the suggested procedure in comparison to the well-known likelihood ratio (LR) DIF test. Our simulation study results indicated the superiority of FMM over the LR DIF test in terms of detection power and illustrated the importance of accounting for latent heterogeneity in zero-inflated data. The empirical data analysis results further supported the use of FMM by flagging additional DIF items over and above the LR test.

  • Research Article
  • 10.61882/emp.2026.2
Using the M-DIF Online Platform: A Practical Tutorial for DIF Magnitude Estimation
  • Feb 1, 2026
  • Educational Methods and Psychometrics
  • Shan Huang + 1 more

This article presents a practical tutorial for using the M-DIF online platform, a no-code interface designed to estimate the magnitude of Differential Item Functioning (DIF). The M-DIF approach defines DIF magnitude as the predicted difference in item difficulty between a focal group and a reference group. It integrates information from multiple established DIF detection methods together with testing-condition indicators such as group sizes and test length, providing a single continuous estimate with an accompanying uncertainty interval. To make this framework accessible to users without programming experience, we developed an interactive Shiny platform that performs magnitude estimation, visualization, and diagnostic reporting. This tutorial illustrates the platform’s core functions using a subset of data from the 2021 Progress in International Reading Literacy Study (PIRLS). Step-by-step examples guide users through uploading data, specifying groups and anchor items, adjusting visualization settings, interpreting magnitude estimates, and reviewing supplementary diagnostics. The platform offers a practical and efficient tool for researchers and practitioners who seek interpretable, magnitude-centered evidence for item review in educational and psychological assessments.

  • Research Article
  • Cite Count Icon 23
  • 10.1177/00131649921970251
A Comparison of Logistic Regression and Analysis of Variance Differential Item Functioning Detection Methods
  • Dec 1, 1999
  • Educational and Psychological Measurement
  • Marjorie L Whitmore + 1 more

Differential item functioning (DIF) detection rates were compared between logistic regression and analysis of variance for dichotomously scored items. These two DIF methods were compared using simulated binary item response data sets of varying test length (20, 40, and 60 items), sample size (200, 400, and 600 examinees), discrimination type (fixed and varying), and relative underlying ability (equal and unequal) between groups under conditions of uniform DIF, nonuniform DIF, combination DIF, and false positive errors. These test conditions were replicated 100 times. For both DIF detection methods, a test length of 20 items was sufficient for satisfactory DIF detection with detection rate increasing as sample size increased. With the exception of uniform DIF, the logistic regression method had higher mean detection rates than the analysis of variance method. Because the type of DIF present in real data is rarely known, the logistic regression method is recommended for most practical applications.

  • Research Article
  • 10.21449/ijate.1250358
Purification procedures used for the detection of gender DIF: Item bias in a foreign language test
  • Dec 23, 2023
  • International Journal of Assessment Tools in Education
  • Serap Büyükkidik

Differential item functioning (DIF) detection was handled based on “Mantel-Haenszel (MH)”, “Simultaneous item bias test (SIBTEST)”, “Lord's chi-square”, “Raju's area” methods when item purification was performed or item purification was not performed using real data in current study. After detecting gender-related DIF, expert opinions were taken for bias study. It is important to conduct the gender bias research in the English test when purification is performed and when purification is not performed, as there were DIF studies, but there were not completely similar bias studies in the literature. The sample of the research consists of 7389 students who took the “Transition from Primary to Secondary Education Exam (TPSEE, referred to as “TEOG” in Turkey)” administered in April 2017. When gender-related DIF analysis was performed with the four methods, the results were found to differ partially. DIF analysis results differed in the different conditions item purification was performed or not. Detection of DIF was indicative of possible bias. In the second stage of the study, the opinions of seven experts were taken for item 11, for which DIF was detected at least at B level based on MH, SIBTEST. As a result of expert opinion, it was found that there was no item bias according to gender in any item in the English test. It is recommended that similar bias studies can be conducted for test developers to be aware of the features that may lead to item bias and to construct unbiased items.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 23
  • 10.4102/rw.v1i1.4
South African teacher proles and emerging teacher factors: €The picture painted by PIRLS 2006
  • May 22, 2010
  • Reading & Writing
  • Surette Van Staden + 1 more

The Progress in International Reading Literacy Study (PIRLS) assessment is an international comparative study of reading skills of Grade Four learners. South Africa’s "rst participation in the study took place in the 2006 cycle (Mullis et al., 2007), with repeat participation planned to take place for PIRLS 2011. PIRLS 2006 results pointed to serious issues of under achievement among South African Grade Four learners, resulting in the adoption of the National Reading Strategy (Department of Education,2008) and the Foundations for Learning Campaign. While some time has passed since the release of the PIRLS 2006 results, participation in PIRLS 2011 would highlight trends and possible progress made since the PIRLS 2006 study. !is paper reports on the analysis of the Grade Four learner achievement in the PIRLS 2006 assessment into the teacher characteristics, use of resources and instructional practices and analyses of the PIRLS 2006 Teacher Questionnaire data. The main findings outlined by this paper reflects the need for teachers’ continued professional development at Intermediate Phase, the need to employ strategies to retain young teachers and the importance of making available good quality reading materials to schools.

  • Research Article
  • Cite Count Icon 23
  • 10.3102/1076998616659371
Detection of Uniform and Nonuniform Differential Item Functioning by Item-Focused Trees
  • Jul 28, 2016
  • Journal of Educational and Behavioral Statistics
  • Moritz Berger + 1 more

Detection of differential item functioning (DIF) by use of the logistic modeling approach has a long tradition. One big advantage of the approach is that it can be used to investigate nonuniform (NUDIF) as well as uniform DIF (UDIF). The classical approach allows one to detect DIF by distinguishing between multiple groups. We propose an alternative method that is a combination of recursive partitioning methods (or trees) and logistic regression methodology to detect UDIF and NUDIF in a nonparametric way. The output of the method are trees that visualize in a simple way the structure of DIF in an item showing which variables are interacting in which way when generating DIF. In addition, we consider a logistic regression method, in which DIF can be induced by a vector of covariates, which may include categorical but also continuous covariates. The methods are investigated in simulation studies and illustrated by two applications.

  • Research Article
  • 10.1177/01466216261451502
Accounting for CAT-Induced Dependency in Differential Item Functioning Detection: A Multilevel Modeling Framework.
  • May 15, 2026
  • Applied psychological measurement
  • Dandan Danielle Chen Kaptur + 3 more

Differential item functioning (DIF) detection is an important yet understudied problem in computerized adaptive testing (CAT). In this article, we proposed a two-level logistic model to improve DIF detection in CAT by explicitly accounting for nuisance effects arising from CAT-induced structural dependency. First, we conceptualized that adaptive item selection induces systematic dependencies among examinees and items through provisional ability estimates, whereas traditional single-level DIF methods assume independent observations and may yield misleading results in CAT settings. Then, using a numeric example and Monte Carlo simulations, we compared our proposed two-level model with competing single-level models under various CAT conditions, manipulating test length, exposure control, ability estimator, DIF type, and DIF prevalence. Item-level Type-I error and statistical power conditional on joint model convergence were reported for each model. We showed that the proposed two-level model has improved control of spurious DIF and competitive power relative to single-level models, particularly with shorter tests and smaller exposure rates. However, we observed that the model convergence varied systematically across simulated conditions, highlighting that inferential accuracy and convergence reliability are intertwined in complex CAT DIF settings. Through this study, we underscored both the promise of multilevel DIF modeling in CAT and the need for future research to jointly evaluate convergence and inferential performance when assessing DIF models.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 5
  • 10.12955/pss.v1.70
THE COMPARISON OF PIRLS, TIMSS, AND PISA EDUCATIONAL RESULTS IN MEMBER STATES OF THE EUROPEAN UNION
  • Nov 16, 2020
  • Proceedings of CBU in Social Sciences
  • Peter Plavčan

The PIRLS (Progress in International Reading Literacy Study), TIMSS (Trends in International Mathematics and Science Study), and PISA (Programme for International Student Assessment) have become gold standards for the international comparison of children’s performances, when aged 10 and 15 years.
 This paper focuses on secondary analysis of basic statistical indicators on reading literacy (PIRLS), as well as the mathematics and scientific literacy (TIMSS) of pupils at 10 years of age, followed by their reading, mathematics and scientific literacy at 15 years of age (PISA). It compares the pupils’ main educational results in PIRLS and TIMSS with their PSA results. PIRLS, TIMSS, and PISA help to identify key problems within pupils’ educational levels in these selected literacies and create effective educational policy measures.
 One aspect of the comparison within the research paper is the aggregate indicator; this is the arithmetic mean of PIRLS and TIMSS results, using pupils’ PIRLS results from 2001, 2006, 2011 and 2016, and TIMSS results from 2007, 2011 and 2015. The other aspect of the comparison is the aggregate indicator; which is the arithmetic mean of pupils’ PISA results for 2006, 2009, 2012 and 2015. A significant relationship was found to exist between the arithmetic mean of pupils’ PIRLS, TIMSS, and PISA results.
 Political and professional policy decisions within schooling affect the early years of pupils’ school attendance. This has a significant impact on their future education at all levels of schooling. The findings of this paper support a hypothesis regarding the effects of pupils’ educational performance and the need for measures to improve education in schools that should be adopted on an ongoing basis.

  • Research Article
  • Cite Count Icon 55
  • 10.3758/s13428-019-01224-2
A regularization approach for the detection of differential item functioning in generalized partial credit models.
  • Mar 18, 2019
  • Behavior Research Methods
  • Gunther Schauberger + 1 more

Most common analysis tools for the detection of differential item functioning (DIF) in item response theory are restricted to the use of single covariates. If several variables have to be considered, the respective method is repeated independently for each variable. We propose a regularization approach based on the lasso principle for the detection of uniform DIF. It is applicable to a broad range of polytomous item response models with the generalized partial credit model as the most general case. A joint model is specified where the possible DIF effects for all items and all covariates are explicitly parameterized. The model is estimated using a penalized likelihood approach that automatically detects DIF effects and provides trait estimates that correct for the detected DIF effects from different covariates simultaneously. The approach is evaluated by means of several simulation studies. An application is presented using data from the children's depression inventory.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 1
  • 10.3389/feduc.2019.00120
Measurement Comparability of Reading in the English and French Canadian Populations: Special Case of the 2011 Progress in International Reading Literacy Study
  • Dec 6, 2019
  • Frontiers in Education
  • Shawna Goodrich + 1 more

The purpose of this study is to examine item equivalence and score comparability of the Progress in International Reading Literacy Study (PIRLS) 2011for the Canadian French and English language groups. Two methods of differential item functioning were conducted to examine item equivalence across thirteen test booklets designed to assess reading literacy in early years of schooling. Four bilingual reviewers with expertise in reading literacy conducted independent, linguistic and cultural reviews to identify both the degree of item equivalence and potential sources of differences between language versions of released items. Results indicate that an average of 25% of items per booklet function differently at the item level. Reviews by experts indicate differences between the two language versions on some items flagged as displaying differential item function (DIF). Some of these were identified to have linguistic differences pointing to differential difficulty levels in the two language versions.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant