Factor Scores in Small Samples: Recommendations and Solutions
Simultaneous estimation of structural and measurement models in structural equation modeling (SEM) is not always tenable in small samples. In such cases, it may be necessary or advantageous to obtain scores. Current scoring recommendations draw predominantly from simulations with sample sizes greater than N = 200. This paper extends these recommendations to small N, directly comparing factor scores to sum scores. In addition, scores computed from an essentially tau-equivalent factor model are introduced as an alternative scoring option aimed at balancing the competing benefits of sum scores and factor scores. Findings largely suggest that factor scores from an essentially tau-equivalent factor model are advantageous when considering convergence and stability, even when not supported by the data, so long as departures from their assumptions are not substantial. They are obtainable when congeneric factor models fail to converge and have similar correlations with true scores compared with typical factor scores in samples at or less than N = 200.
- Research Article
58
- 10.3758/s13428-022-02016-x
- Nov 17, 2022
- Behavior Research Methods
Commentary in Widaman and Revelle (2022) argued that sum scoring is justified as long as unidimensionality holds because sum score reliability is defined. My response begins with a review of the literature supporting the perspective we adopted in the original article. I then conduct simulation studies to assess the psychometric properties of sum scores created using Widaman and Revelle's justification relative to scores created by the weighted factor score approach in the original article. In my simulations, I generate data where sum and factor scores are correlated at 0.96 or 0.98 because high factor-sum score correlations are often used to support the contention that sum and factor scores have interchangeable psychometric properties. I explore (a) correlations between estimated scores and true scores, (b) classification accuracy of sum and factor scores, and (c) reliability of sum and factor scores. Results show that factor scores have (a) higher correlations with true scores (Δ = 0.02-0.04), (b) higher sensitivity (Δ = 4-8 percentage points), and (c) higher reliability (Δ = 0.04-0.07). Factor score performance metrics also have less sampling variability in most conditions. Psychometric properties of sum scores-even when highly correlated with factor scores-remain less desirable than those of factor scores. Additional considerations like models with multiple factors and measurement invariance are also discussed. Essentially, even if accepting Widaman and Revelle's justification for sum scoring, it is uncertain whether researchers generally would want to sum score after fitting a factor analysis unless sum and factor scores correlate at (and not merely close to) 1.00.
- Research Article
20
- 10.1017/s1092852923000858
- Mar 1, 2023
- CNS Spectrums
transfer, create conditions for the establishment of farmers' behavioral psychological contracts in the process of agricultural land transfers, and guide farmers to establish relationship psychological contracts. The second is to improve the market system, properly cultivate and develop agricultural land transfer intermediaries, reduce transaction costs, and reduce the probability of farmers' psychological contracts being broken. The third is to guide farmers to establish a positive agricultural land transfer psychology based on their resource endowments such as labor force quality and cultural quality, and encourage farmers to make agricultural land transfer decisions such as subcontracting, leasing, reselling, and interchanging.
- Research Article
368
- 10.54055/ejtr.v6i2.134
- Oct 1, 2013
- European Journal of Tourism Research
Primer on Partial Least Squares Structural Equation Modeling (PLS-SEM)In view of its essential role in knowledge creation, multivariate data analysis prevails in the social sciences literature. The field of tourism is not an exception, specifically in the widely adoption of structural equation modeling (SEM), a multivariate technique, by tourism researchers over the past decade. While there are two major types of SEM including covariance-based SEM (CB-SEM) and variance-based SEM (PLS-SEM), the former dominated previous tourism research. However, increasing use of PLS-SEM in tourism research has been witnessed in recent years. This upward trend is likely to persist in the near future given the growing popularity of PLS-SEM in other social sciences domains like marketing, strategic management, and management information system, as specified in the preface of the book. Indeed, PLS-SEM, in relative to CB- SEM, provides more flexibility in handling of data. For instance, PLS-SEM is well-suited for accommodating small sample sizes and complex model, fortesting a model containing both formative and reflective constructs, and for handling single-item measures. To this end, the timely introduction of the book A Primer on Partial Least Squares Structural Equation Modeling (PLS-SEM) helps tourism researchers stand at the front edge of the SEM technique and make effective use of the PLS-SEM in data analysis. Additionally, the book illustrates the application of PLS-SEM with a free downloadable software namely SmartPLS which is essential to extend the application of PLS-SEM in tourism research.Authored by Hair, Hult, Ringo, and Sarstedt, the book consists of eight chapters. To equip the readers with the basic knowledge of PLS- SEM, Chapter 1 delineates the meaning of SEM and its relationship with multivariate data analysis, followed by a description of the major elements in multivariate data analysis. Then the basic elements of PLS-SEM are explained. Finally, PLS-SEM is distinguished from its counterpart namely CB-SEM while the major characteristics of PLS-SEM and the conditions where the PLS-SEM are more adequate than CB-SEM and vice versa are discussed. To step in the application of PLS- SEM, Chapter 2 firstly explicates the concepts in structural model specification including mediation, moderation, and higher-order models. Then specification of measurement model is explained with a special focus on the differences between reflective and formative measures. After that, the issues that need to be addressed after data collection are discussed. The chapter ends by creating the model in the SmartPLS is illustrated. With an established model, Chapter 3 focuses on model estimation. The chapter explains the algorithm underpinning the estimation and the statistical properties of the PLS-SEM method, as well as the options and parameter settings for running the algorithm. Following that, the issues about interpretation of results are explained. The final section illustrates the execution of model estimation in the SmartPLS.Based on the model estimation, empirical measures of the measurement and structural models are derived, where evaluation of the models takes place. Chapter 4 exhibits the major steps in model evaluation in the beginning. Thereafter, the chapter explains the evaluation of reflective measurement models according to three major criteria including internal consistency reliability, convergent validity, and discriminant validity, followed by an illustration with the SmartPLS. Chapter 5 explains the assessment of formative measurement models with respect to the criteria of convergent validity, collinearity, and significance and relevance of the formative indicators. The chapter also elucidates the basic concepts of bootstrapping which is used to examine the statistical significance of estimates in PLS- SEM. An illustration of the assessment of formative measurement model in the SmartPLS follows. Chapter 6 continues the topic on model evaluation by focusing on the assessment of structural model. …
- Research Article
107
- 10.1080/00273171.2012.730072
- Jan 1, 2013
- Multivariate Behavioral Research
Factor score estimation is a controversial topic in psychometrics, and the estimation of factor scores from exploratory factor models has historically received a great deal of attention. However, both confirmatory factor models and the existence of missing data have generally been ignored in this debate. This article presents a simulation study that compares the reliability of sum scores, regression-based and expected posterior methods for factor score estimation for confirmatory factor models in the presence of missing data. Although all methods perform reasonably well with complete data, expected posterior-weighted (full) maximum likelihood methods are significantly more reliable than sum scores and regression estimators in the presence of missing data. Factor score reliability for complete data can be predicted by Guttman's 1955 formula for factor communality. Furthermore, factor score reliability for incomplete data can be reasonably approximated by communality raised to the power. An empirical demonstration shows that the full maximum likelihood method best preserves the relationship between nicotine dependence and a genetic predictor under missing data. Implications and recommendations for applied research are discussed.
- Research Article
6
- 10.19184/mims.v19i2.17272
- Sep 1, 2019
- Majalah Ilmiah Matematika dan Statistika
Structural Equation Model (SEM) is a statistical technique with simultaneous processing involves measurement errors, indicator variables, and latent variables. SEM is used to test hypotheses that state the relationships between latent variables when latent variables have been assessed through each of the indicator variables. Multiple Group SEM is a basic model analysis that uses more than one sample. This analysis aims to determine whether the components or models of measurement and structural models are invariant for the two sample groups. In this study, the data generated by some requirements. First, the data generated with sample size n = 250. The first generated data is homogeneous data where the measurement model is the same as the structural model in group 1 and group 2, while the second data is non-homogeneous data where the measurement model and the structural model in group 1 and group 2 is not the same. The data was analyzed using the help of the lavaan package available in R to obtain SEM estimation results and Goodness of Fit Model from some data that was formed. From the results of the merger of the two groups, it shows that the invariant of the two models with the largest df (63) which is Fit Mean model states the simplest model. However, the smallest df (48) with Fit.configural model states the most complex model.
 Keywords: SEM, Multiple Group, R Program
- Research Article
- 10.64753/jcasc.v11i1.4697
- Mar 19, 2026
- Journal of Cultural Analysis and Social Change
This study aimed to mediate the effect of employee motivation on the relationship between selected leadership styles and employee performance in selected financial institutions in Mthatha, Eastern Cape. Using a quantitative approach, data were collected via self-administered online questionnaires and manual copies. A total of 155 respondents were given the questionnaires. Inferential analysis was conducted with Smart PLS and Structural Equation Modelling to assess multivariate causal relationships. Findings indicated that the selected leadership styles significantly positively correlated with employee motivation and performance. Partial Least Squares Structural Equation Modelling (PLS-SEM) implemented through SmartPLS offers robustness, particularly in studies with complex mediating relationships and relatively small sample sizes (Hair, 2014). SEM allows simultaneous estimation of measurement and structural models, which is particularly valuable in industrial psychology studies like this one. The study recommends that financial institutions continue to design and implement leadership development programs tailored to foster motivational leadership styles and improve employee motivation and performance. These styles have demonstrated a strong positive effect on employee motivation, significantly improving employee performance.
- Research Article
36
- 10.1177/0013164419854492
- Jun 17, 2019
- Educational and Psychological Measurement
Recently, quantitative researchers have shown increased interest in two-step factor score regression (FSR) approaches to structural model estimation. A particularly promising approach proposed by Croon involves first extracting factor scores for each latent factor in a larger model, then correcting the variance-covariance matrix of the factor scores for bias before using this matrix as input data in a subsequent regression analysis or path model. Although not immediately obvious, Croon's bias correction formulas are predicated upon the standard assumption of conditionally independent uniquenesses (measurement residuals). To our knowledge, the method's performance has never been evaluated under conditions in which this assumption is violated. In the present research, we rederive Croon's formulas for the case of correlated uniqueness and present the results of two Monte Carlo simulations comparing the method's performance with standard methods when the unique factors were correlated in the population model. In our simulations, our proposed Croon FSR approaches outperformed methods that blindly assumed conditionally independent uniquenesses (e.g., uncorrected FSR, traditional Croon FSR, structural equation modeling [SEM] using standard specification), performed comparably to a correctly specified SEM, and outperformed SEMs that correctly specified the unique factor covariances but misspecified the structural model. We discuss the implications of our results for substantive researchers.
- Research Article
5
- 10.3102/1076998620911920
- Apr 8, 2020
- Journal of Educational and Behavioral Statistics
We address measurement error bias in propensity score (PS) analysis due to covariates that are latent variables. In the setting where latent covariate X is measured via multiple error-prone items W, PS analysis using several proxies for X—the W items themselves, a summary score (mean/sum of the items), or the conventional factor score (i.e., predicted value of X based on the measurement model)—often results in biased estimation of the causal effect because balancing the proxy (between exposure conditions) does not balance X. We propose an improved proxy: the conditional mean of X given the combination of W, the observed covariates Z, and exposure A, denoted [Formula: see text]. The theoretical support is that balancing [Formula: see text] (e.g., via weighting or matching) implies balancing the mean of X. For a latent X, we estimate [Formula: see text] by the inclusive factor score (iFS)—predicted value of X from a structural equation model that captures the joint distribution of [Formula: see text] given Z. Simulation shows that PS analysis using the iFS substantially improves balance on the first five moments of X and reduces bias in the estimated causal effect. Hence, within the proxy variables approach, we recommend this proxy over existing ones. We connect this proxy method to known results about valid weighting/matching functions. We illustrate the method in handling latent covariates when estimating the effect of out-of-school suspension on risk of later police arrests using National Longitudinal Study of Adolescent to Adult Health data.
- Research Article
- 10.1177/00131644251399588
- Jan 4, 2026
- Educational and psychological measurement
Researchers in the behavioral, educational, and social sciences often aim to analyze relationships among latent variables. Structural equation modeling (SEM) is widely regarded as the gold standard for this purpose. A straightforward alternative for estimating the structural model parameters is uncorrected factor score regression (UFSR), where factor scores are first computed and then employed in regression or path analysis. Unfortunately, the most commonly used factor scores (i.e., Regression and Bartlett factor scores) may yield biased estimates and invalid inferences when using this approach. In recent years, factor score regression (FSR) has enjoyed several methodological advancements to address this inconsistency. Despite these advancements, the use of FSR with correlation-preserving factor scores, here termed consistent factor score regression (cFSR), has received limited attention. In this paper, we revisit cFSR and compare its advantages and disadvantages relative to other recent FSR and SEM methods. We conducted an extensive simulation study comparing cFSR with other estimation approaches, assessing their performance in terms of convergence rate, bias, efficiency, and type I error rate. The findings indicate that cFSR outperforms UFSR while maintaining the conceptual simplicity of UFSR. We encourage behavioral, educational, and social science researchers to avoid UFSR and adopt cFSR as an alternative to SEM.
- Supplementary Content
91
- 10.1108/ejm-04-2023-0307
- Feb 8, 2024
- European Journal of Marketing
PurposeThe purpose of this paper is to assess the appropriateness of equal weights estimation (sumscores) and the application of the composite equivalence index (CEI) vis-à-vis differentiated indicator weights produced by partial least squares structural equation modeling (PLS-SEM).Design/methodology/approachThe authors rely on prior literature as well as empirical illustrations and a simulation study to assess the efficacy of equal weights estimation and the CEI.FindingsThe results show that the CEI lacks discriminatory power, and its use can lead to major differences in structural model estimates, conceals measurement model issues and almost always leads to inferior out-of-sample predictive accuracy compared to differentiated weights produced by PLS-SEM.Research limitations/implicationsIn light of its manifold conceptual and empirical limitations, the authors advise against the use of the CEI. Its adoption and the routine use of equal weights estimation could adversely affect the validity of measurement and structural model results and understate structural model predictive accuracy. Although this study shows that the CEI is an unsuitable metric to decide between equal weights and differentiated weights, it does not propose another means for such a comparison.Practical implicationsThe results suggest that researchers and practitioners should prefer differentiated indicator weights such as those produced by PLS-SEM over equal weights.Originality/valueTo the best of the authors’ knowledge, this study is the first to provide a comprehensive assessment of the CEI’s usefulness. The results provide guidance for researchers considering using equal indicator weights instead of PLS-SEM-based weighted indicators.
- Research Article
2
- 10.1080/19439962.2020.1716908
- Feb 14, 2020
- Journal of Transportation Safety & Security
The severity of traffic barrier crashes has been modeled in the literature based on human, road, and traffic barrier factors. However, all these factors interact in a complicated way so the relationship between these factors still remains unclear. A structural equation model (SEM) can be used to capture the intricate relationships between the contributory factors and the latent or unobservable factors. This study was conducted using the SEM to model the complicated relationships between the confounding factors and traffic barrier crash severity. Due to possible differences across different road classifications, the SEM was applied separately to the highway and interstate systems. The SEM involves the measurement or confirmatory factor analysis (CFA) model and the structural model involving the interrelationships between the factors. This study evaluates traffic barrier crash severity in terms of numbers of death, injury, and severity of crashes. This study examined the nature and causes of severe traffic barrier crashes in Wyoming. The results indicated that road conditions, traffic barrier types, risky driving, and force are factors that need to be considered in addressing the severity of traffic barrier crashes. The methodology in this study also addresses categorical predictors and model selection techniques for specifying the measurement model and the structural model which make up the SEM. A careful sequential modeling framework was used to build the structural model characterizing the interrelationships among the latent variables as well as evaluation of the fitted SEM based on goodness of fit indices.
- Dissertation
- 10.21248/gups.87712
- Jan 1, 2024
In the social sciences, relations among latent variables are of great interest. These associations are not necessarily linear, and selecting the correct model is of significant importance; otherwise, the inferences drawn from these models might be arbitrarily wrong. Latent variables cannot be measured directly but are operationalized indirectly via multiple indicators. Structural Equation Modeling (SEM) is used to examine the associations between latent variables through the correlational structure in multiple measurements (Bollen, 1989). The model selection process is guided by (robust) model fit tests, such as the (robust) χ2-test (e.g., Satorra & Bentler, 2010) in linear SEM analysis. Although there are extensions for specific quadratic and interaction SEM (QISEM, Büchner & Klein, 2020), these model fit tests are not generally applicable to nonlinear SEM (NLSEM) due to the lack of a saturated comparison model (Büchner & Klein, 2020; Mooijaart & Satorra, 2009). In regression analysis, non-parametric trend estimates and scatterplots can be used to examine the structural form and guide model selection. However, these are not suitable for NLSEM due to measurement errors in the observations. Latent variables cannot be observed directly, but estimates for these are available by using factor scores. Factor scores are transformations of the data that estimate the latent variables with approximation error. Still, factor scores are indeterminate (Grice, 2001), meaning that several approaches will result in different results. However, as the number of measurements increases, the approximation error diminishes, and factor scores become close estimates of the latent variables. In this thesis in Manuscript A, a simple set of assumptions is derived to identify the trends in NLSEM using linear factor scores. These trends are described by the conditional expectation of endogenous variables given exogenous variables. Identification means that the trends can be estimated from observed data. Linear factor scores are simple linear transformations of the data and therefore are straightforward to compute. However, due to the approximation error in linear factor scores, these assumptions are asymptotic in nature, as they are only applicable for a large number of measurements per latent variable, as then the prediction error becomes negligible or can be explicitly modeled. In contrast to previous identification results (Kelava et al., 2017), these asymptotic assumptions used in this thesis allow for cross-relations between the measurements within the exogenous part of the model and within the endogenous part of the model. This implies that cross-loadings and residual covariances are permitted within measurements of the exogenous latent variables (ξ) and within measurements of the endogenous latent variables (η), while cross-relations between measurements of the exogenous part and the endogenous part of the model are disallowed. Independence assumptions on the measurement errors of the exogenous and endogenous parts of the model, along with independence from the latent variables, imply that the Bartlett (1937) factor score (BFS) has a specific structure that connects it to the literature on non-parametric regression with measurement error (Delaigle et al., 2009; Huang & Zhou, 2017). Although, in this literature, the distribution of the measurement error is assumed to be known, this similarity, combined with continuity assumptions on the nonlinear trend, can be used to identify the conditional expectation. This is shown mathematically in Manuscript A. Consequently, trends in NLSEM can be estimated using BFS under the given assumptions either by ignoring the approximation error or by explicitly modeling the approximation error. For a large number of measurements in the exogenous part of the model, the approximation error variance of the BFS on the exogenous part of the model is small. Therefore, this approximation error can be ignored by inputting BFS into non-parametric regression methods such as the LOESS (locally estimated scatterplot smoothing, Cleveland, 1979, 1981). Since BFS are linear transformations of the data, a central limits effect occurs on the approximation error.
- Single Book
6348
- 10.4324/9781315749105
- Dec 22, 2015
This book provides the reader with a review of correlation and covariance among variables, followed by multiple regression and path analysis techniques to better understand the building blocks of structural equation modelling. The concepts behind measurement models are introduced to illustrate how measurement error impacts statistical analyses, and structural models are presented that indicate how latent variable relationships can be established. Examples are included throughout to make the concepts clear to the reader. The structural equation modelling examples are presented using either EQS5.0 or LISREL8-SIMPLIS programming language, both of which have an easy-to-use set of commands to specify measurement and strucural models. No complicated programming is required, nor does the reader need an advanced understanding of statistics of matrix algebra. A goal in writing this volume was to focus conceptually on the steps one takes in analyzing theoretical models. These steps encompass: specifying a model based upon theory or prior research; determining whether the model can be identified to have unique estimates for variables in the model; selecting an appropriate estimation method based on the distributional assumptions of variables; testing the model and interpreting fit indices; and finally respecifying a model based on suggested modification indices, which involves adding or dropping paths in the model to obtain a better model fit. The resources and references provided in this book should equip faculty, students and researchers to enhance their working knowledge of structural equation modelling. Not intended as an in-depth presentation of statistics or factor analysis, this text focuses on the basic ideas and principles behind structural equation modelling. Assuming that the reader has a basic understanding of correlation, the authors have built upon this understanding to present these basic ideas and principles.
- Conference Article
- 10.2991/icssr-13.2013.8
- Jan 1, 2013
In this paper, we have discussed and analyzed the impact on the university library service image using questionnaire data of the college students in Zhejiang Area based on the Structural Equation Model (SEM). The determinants include the service time, the service consciousness, the service space and the service attitude. And not only that, but this study put forward tactics for the improve university library service image based on this paper's result.
- Research Article
7
- 10.1177/014662168901300311
- Sep 1, 1989
- Applied Psychological Measurement
An alternative procedure for estimating structural equations models is described. The two-stage proce dure, Path Analysis of Covariance Matrix (PACM), sep arately estimates the measurement and structural models using standard least-squares procedures. PACM was empirically compared to simultaneous maximum likelihood estimation of measurement and structural models using LISREL. PACM produced results similar to LISREL in many cases; it also seems to have advan tages when dealing with large-scale problems, model misspecifications, collinearity among indicators, and missing data.