Measure Length in Research Settings: Effects on Cognitive and Non-Cognitive Scores and Implications for Psychometric Properties
ABSTRACT When administering measures, researchers are often concerned about the psychological burden of measure length (i.e., the number of items) on participants and its potential effect on test scores and resulting psychometric properties. The present work experimentally examined this issue in two studies using two independent samples of undergraduate students and a between-subjects design whereby participants of each sample were randomly assigned to complete either a short or long measure. Study 1 (N = 398) examined the effect of measure length on careless responding on a non-cognitive measure and the resulting psychometric properties (i.e., internal consistency and construct validity). Study 2 (N = 426) examined the effect of measure length on performance scores on a cognitive measure and the same resulting psychometric properties. Measure length had no effect on performance on the cognitive measure, and for the non-cognitive measure had an inconsistent effect on careless responding. Pertaining to psychometric properties, measure length had no meaningful effect on the internal consistency of both measures. The effect on construct validity depended on the criterion, but on average, was small to negligible. Overall, the results provide little evidence to support researchers’ concerns about the potential adverse psychological effects associated with using longer measures.
- Front Matter
40
- 10.1111/add.13221
- Dec 14, 2015
- Addiction
Keywords: Careless responding; data cleaning; data integrity; invalid responding; online survey; research outcome quality
- Research Article
- 10.1097/gme.0000000000002494
- Mar 11, 2025
- Menopause (New York, N.Y.)
Many midlife women report cognitive issues when they transition through menopause. These cognitive complaints affect women's mental health and quality of life. However, the current understanding of women's cognitive experiences during the menopause transition has been limited by the lack of validated self-reported cognitive measures. This systematic review aimed to identify existing self-reported, or subjective, cognitive measures used in menopause research and evaluate their psychometric properties and applicability. Three databases, Medline, Embase, and PsycINFO, were searched in March 2024 with no restriction on publication year. Studies investigating women transitioning into postmenopause and with cognitive experiences measured using validated subjective cognitive measures were selected. The assessment of psychometric properties and applicability of included measures was conducted based on their development process and their performance in the menopause studies selected. Twenty-eight menopause studies involving 15 measures were included. Included measures showed adequate content validity, internal consistency, and construct validity when they were developed, yet other psychometric properties were either poor or not reported. Hence, the overall performance of included measures was generally moderate to poor. Information relating to psychometric properties of included measures in menopause studies was also lacking, indicating doubtful applicability. Poor psychometric properties or the lack of psychometric assessment of existing subjective cognitive measures may indicate doubt or uncertainty regarding their applicability in women transitioning through menopause. This review recommends the use of subjective cognitive measures that assess more than one cognitive domain, as well as further assessment of the psychometric properties of these measures before their use in menopause research or clinical settings, particularly those measures initially developed for clinical practice. It also highlights the need for future development of a subjective cognitive measure for women transitioning through menopause to improve the current understanding of their cognitive challenges.
- Research Article
25
- 10.1027/1015-5759/a000609
- Jul 1, 2020
- European Journal of Psychological Assessment
When constructing a questionnaire to assess a psychological construct, one important decision researchers have to make is how to collect responses from test takers; that is, which response format to implement.We argued in a previous editorial published in the European Journal of Psychological Assessment (EJPA) that this decision deserves more attention and should be an explicit step in the test construction process (Wetzel & Greiff, 2018).The reason for this is that it can be a consequential decision that influences the validity of conclusions we draw about test takers' trait levels or about relations between constructs and criteria (Brown & Maydeu-Olivares, 2013; Wetzel & Frick, 2020).In this editorial, which can be considered a followup to the first one, we will take a closer look at two response formats 1 : rating scales (RS), the current default in most questionnaires, and the multidimensional forced-choice (MFC) format, an alternative that is currently the focus of a considerable body of research.We will first define the two formats and point out some of their advantages and disadvantages.Then, we will provide a summary and evaluation of research comparing RS and MFC.Third, we will draw some preliminary conclusions on the feasibility of applying MFC as an alternative to RS. Fourth, we will point out some open research questions.We will end with some recommendations and implications for readers and authors of EJPA.In this editorial, the overall goal is to give researchers and test users an overview of the current state of the research on RS versus MFC and to provide guidance on the feasibility of applying MFC in research on psychological assessment.1 The multidimensional forced-choice format is both an item and a response format.For simplicity in the comparison with rating scales, we refer to it as response format.
- Research Article
- 10.3389/fpsyg.2026.1815225
- Jan 1, 2026
- Frontiers in Psychology
IntroductionCareless responding, in which survey participants fail to attend to item content, is a well-recognized threat to the quality of self-report data. Although prevalence estimates commonly fall between 8 and 12% in student samples, the extent to which careless responding distorts the psychometric properties of attitude measures has received limited attention, particularly outside Western samples.MethodsThe present study investigated the prevalence and psychometric consequences of careless responding in a paper-and-pencil sample of 1,112 Turkish university students who completed the sustainable development awareness scale. An instructed response item embedded within the scale identified 126 respondents (11.33%) as careless, and two post-hoc indicators (longstring and even-odd consistency) converged with this classification. Parallel analyses on the unscreened and screened samples were conducted to evaluate effects on internal consistency, confirmatory factor analysis, multigroup measurement invariance, criterion-related validity correlations, and an item-level Composite Sensitivity Index (CSI).ResultsInternal consistency was higher in the screened sample, with the gains concentrated in the two subscales containing reverse-coded items. Confirmatory factor analysis indices trended in the direction of better fit after screening, and criterion-validity correlations with related constructs were modestly larger in the screened sample at both the manifest and latent levels. Multigroup measurement invariance testing across attentive and careless responders supported metric but not scalar invariance, pointing to systematic intercept differences consistent with acquiescent responding among careless respondents. The composite sensitivity index combining changes in item means, item-total correlations, and factor loadings was used to rank items by vulnerability to careless responding, with results robust to an alternative standardized aggregation. All six reverse-coded items appeared among the ten most sensitive items, despite constituting only one-sixth of the scale, and the item immediately following the attention check showed a notably elevated sensitivity pattern that we interpret as a tentative hypothesis worth further investigation.DiscussionThese findings underscore the importance of routine screening for careless responding and offer practical guidance on the placement of attention checks and the use of reverse-coded items.
- Research Article
10
- 10.1037/adb0000924
- Feb 1, 2024
- Psychology of addictive behaviors : journal of the Society of Psychologists in Addictive Behaviors
The prevalence of research conducted online in the addiction field has increased rapidly over the past decade. However, little focus has been given to careless responding in these online studies, despite the issues it may cause for statistical inference and generalizability. Our aim was to examine whether alcohol use is associated with careless responses. Raw data were requested from online studies examining alcohol use and related problems which also addressed careless responding. We obtained 13 data sets of 12,237 participants (Mage = 42.16, SD = 15.65, 50.5% female). The sample had an average Alcohol Use Disorders Identification Test (AUDIT) score of 10.88 (SD = 7.77). Predictors included demographic information (age, gender) and AUDIT total scores. The primary outcome was whether an individual was classed as a careless responder, for example, by failing an explicit attention check question. AUDIT total scores were associated with careless responding (OR = 1.07, 95% CI [1.06, 1.08], p < .001). Hazardous drinking or worse was associated with 2.21 greater odds (OR = 2.21, 95% CI [1.81, 2.71] of careless responding, whereas harmful drinking or worse was associated with 3.43 greater odds (OR = 3.43, 95% CI [2.83, 4.17]) and probable dependence was associated with 3.63 greater odds (OR = 3.63, 95% CI [2.95, 4.48]). Alcohol use and related problems are positively associated with careless responding in online research. Removal of individuals identified as careless responders may lead to issues of generalizability, and more care should be taken to identify and handle careless responder data. (PsycInfo Database Record (c) 2024 APA, all rights reserved).
- Research Article
13
- 10.3758/s13428-023-02074-9
- Feb 3, 2023
- Behavior research methods
It is common to model responses to surveys within latent variable frameworks (e.g., item response theory [IRT], confirmatory factor analysis [CFA]) and use model fit indices to evaluate model-data congruence. Unfortunately, research shows that people occasionally engage in careless responding (CR) when completing online surveys. While CR has the potential to negatively impact model fit, this issue has not been systematically explored. To better understand the CR-fit linkage, two studies were conducted. In study 1, participants' response behaviors were experimentally shaped and used to embed aspects of a comprehensive simulation (study 2) with empirically informed data. For this simulation, 144 unique conditions (which varied the sample size, number of items, CR prevalence, CR severity, and CR type), two latent variable models (IRT, CFA), and six model fit indices (χ2, RMSEA, SRMSR [CFA] and M2, RMSEA, SRMSR [IRT]), were examined. The results indicated that CR deteriorates model fit under most circumstances, though these effects are nuanced, variable, and contingent on many factors. These findings can be leveraged by researchers and practitioners to improve survey methods, obtain more accurate survey results, develop more precise theories, and enable more justifiable data-driven decisions.
- Research Article
- 10.1016/j.jad.2025.120579
- Feb 1, 2026
- Journal of affective disorders
Psychosocial functioning in older adults with bipolar disorder: Spanish validation of the Functioning Assessment Short Test for Older adults (FAST-O).
- Research Article
9
- 10.1007/s10862-016-9573-7
- Oct 24, 2016
- Journal of Psychopathology and Behavioral Assessment
Cognitive factors, including beliefs, thoughts and assumptions have been found to play an important role in the development and maintenance of Social Anxiety Disorder. Trait cognitive self-report measures of social anxiety are widely used in research and clinical settings. It is imperative that only measures with good psychometric properties are used in order to interpret assessment scores accurately, and to make valid and reliable conclusions. The present systematic review evaluated the psychometric properties of trait cognitive self-report measures of social anxiety. Relevant studies were identified via a comprehensive and systematic search of academic databases. The reported psychometric properties of included studies were analysed by applying an appraisal of adequacy tool developed by Terwee et al. (2007). Of the 3091 studies identified, 50 studies met the inclusion criteria, and they included 21 measures. Included studies demonstrated that a number of measures had some adequate psychometric properties, however, no measure fulfilled criteria for all psychometric properties according to the appraisal tool. Findings highlight the need to further establish the psychometric properties of cognitive self-report measures of social anxiety in clinical and research settings through additional empirical studies.
- Research Article
44
- 10.1037/pha0000546
- Aug 1, 2022
- Experimental and Clinical Psychopharmacology
Crowdsourcing-the process of using the internet to outsource research participation to "workers"-has considerable benefits, enabling research to be conducted quickly, efficiently, and responsively, diversifying participant recruitment, and allowing access to hard-to-reach samples. One of the biggest threats to this method of online data collection however is the prevalence of careless responders who can significantly affect data quality. The aims of this preregistered systematic review and meta-analysis were: (a) to examine the prevalence of screening for careless responding in crowdsourced alcohol-related studies; (b) to examine the pooled prevalence of careless responding; and (c) to identify any potential moderators of careless responding across studies. Our review identified 96 eligible studies (∼126,130 participants), of which 51 utilized at least one measure of careless responding, 53.2%, 95% CI [42.7%-63.3%]; ∼75,334 participants. Of these, 48 reported the number of participants identified by careless responding method(s) and the pooled prevalence rate was ∼11.7%, 95% CI [7.6%-16.5%]. Studies using the MTurk platform identified more careless responders compared to other platforms, and the number of careless response items was positively associated with prevalence rates. The most common measure of careless responding was an attention check question, followed by implausible response times. We suggest that researchers plan for such attrition when crowdsourcing participants and provide practical recommendations for handling and reporting careless responding in alcohol research. (PsycInfo Database Record (c) 2022 APA, all rights reserved).
- Research Article
- 10.12982/jams.2022.004
- Jan 2, 2023
- Journal of Associated Medical Sciences
Background: Occupational therapy (OT) cognitive interventions requires a standardised cognitive outcome measure to help explain the effectiveness of the interventions. Now, there is a lack of measures to use for Thai older adults with cognitive impairments. Therefore, a new Occupational Therapy Cognitive Outcome Measure (OTCOM) for Thai older adults with cognitive impairments was developed to support evidence-based OT cognitive interventions. Objectives: To examine the psychometric properties including internal consistency, inter-rater and intra-rater reliability, known-group construct validity, concurrent validity, and responsiveness of the OTCOM. Materials and methods: A prospective cohort design was used in this study. One hundred and ten older adults; sixty-one older adults with cognitive impairments and forty-nine older adults without cognitive impairments, were recruited. The Cronbach’s alpha coefficient was calculated for internal consistency. Intraclass correlation coefficient (ICC) was used to analyse rater reliability. Analyses of concurrent and known-group construct validity were done using Pearson correlation and independent t-test, respectively. Both effect size (ES) and standardised response mean (SRM) were calculated for responsiveness of the OTCOM. Results: The results showed good internal consistency (α=0.88), and excellent inter-rater and intra-rater reliability (ICC=0.99). A high correlation between the OTCOM and the Dynamic Lowenstein Occupational Therapy Cognitive Assessment-Geriatric (DLOTCA-G) and the Thai Cognitive-Perceptual Test (Thai-CPT) was found, indicating good concurrent validity. There was a significant difference between older adults with cognitive impairments and without cognitive impairments, suggesting good construct validity by the known-group method. Responsiveness was shown as large ES and SRM in the total score. Conclusion: The OTCOM showed good psychometric properties, making it useful in OT practice after revisions
- Research Article
113
- 10.3310/hta16010
- Jan 1, 2012
- Health Technology Assessment
To produce a robust measure of social inclusion [Social and Community Opportunities Profile (SCOPE)] that is multidimensional and captures multiple life domains; incorporates objective and subjective indicators of inclusion; has sound psychometric properties including responsiveness; facilitates benchmark comparisons with normative general population and mental health samples [including common mental disorder (CMD) and severe mental illness groups]; can be used with people with mental health problems receiving support from mental health services or not; and can be used across a range of community service settings. Phase I: conceptual framework developed from a review of the literature and concept mapping. Phase II: questionnaire developed including UK national population surveys and other normative data. Pre-testing using cognitive appraisal and evaluation then pilot testing in a small convenience sample. Preliminary testing (following modification) in community (n = 252) and mental health service users (MHSUs) samples (n = 43). Data reduction including factor analysis and Mokken scaling for polytomous item response analysis then psychometric evaluation, including internal consistency and discriminant and construct validity. Test-retest reliability assessed in a convenience sample of students (n = 119). Final testing in clinical services including psychometric evaluation and responsiveness testing. The community sample was set in participants' households across the UK. The MHSU sample was set in a south Wales resource centre. The student sample was set in a university. The community sample was randomly selected from the postal address file in five areas in England and Wales. Forty people in this sample were subgrouped as having a CMD based on their responses to the Mental Health Index five items. Two MHSU samples were obtained from existing services. Psychometric testing on the field data from the SCOPE long version demonstrated good internal consistency of all scales (alpha ≥ 0.7), good construct validity, with SCOPE scales correlating highly with each other sharing between 40% and 61% of variance and a close but lesser association with community participation and social capital. Chi-squared tests on objective items and analysis of variance between groups on SCOPE scales demonstrated good discriminant validity between different mental health groups (and better than the Mokken scaling results). Acceptability was good, with 77% of the service user sample finding the SCOPE domains relevant. The number of items in SCOPE decreased from 121 to 48 following data reduction. Scales in the short version of SCOPE retained reasonable internal consistency (alpha between 0.60 and 0.75). Test-retest reliability demonstrated reliability over time, with strong associations between all items over a 2-week period. Repeating the discriminant validity tests on the short version demonstrates good discriminant validity between the mental health groups. Acceptability improved, with 90% of the sample describing questions as relevant to them. The main aim of producing an instrument with good psychometric properties for use in research and clinical settings, namely the SCOPE short version, was achieved. Ongoing data collection will enable responsiveness testing in the future. Further research is needed including larger samples of minority and disadvantaged groups, including those with physical illnesses and disabilities, and specific mental health diagnostic groups. The National Institute for Health Research Health Technology Assessment programme.
- Research Article
57
- 10.1007/s11136-017-1767-2
- Dec 16, 2017
- Quality of Life Research
Quality of life (QoL) measurement relies upon participants providing meaningful responses, but not all respondents may pay sufficient attention when completing self-reported QoL measures. This study examined the impact of careless responding on the reliability and validity of Internet-based QoL assessments. Internet panelists (n = 2000) completed Patient-Reported Outcomes Measurement Information System (PROMIS®) short-forms (depression, fatigue, pain impact, applied cognitive abilities) and single-item QoL measures (global health, pain intensity) as part of a larger survey that included multiple checks of whether participants paid attention to the items. Latent class analysis was used to identify groups of non-careless and careless responders from the attentiveness checks. Analyses compared psychometric properties of the QoL measures (reliability of PROMIS short-forms, correlations among QoL scores, "known-groups" validity) between non-careless and careless responder groups. Whether person-fit statistics derived from PROMIS measures accurately discriminated careless and non-careless responders was also examined. About 7.4% of participants were classified as careless responders. No substantial differences in the reliability of PROMIS measures between non-careless and careless responder groups were observed. However, careless responding meaningfully and significantly affected the correlations among QoL domains, as well as the magnitude of differences in QoL between medical and disability groups (presence or absence of disability, depression diagnosis, chronic pain diagnosis). Person-fit statistics significantly and moderately distinguished between non-careless and careless responders. The results support the importance of identifying and screening out careless responders to ensure high-quality self-report data in Internet-based QoL research.
- Research Article
- 10.1017/psy.2025.10041
- Sep 1, 2025
- Psychometrika
Visual Analogue scales (VASs) are increasingly popular in psychological, social, and medical research. However, VASs can also be more demanding for respondents, potentially leading to quicker disengagement and a higher risk of careless responding. Existing mixture modeling approaches for careless response detection have so far only been available for Likert-type and unbounded continuous data but have not been tailored to VAS data. This study introduces and evaluates a model-based approach specifically designed to detect and account for careless respondents in VAS data. We integrate existing measurement models for VASs with mixture item response theory models for identifying and modeling careless responding. Simulation results show that the proposed model effectively detects careless responding and recovers key parameters. We illustrate the model’s potential for identifying and accounting for careless responding using real data from both VASs and Likert scales. First, we show how the model can be used to compare careless responding across different scale types, revealing a higher proportion of careless respondents in VAS compared to Likert scale data. Second, we demonstrate that item parameters from the proposed model exhibit improved psychometric properties compared to those from a model that ignores careless responding. These findings underscore the model’s potential to enhance data quality by identifying and addressing careless responding.
- Research Article
36
- 10.3389/fpsyg.2019.01258
- Jun 14, 2019
- Frontiers in Psychology
The current research investigates the impact of careless responding on factorial analytic results and construct validity with real data. Results showed that inclusion of careless respondents in data analysis distorts factor loading pattern and hinders recovery of theoretical existing factors. Careless respondents also blur the distinction of theoretically distinct factors, resulting in higher inter-factor correlations. That careless responding may threaten convergent validity also receives limited support. Researchers are advised to exclude careless respondents before statistical analysis.
- Research Article
- 10.1186/s41687-025-00910-4
- Jun 20, 2025
- Journal of Patient-Reported Outcomes
BackgroundPatient-reported outcome measures (PROMs) are important for assessing premenstrual syndrome (PMS) and premenstrual dysphoric disorder (PMDD) to effectively capture subjective symptom burden and evaluate treatment effectiveness in clinical and research settings. This systematic review evaluated the psychometric properties of PROMs used to assess PMS/PMDD in Japan.MethodologyA systematic literature search was conducted in the MEDLINE, CINAHL, Cochrane Library, and Ichushi-Web databases. The COnsensus-based Standards for the selection of health Measurement Instruments (COSMIN) methodology was used to assess the methodological quality and measurement properties of the included PROMs.ResultsA total of 13 studies that evaluated 12 versions of 11 unique PROMs were included. PROMs were categorized as recall-based (n = 9, 69%) or daily recording scales (n = 4, 31%). The structural validity and internal consistency were relatively well evaluated for most scales. However, evidence was limited for other measurement properties such as reliability, criterion validity, and construct validity. None of the scales reported all psychometric properties outlined by COSMIN. The New Short-Form of the Premenstrual Symptoms Questionnaire and the Japanese version of the Daily Record of Severity of Problems demonstrated sufficient structural validity and internal consistency, although the quality of evidence for other properties was indeterminate.ConclusionsAlthough some PROMs demonstrated promising psychometric properties, further validation studies are required for most scales. The development of innovative scales with robust measurement properties is essential for advancing the assessment of PMS/PMDD in Japanese clinical and research settings. Careful consideration of the characteristics of each PROM is necessary when selecting instruments for specific purposes.