Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Automating creativity assessment in engineering design: A psychometric validation of AI ‐generated items of the design problem task

  • TL;DR
  • Abstract
  • Literature Map
  • Similar Papers
TL;DR

This study evaluates the use of large language models for automatic generation of engineering design problem items, finding that AI-generated items achieved higher content validity than expert-created ones and demonstrated reliable psychometric properties, supporting scalable assessment of engineering creativity.

Abstract
Translate article icon Translate Article Star icon

Abstract Background Creativity is essential for engineering design, yet its assessment remains challenging due to the resource‐intensive nature of traditional evaluation methods. Purpose/Hypothesis(es) This study investigates the potential of automatic item generation (AIG) using large language models (LLMs) to create psychometrically sound items for the design problem task (DPT), which measures creative thinking in engineering. Design/Method We developed and validated engineering design problems across three domains: ability difference and limitations (e.g., assisting people with learning impairments), transportation and mobility (e.g., reducing traffic congestion in mega cities), and social environments and systems (e.g., improving access to clean water in remote areas). The study comprised three phases with samples matched on race and ethnicity: (1) content validation with a diverse sample of 40 engineers evaluating item clarity and validity; (2) item administration to 462 engineering students; and (3) response evaluation by 65 expert raters assessing originality and effectiveness. Results Results demonstrated that LLM‐generated items achieved comparable or higher content validity rates than expert‐written items (43% vs. 20% success). Bayesian confirmatory factor analysis supported a unidimensional model for fluency, originality, and effectiveness scores, with excellent reliability estimates (.92–.95). While fluency showed minimal correlation with originality ( r = −.11) and effectiveness ( r = −.04), originality and effectiveness were strongly positively correlated ( r = .73). Conclusions The present research advances our understanding of automated assessment generation in engineering education, provides empirical evidence for the psychometric properties of AI‐generated engineering creativity tasks, and offers a scalable approach for measuring creative thinking in engineering classrooms.

Similar Papers
  • Research Article
  • Cite Count Icon 12
  • 10.21449/ijate.1602294
A review of automatic item generation techniques leveraging large language models
  • Jun 1, 2025
  • International Journal of Assessment Tools in Education
  • Bin Tan + 4 more

This study reviews existing research on the use of large language models (LLMs) for automatic item generation (AIG). We performed a comprehensive literature search across seven research databases, selected studies based on predefined criteria, and summarized 60 relevant studies that employed LLMs in the AIG process. We identified the most commonly used LLMs in current AIG literature, their specific applications in the AIG process, and the characteristics of the generated items. We found that LLMs are flexible and effective in generating various types of items across different languages and subject domains. However, many studies have overlooked the quality of the generated items, indicating a lack of a solid educational foundation. Therefore, we share two suggestions to enhance the educational foundation for leveraging LLMs in AIG, advocating for interdisciplinary collaborations to exploit the utility and potential of LLMs.

  • Research Article
  • Cite Count Icon 16
  • 10.1016/j.tsc.2023.101364
Empowering students'engineering thinking: An empirical study of integrating engineering into science class at junior secondary schools
  • Jun 28, 2023
  • Thinking Skills and Creativity
  • Xiaohong Zhan + 4 more

Empowering students'engineering thinking: An empirical study of integrating engineering into science class at junior secondary schools

  • Research Article
  • Cite Count Icon 1
  • 10.1016/j.chbr.2026.100964
Automatic Item Generation for Personality Situational Judgment Tests with Large Language Models
  • Feb 1, 2026
  • Computers in Human Behavior Reports
  • Chang-Jin Li + 3 more

Automatic Item Generation for Personality Situational Judgment Tests with Large Language Models

  • Research Article
  • Cite Count Icon 4
  • 10.3390/educsci15081029
AI-Assisted Exam Variant Generation: A Human-in-the-Loop Framework for Automatic Item Creation
  • Aug 11, 2025
  • Education Sciences
  • Charles Macdonald Burke

Educational assessment relies on well-constructed test items to measure student learning accurately, yet traditional item development is time-consuming and demands specialized psychometric expertise. Automatic item generation (AIG) offers template-based scalability, and recent large language model (LLM) advances promise to democratize item creation. However, fully automated approaches risk introducing factual errors, bias, and uneven difficulty. To address these challenges, we propose and evaluate a hybrid human-in-the-loop (HITL) framework for AIG that combines psychometric rigor with the linguistic flexibility of LLMs. In a Spring 2025 case study at Franklin University Switzerland, the instructor collaborated with ChatGPT (o4-mini-high) to generate parallel exam variants for two undergraduate business courses: Quantitative Reasoning and Data Mining. The instructor began by defining “radical” and “incidental” parameters to guide the model. Through iterative cycles of prompt, review, and refinement, the instructor validated content accuracy, calibrated difficulty, and mitigated bias. All interactions (including prompt templates, AI outputs, and human edits) were systematically documented, creating a transparent audit trail. Our findings demonstrate that a HITL approach to AIG can produce diverse, psychometrically equivalent exam forms with reduced development time, while preserving item validity and fairness, and potentially reducing cheating. This offers a replicable pathway for harnessing LLMs in educational measurement without sacrificing quality, equity, or accountability.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 5
  • 10.1007/s10869-025-10067-y
AI-powered Automatic Item Generation for Psychological Tests: A Conceptual Framework for an LLM-based Multi-Agent AIG System
  • Aug 26, 2025
  • Journal of Business and Psychology
  • Philseok Lee + 2 more

Large Language Models (LLMs) are transforming industrial-organizational psychology and human resource management, with one of their most promising applications being automatic item generation (AIG) for psychological test development. Although recent advances in LLM-based AIG—particularly for non-cognitive assessments such as personality— show significant potential, ensuring rigorous quality control remains a persistent challenge. This study introduces a novel AIG framework, the LLM-based Multi-agent AIG system (LM-AIG), where each agent is responsible for different stages of item development, including item generation, content review, linguistic evaluation, bias assessment, and item revision. The LM-AIG also incorporates human feedback to enhance item quality. We implemented the LM-AIG framework using the open-source tool AutoGen to generate items assessing attitudes toward the use of AI in the workplace. To evaluate the quality of the generated items, we conducted an empirical study based on structured ratings from human raters, assessing construct relevance, linguistic clarity, appropriate language level, contextual specificity, and potential bias. This paper further discusses the role of human-in-the-loop mechanisms within the LM-AIG system and outlines future research directions.

  • Research Article
  • Cite Count Icon 3
  • 10.21831/reid.v10i2.76864
Automatic generation of physics items with Large Language Models (LLMs)
  • Oct 12, 2024
  • REID (Research and Evaluation in Education)
  • Moses Oluoke Omopekunola + 1 more

High-quality items are essential for producing reliable and valid assessments, offering valuable insights for decision-making processes. As the demand for items with strong psychometric properties increases for both summative and formative assessments, automatic item generation (AIG) has gained prominence. Research highlights the potential of large language models (LLMs) in the AIG process, noting the positive impact of generative AI tools like ChatGPT on educational assessments, recognized for their ability to generate various item types across different languages and subjects. This study fills a research gap by exploring how AI-generated items in secondary/high school physics aligned with educational taxonomy. It utilizes Bloom's taxonomy, a well-known framework for designing and categorizing assessment items across various cognitive levels, from low to high. It focuses on a preliminary assessment of LLMs ability to generate physics items that match the Bloom’s taxonomy application level. Two leading LLMs, ChatGPT (GPT-4) and Gemini, were chosen for their strong performance in creating high-quality educational content. The research utilized various prompts to generate items at different cognitive levels based on Bloom's taxonomy. These items were assessed using multiple criteria: clarity, accuracy, absence of misleading content, appropriate complexity, correct language use, alignment with the intended level of Bloom's taxonomy, solvability, and assurance of a single correct answer. The findings indicated that both ChatGPT and Gemini were skilled at generating physics assessment items, though their effectiveness varied based on the prompting methods used. Instructional prompts, particularly, resulted in excellent outputs from both models, producing items that were clear, precise, and consistently aligned with the Application level of Bloom's taxonomy.

  • Research Article
  • Cite Count Icon 2
  • 10.1111/jedm.12420
Algorithmic Bias in BERT for Response Accuracy Prediction: A Case Study for Investigating Population Validity
  • Oct 27, 2024
  • Journal of Educational Measurement
  • Guher Gorgun + 1 more

Pretrained large language models (LLMs) have gained popularity in recent years due to their high performance in various educational tasks such as learner modeling, automated scoring, automatic item generation, and prediction. Nevertheless, LLMs are black box approaches where models are less interpretable, and they may carry human biases and prejudices because historical human data have been used for pretraining these large‐scale models. For these reasons, the prediction tasks based on LLMs require scrutiny to ensure that the prediction models are fair and unbiased. In this study, we used BERT—a pretrained encoder‐only LLM for predicting response accuracy using action sequences extracted from the 2012 PIAAC assessment. We selected three countries (i.e., Finland, Slovakia, and the United States) representing different performance levels in the overall PIAAC assessment. We found promising results for predicting response accuracy using the fine‐tuned BERT model. Additionally, we examined algorithmic bias in the prediction models trained with different countries. We found differences in model performance, suggesting that some trained models are not free from bias, and thus the models are less generalizable across countries. Our results highlighted the importance of investigating algorithmic fairness in prediction models utilizing algorithmic systems to ensure models are bias‐free.

  • Research Article
  • Cite Count Icon 26
  • 10.1111/2041-210x.14325
Harnessing large language models for coding, teaching and inclusion to empower research in ecology and evolution
  • May 2, 2024
  • Methods in Ecology and Evolution
  • Natalie Cooper + 4 more

Large language models (LLMs) are a type of artificial intelligence (AI) that can perform various natural language processing tasks. The adoption of LLMs has become increasingly prominent in scientific writing and analyses because of the availability of free applications such as ChatGPT. This increased use of LLMs not only raises concerns about academic integrity but also presents opportunities for the research community. Here we focus on the opportunities for using LLMs for coding in ecology and evolution. We discuss how LLMs can be used to generate, explain, comment, translate, debug, optimise and test code. We also highlight the importance of writing effective prompts and carefully evaluating the outputs of LLMs. In addition, we draft a possible road map for using such models inclusively and with integrity. LLMs can accelerate the coding process, especially for unfamiliar tasks, and free up time for higher level tasks and creative thinking while increasing efficiency and creative output. LLMs also enhance inclusion by accommodating individuals without coding skills, with limited access to education in coding, or for whom English is not their primary written or spoken language. However, code generated by LLMs is of variable quality and has issues related to mathematics, logic, non‐reproducibility and intellectual property; it can also include mistakes and approximations, especially in novel methods. We highlight the benefits of using LLMs to teach and learn coding, and advocate for guiding students in the appropriate use of AI tools for coding. Despite the ability to assign many coding tasks to LLMs, we also reaffirm the continued importance of teaching coding skills for interpreting LLM‐generated code and to develop critical thinking skills. As editors of MEE, we support—to a limited extent—the transparent, accountable and acknowledged use of LLMs and other AI tools in publications. If LLMs or comparable AI tools (excluding commonly used aids like spell‐checkers, Grammarly and Writefull) are used to produce the work described in a manuscript, there must be a clear statement to that effect in its Methods section, and the corresponding or senior author must take responsibility for any code (or text) generated by the AI platform.

  • PDF Download Icon
  • Research Article
  • 10.37256/ser.5120243559
Exploring the Effectiveness of Reverse Engineering Pedagogy in a Culture-Oriented STEAM Course: from Heritage to Creativity
  • Mar 19, 2024
  • Social Education Research
  • Zhihua Lin + 2 more

This research innovatively combines culture-based Culture-STEAM (C-STEAM) with Reverse Engineering pedagogy, aiming to explore the effectiveness of C-STEAM Reverse Engineering Teaching Mode on students’ Creative Thinking, Cultural Competence and Engineering Thinking. The study was conducted in an STEAM course of 90 undergraduate students major in Educational Technology. Based on the observation of the pilot study (Study 1), we refined the research design and accessed its efficacy through quasi-experimental design in the formal study (Study2). The results indicate that the experimental group outperformed the control group in aspects of Creative Thinking. Such as the Rationality, Concreteness, Generate Precise Ideas and Improve Idea. Both the experimental and control groups exhibited significant improvement in terms of the Value, Accuracy, and Logicality of Creative Thinking, overall Cultural Competence, Cultural Understanding, Cultural Identity, and Engineering Thinking. However, the two groups of students did not show significant progress in Generate Diverse Ideas, Flexibility and Evaluate Idea within Creative Thinking, as well as Cultural Practice within Cultural Competence. The study underscores the value of the C-STEAM Reverse Engineering Teaching Mode, particularly in enhancing students’ Creative Thinking, Cultural Competence, and Engineering Thinking, which provides a reference example for teachers to cultivate students’ ability. This study makes up for the gaps in C-STEAM related teaching mode innovation and explores possible directions for the future development of STEAM education. Nevertheless, it necessitates further refinement exploration of its internal mechanism.

  • Supplementary Content
  • Cite Count Icon 6
  • 10.48550/arxiv.2306.10062
Revealing the structure of language model capabilities
  • Jun 14, 2023
  • arXiv (Cornell University)
  • Ryan Burnell + 3 more

Building a theoretical understanding of the capabilities of large language models (LLMs) is vital for our ability to predict and explain the behavior of these systems. Here, we investigate the structure of LLM capabilities by extracting latent capabilities from patterns of individual differences across a varied population of LLMs. Using a combination of Bayesian and frequentist factor analysis, we analyzed data from 29 different LLMs across 27 cognitive tasks. We found evidence that LLM capabilities are not monolithic. Instead, they are better explained by three well-delineated factors that represent reasoning, comprehension and core language modeling. Moreover, we found that these three factors can explain a high proportion of the variance in model performance. These results reveal a consistent structure in the capabilities of different LLMs and demonstrate the multifaceted nature of these capabilities. We also found that the three abilities show different relationships to model properties such as model size and instruction tuning. These patterns help refine our understanding of scaling laws and indicate that changes to a model that improve one ability might simultaneously impair others. Based on these findings, we suggest that benchmarks could be streamlined by focusing on tasks that tap into each broad model ability.

  • Research Article
  • Cite Count Icon 9
  • 10.3389/frsps.2025.1460277
Validating the use of large language models for psychological text classification
  • Feb 21, 2025
  • Frontiers in Social Psychology
  • Hannah L Bunt + 3 more

Large language models (LLMs) are being used to classify texts into categories informed by psychological theory (“psychological text classification”). However, the use of LLMs in psychological text classification requires validation, and it remains unclear exactly how psychologists should prompt and validate LLMs for this purpose. To address this gap, we examined the potential of using LLMs for psychological text classification, focusing on ways to ensure validity. We employed OpenAI's GPT-4o to classify (1) reported speech in online diaries, (2) other-initiations of conversational repair in Reddit dialogues, and (3) harm reported in healthcare complaints submitted to NHS hospitals and trusts. Employing a two-stage methodology, we developed and tested the validity of the prompts used to instruct GPT-4o using manually labeled data (N = 1,500 for each task). First, we iteratively developed three types of prompts using one-third of each manually coded dataset, examining their semantic validity, exploratory predictive validity, and content validity. Second, we performed a confirmatory predictive validity test on the final prompts using the remaining two-thirds of each dataset. Our findings contribute to the literature by demonstrating that LLMs can serve as valid coders of psychological phenomena in text, on the condition that researchers work with the LLM to secure semantic, predictive, and content validity. They also demonstrate the potential of using LLMs in rapid and cost-effective iterations over big qualitative datasets, enabling psychologists to explore and iteratively refine their concepts and operationalizations during manual coding and classifier development. Accordingly, as a secondary contribution, we demonstrate that LLMs enable an intellectual partnership with the researcher, defined by a synergistic and recursive text classification process where the LLM's generative nature facilitates validity checks. We argue that using LLMs for psychological text classification may signify a paradigm shift toward a novel, iterative approach that may improve the validity of psychological concepts and operationalizations.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 25
  • 10.1007/s11633-025-1546-4
Assessing and Understanding Creativity in Large Language Models
  • Apr 28, 2025
  • Machine Intelligence Research
  • Yunpu Zhao + 3 more

In the field of natural language processing, the rapid development of large language model (LLM) has attracted increasing attention. LLMs have shown a high level of creativity in various tasks, but the methods for assessing such creativity are inadequate. Assessment of LLM creativity needs to consider differences from humans, requiring multiple dimensional measurement while balancing accuracy and efficiency. This paper aims to establish an efficient framework for assessing the level of creativity in LLMs. By adapting the modified Torrance tests of creative thinking, the research evaluates the creative performance of various LLMs across 7 tasks, emphasizing 4 criteria including fluency, flexibility, originality, and elaboration. In this context, we develop a comprehensive dataset of 700 questions for testing and an LLM-based evaluation method. In addition, this study presents a novel analysis of LLMs’ responses to diverse prompts and role-play situations. We found that the creativity of LLMs primarily falls short in originality, while excelling in elaboration. In addition, the use of prompts and role-play settings of the model significantly influence creativity. Additionally, the experimental results also indicate that collaboration among multiple LLMs can enhance originality. Notably, our findings reveal a consensus between human evaluations and LLMs regarding the personality traits that influence creativity. The findings underscore the significant impact of LLM design on creativity and bridge artificial intelligence and human creativity, offering insights into LLMs’ creativity and potential applications.

  • Research Article
  • Cite Count Icon 1
  • 10.1177/00131644251355485
Human Expertise and Large Language Model Embeddings in the Content Validity Assessment of Personality Tests.
  • Aug 14, 2025
  • Educational and psychological measurement
  • Nicola Milano + 2 more

In this article, we explore the application of Large Language Models (LLMs) in assessing the content validity of psychometric instruments, focusing on the Big Five Questionnaire (BFQ) and Big Five Inventory (BFI). Content validity, a cornerstone of test construction, ensures that psychological measures adequately cover their intended constructs. Using both human expert evaluations and advanced LLMs, we compared the accuracy of semantic item-construct alignment. Graduate psychology students employed the Content Validity Ratio to rate test items, forming the human baseline. In parallel, state-of-the-art LLMs, including multilingual and fine-tuned models, analyzed item embeddings to predict construct mappings. The results reveal distinct strengths and limitations of human and AI approaches. Human validators excelled in aligning the behaviorally rich BFQ items, while LLMs performed better with the linguistically concise BFI items. Training strategies significantly influenced LLM performance, with models tailored for lexical relationships outperforming general-purpose LLMs. Here we highlight the complementary potential of hybrid validation systems that integrate human expertise and AI precision. The findings underscore the transformative role of LLMs in psychological assessment, paving the way for scalable, objective, and robust test development methodologies.

  • Conference Article
  • 10.54941/ahfe1006042
Leveraging LLMs to emulate the design processes of different cognitive styles
  • Jan 1, 2025
  • AHFE international
  • Xiyuan Zhang + 5 more

Cognitive styles, which shape designers’ thinking, problem-solving, and decision-making, influence strategies and preferences in design tasks. In team collaboration, diversity cognitive styles enhance problem-solving efficiency, foster creativity, and improve team performance.The ‘Co-evolution of problem–solution’ model serves as a key theoretical framework for understanding differences in designers’ cognitive styles. Based on this model, designers can be categorized into two cognitive styles: problem-driven and solution-driven. Problem-driven designers prioritize structuring the problem before developing solutions, while solution-driven designers generate solutions when design problems still ill-defined, and then work backward to define the problem. Designers with different expertise and disciplinary backgrounds exhibit distinct cognitive style tendencies. Different cognitive styles also adapt differently to design tasks, excelling in some more than others.As a rapidly advancing technology, large language models (LLMs) have shown considerable potential in the field of design. Their powerful generative capabilities position them as potential collaborators in design teams, emulating different cognitive styles. These emulations aim to bridge cognitive differences among team members, enable designers to leverage their individual strengths, and ultimately produce more feasible and high-quality design solutions.However, previous studies have been limited in leveraging LLMs to directly generate design outcomes based on different cognitive styles, neglecting the emulation of the design process itself. In fact, the evolutionary development between problem and solution spaces better reflects the core differences in cognitive styles. Moreover, communication and collaboration within design teams extend beyond simply exchanging solutions, but span multiple stages of the design process—from problem analysis, idea generation, to evaluation. To better integrate LLMs into design teams, it is necessary to consider the emulation of the design cognition process.To this end, our study, based on the cognitive style taxonomy proposed by Dorst and Cross (2001), explores how LLMs can be used to emulate the design processes of problem-driven and solution-driven designers. We develop a zero-shot chain-of-thought (CoT)-based prompting strategy that enables LLMs to emulate the step-by-step cognitive flow of both design styles. The prompt design is inspired by Jiang et al. (2014) and Chen et al. (2023), who analyzed cognitive differences in conceptual design process using the FBS ontology model. Furthermore, to evaluate the effectiveness of LLMs in emulating cognitive styles, this study establishes a three-dimentional evaluation metrics: static distribution (the proportion and preference of cognitive issues), dynamic transformation (behavioral transition patterns), and the creativity of the design outcomes. Using previous studies identified human design behaviours as a benchmark, we compare the cognitive styles emulated by LLMs under different design constraints against human performance to assess their alignment and differences.The results show that LLM-generated design processes align well with human cognitive styles, effectively emulate static cognitive characteristics. Moreover, enhancing novelty and integrity in solutions and demonstrating superior creativity compared to baseline methods. However, LLMs lack the fully complex nonlinear transitions between problem and solution spaces observed in human designers.This process-based emulation has the potential to enhance the application of LLMs in design teams, enabling them to not only serve as tools for generating solutions but also provide support for collaboration during key stages of the design process. Future research should enhance LLMs' reasoning flexibility through fine-tuning or the GoT approach and explore their impact on human-AI collaboration across diverse design tasks to refine their role in design teams.

  • Research Article
  • Cite Count Icon 3
  • 10.4274/dir.2025.253407
Artificial intelligence in radiology examinations: a psychometric comparison of question generation methods.
  • Jul 21, 2025
  • Diagnostic and interventional radiology (Ankara, Turkey)
  • Emre Emekli + 1 more

This study aimed to evaluate the usability of artificial intelligence (AI)-based question generation methods-Chat Generative Pre-trained Transformer (ChatGPT)-4o (a non-template-based large language model) and a template-based automatic item generation (AIG) method-in the context of radiology education. The primary objective was to compare the psychometric properties, perceived quality, and educational applicability of generated multiple-choice questions (MCQs) with those written by a faculty member. Fifth-year medical students who participated in the radiology clerkship at Eskişehir Osmangazi University were invited to take a voluntary 15-question examination covering musculoskeletal and rheumatologic imaging. The examination included five MCQs from each of three sources: a radiologist educator, ChatGPT-4o, and the template-based AIG method. Student responses were evaluated in terms of difficulty and discrimination indices. Following the examination, students rated each question using a Likert scale based on clarity, difficulty, plausibility of distractors, and alignment with learning goals. Correlations between students' examination performance and their theoretical/practical radiology grades were analyzed using Pearson's correlation method. A total of 115 students participated. Faculty-written questions had the highest mean correct response rate (2.91 ± 1.34), followed by template-based AIG (2.32 ± 1.66) and ChatGPT-4o (2.3 ± 1.14) questions (P < 0.001). The mean difficulty index was 0.58 for faculty, and 0.46 for both template- based AIG and ChatGPT-4o. Discrimination indices were acceptable (≥0.2) or very good (≥0.4) for template-based AIG questions. In contrast, four of the ChatGPT-generated questions were acceptable, and three were very good. Student evaluations of questions and the overall examination were favorable, particularly regarding question clarity and content alignment. Examination scores showed a weak correlation with practical examination performance (P = 0.041), but not with theoretical grades (P = 0.652). Both the ChatGPT-4o and template-based AIG methods produced MCQs with acceptable psychometric properties. While faculty-written questions were most effective overall, AI-generated questions- especially those from the template-based AIG method-showed strong potential for use in radiology education. However, the small number of items per method and the single-institution context limit the robustness and generalizability of the findings. These results should be regarded as exploratory, and further validation in larger, multicenter studies is required. AI-based question generation may potentially support educators by enhancing efficiency and consistency in assessment item creation. These methods may complement traditional approaches to help scale up high-quality MCQ development in medical education, particularly in resource-limited settings; however, they should be applied with caution and expert oversight until further evidence is available, especially given the preliminary nature of the current findings.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant