Using Generative AI to Appraise the Quality of Medical Education Research Studies: Agreement Between AI-Generated and Human MERSQI Scores.
The increasing volume of medical education research necessitates efficient, reliable, and scalable methods for conducting quality appraisals. The Medical Education Research Study Quality Instrument (MERSQI) is a widely used tool, although its manual scoring process remains resource-intensive. This study evaluated how well large language models (LLMs) appraise medical education research using the MERSQI tool in comparison with human judges. Three LLMs (GPT-5, Claude Sonnet 4, and Gemini 2.5 Pro) assigned MERSQI domain scores to 1423 medical education research articles. The authors compared AI-generated scores with human-generated scores using intraclass correlation coefficients (ICCs) across the six MERSQI domains. They evaluated the agreement between AI- and human-generated MERSQI composite scores using Bland-Altman plots. Domain-level ICC values ranged from fair (0.24) to near perfect (0.81), with the lowest agreement observed in the 'sampling,' 'validity evidence,' and 'data analysis' domains. No single LLM consistently outperformed the others across all domains. Composite score agreement with human ratings was substantial and similar across LLMs (ICC range: 0.65-0.69). GPT-5 produced slightly lower composite scores than humans, while Claude Sonnet 4 and Gemini 2.5 Pro produced higher scores, with Gemini showing the largest deviation. The Bland-Altman plots for Gemini 2.5 Pro suggested proportional bias, indicating its agreement with human scores varied across the range of study quality. These LLMs demonstrated substantial agreement with human raters for MERSQI composite scores, but domain-level agreement varied. Systematic differences in scoring patterns highlight the need for human oversight and additional calibrations before integrating LLMs into systematic review appraisal workflows.
- # Medical Education Research Study Quality Instrument
- # Large Language Models
- # Medical Education Research
- # Education Research Study Quality Instrument
- # Medical Education Research Study Quality
- # Medical Education Research Studies
- # Intraclass Correlation Coefficients Range
- # Validity Evidence
- # Intraclass Correlation Coefficients
- # Data Analysis
- Research Article
9
- 10.1080/10872981.2024.2308359
- Jan 24, 2024
- Medical Education Online
Background: The medical education research study quality instrument (MERSQI) was designed to appraise medical education research quality based on study design criteria. As with many such tools, application of the results may have unintended consequences. This study applied the MERSQI to published medical education research identified in a bibliometric analysis. Methods: A bibliometric analysis identified highly cited articles in medical education that two authors independently evaluated using the MERSQI. After screening duplicate or non-research articles, the authors reviewed 21 articles with the quality instrument. Initially, five articles were reviewed independently and results were compared to ensure agreed upon understanding of the instrument items. The remainder of the articles were independently reviewed. Overall scores for the articles were analyzed with a paired samples t-test and individual item ratings were analyzed for inter-rater reliability. Results: There was a significant difference in mean MERSQI score between reviewers. Inter-rater reliability for MERSQI items labeled response rate, validity and outcomes were considered unacceptable. Conclusions: Based on these results there is evidence that MERSQI items can be significantly influenced by interpretation, which lead to a difference in scoring. The MERSQI is a useful guide for identifying research methodologies. However, it should not be used to make judgments on the overall quality of medical education research methodology in its current format. The authors make specific recommendations for how the instrument could be revised for greater clarity and accuracy.
- Research Article
2
- 10.1186/s12909-022-03301-1
- Apr 1, 2022
- BMC Medical Education
BackgroundAs a community of practice (CoP), medical education depends on its research literature to communicate new knowledge, examine alternative perspectives, and share methodological innovations. As a key route of communication, the medical education CoP must be concerned about the rigor and validity of its research literature, but prior studies have suggested the need to improve medical education research quality. Of concern in the present study is the question of how responsive the medical education research literature is to changes in the CoP. We examine the nature and extent of changes in the quality of medical education research over a decade, using a widely cited study of research quality in the medical education research literature as a benchmark to compare more recent quality indicators.MethodsA bibliometric analysis was conducted to examine the methodologic quality of quantitative medical education research studies published in 13 selected journals from September 2013 to December 2014. Quality scores were calculated for 482 medical education studies using a 10-item Medical Education Research Study Quality Instrument (MERSQI) that has demonstrated strong validity evidence. These data were compared with data from the original study for the same journals in the period September 2002 to December 2003. Eleven investigators representing 6 academic medical centers reviewed and scored the research studies that met inclusion and exclusion criteria. Primary outcome measures include MERSQI quality indicators for 6 domains: study design, sampling, type of data, validity, data analysis, and outcomes.ResultsThere were statistically significant improvements in four sub-domain measures: study design, type of data, validity and outcomes. There were no changes in sampling quality or the appropriateness of data analysis methods. There was a small but significant increase in the use of patient outcomes in these studies.ConclusionsOverall, we judge this as equivocal evidence for the responsiveness of the research literature to changes in the medical education CoP. This study identified areas of strength as well as opportunities for continued development of medical education research.
- Research Article
18
- 10.1186/s12909-017-1048-3
- Nov 9, 2017
- BMC Medical Education
BackgroundThere is little evidence regarding the comparative quality of abstracts and articles in medical education research. The Medical Education Research Study Quality Instrument (MERSQI), which was developed to evaluate the quality of reporting in medical education, has strong validity evidence for content, internal structure, and relationships to other variables. We used the MERSQI to compare the quality of reporting for conference abstracts, journal abstracts, and published articles.MethodsThis is a retrospective study of all 46 medical education research abstracts submitted to the Society of General Internal Medicine 2009 Annual Meeting that were subsequently published in a peer-reviewed journal. We compared MERSQI scores of the abstracts with scores for their corresponding published journal abstracts and articles. Comparisons were performed using the signed rank test.ResultsOverall MERSQI scores increased significantly for published articles compared with conference abstracts (11.33 vs 9.67; P < .001) and journal abstracts (11.33 vs 9.96; P < .001). Regarding MERSQI subscales, published articles had higher MERSQI scores than conference abstracts in the domains of sampling (1.59 vs 1.34; P = .006), data analysis (3.00 vs 2.43; P < .001), and validity of evaluation instrument (1.04 vs 0.28; P < .001). Published articles also had higher MERSQI scores than journal abstracts in the domains of data analysis (3.00 vs 2.70; P = .004) and validity of evaluation instrument (1.04 vs 0.26; P < .001).ConclusionsTo our knowledge, this is the first study to compare the quality of medical education abstracts and journal articles using the MERSQI. Overall, the quality of articles was greater than that of abstracts. However, there were no significant differences between abstracts and articles for the domains of study design and outcomes, which indicates that these MERSQI elements may be applicable to abstracts. Findings also suggest that abstract quality is generally preserved from original presentation to publication.
- Research Article
715
- 10.1097/acm.0000000000000786
- Aug 1, 2015
- Academic Medicine
The Medical Education Research Study Quality Instrument (MERSQI) and the Newcastle-Ottawa Scale-Education (NOS-E) were developed to appraise methodological quality in medical education research. The study objective was to evaluate the interrater reliability, normative scores, and between-instrument correlation for these two instruments. In 2014, the authors searched PubMed and Google for articles using the MERSQI or NOS-E. They obtained or extracted data for interrater reliability-using the intraclass correlation coefficient (ICC)-and normative scores. They calculated between-scale correlation using Spearman rho. Each instrument contains items concerning sampling, controlling for confounders, and integrity of outcomes. Interrater reliability for overall scores ranged from 0.68 to 0.95. Interrater reliability was "substantial" or better (ICC > 0.60) for nearly all domain-specific items on both instruments. Most instances of low interrater reliability were associated with restriction of range, and raw agreement was usually good. Across 26 studies evaluating published research, the median overall MERSQI score was 11.3 (range 8.9-15.1, of possible 18). Across six studies, the median overall NOS-E score was 3.22 (range 2.08-3.82, of possible 6). Overall MERSQI and NOS-E scores correlated reasonably well (rho 0.49-0.72). The MERSQI and NOS-E are useful, reliable, complementary tools for appraising methodological quality of medical education research. Interpretation and use of their scores should focus on item-specific codes rather than overall scores. Normative scores should be used for relative rather than absolute judgments because different research questions require different study designs.
- Research Article
- 10.1080/0142159x.2024.2385678
- Aug 3, 2024
- Medical Teacher
What is the educational challenge? The Medical Education Research Study Quality Instrument (MERSQI) is widely used to evaluate the quality of quantitative research in medical education. It has strong evidence of validity and is endorsed by guidelines. However, the manual appraisal process is time-consuming and resource-intensive, highlighting the need for more efficient methods. What are the proposed solutions? We propose to use ChatGPT to evaluate the quality of medical education research with the MERSQI and compare its scoring with those of human evaluators. What are the potential benefits to a broader global audience? Using ChatGPT to evaluate medical education research with the MERSQI can decrease the resources required for quality appraisal. This allows faster summaries of evidence, reducing the workload of researchers, editors, and educators. Furthermore, ChatGPTs’ capability to extract supporting excerpts provides transparency and may have the potential for data extraction and training new medical education researchers. What are the next steps? We plan to continue evaluating medical education research with ChatGPT using the MERSQI and other instruments to determine its feasibility in this realm. Moreover, we plan to investigate which types of studies ChatGPT performs best in.
- Research Article
890
- 10.1001/jama.298.9.1002
- Sep 5, 2007
- JAMA
Methodological shortcomings in medical education research are often attributed to insufficient funding, yet an association between funding and study quality has not been established. To develop and evaluate an instrument for measuring the quality of education research studies and to assess the relationship between funding and study quality. Internal consistency, interrater and intrarater reliability, and criterion validity were determined for a 10-item medical education research study quality instrument (MERSQI). This was applied to 210 medical education research studies published in 13 peer-reviewed journals between September 1, 2002, and December 31, 2003. The amount of funding obtained per study and the publication record of the first author were determined by survey. Study quality as measured by the MERSQI (potential maximum total score, 18; maximum domain score, 3), amount of funding per study, and previous publications by the first author. The mean MERSQI score was 9.95 (SD, 2.34; range, 5-16). Mean domain scores were highest for data analysis (2.58) and lowest for validity (0.69). Intraclass correlation coefficient ranges for interrater and intrarater reliability were 0.72 to 0.98 and 0.78 to 0.998, respectively. Total MERSQI scores were associated with expert quality ratings (Spearman rho, 0.73; 95% confidence interval [CI], 0.56-0.84; P < .001), 3-year citation rate (0.8 increase in score per 10 citations; 95% CI, 0.03-1.30; P = .003), and journal impact factor (1.0 increase in score per 6-unit increase in impact factor; 95% CI, 0.34-1.56; P = .003). In multivariate analysis, MERSQI scores were independently associated with study funding of $20 000 or more (0.95 increase in score; 95% CI, 0.22-1.86; P = .045) and previous medical education publications by the first author (1.07 increase in score per 20 publications; 95% CI, 0.15-2.23; P = .047). The quality of published medical education research is associated with study funding.
- Research Article
56
- 10.4300/jgme-d-11-00083.1
- Jun 1, 2011
- Journal of Graduate Medical Education
Past decisions about teaching often were based on the "PHOG" approach: "prejudices, hunches, opinions, and guesses."1 In the last decade, major advancements have occurred in the development and understanding of new evidence to guide medical education decisions. The formation of the Best Evidence in Medical Education (BEME) international working groups is an example of this new approach.2 BEME work groups systematically search for studies to answer key questions, with a rigorous approach to evaluating the quality of evidence. Other groups have examined the quality of methods and of reporting in English-language education research studies.3–10 Despite these developments, many decisions in medical education are still based on "persuasion and politics."11One of the primary goals of the Journal of Graduate Medical Education (JGME) is to improve the quality of graduate medical education research. Systematic reviews of education research have identified areas of concern.3–6,10 One of our strategies will be to improve readers' understanding of these areas. In each issue, the Journal plans to provide a summary about one aspect of research quality. For this issue, the subject is reliability and validity of assessment instruments used for research outcomes (pp 119–120). This topic is particularly relevant to our readers, as validity and reliability evidence for assessments is routinely underreported in manuscripts submitted to JGME.In this editorial, we introduce areas of concern in the quantitative methodologies delineated in systematic reviews of English-language publications and instruments available to examine the quality of education research. These instruments include the Medical Education Research Study Quality Instrument (MERSQI),5,7 the BEME global scale,3 and the Newcastle–Ottawa Scale (NOS) for assessing quality of nonrandomized studies.12 These instruments are based in part on Kirkpatrick's hierarchy of educational outcomes,3,13 which provides a valuable conceptual framework for planning and evaluating educational initiatives. Standards are also available to assess the quality of methods reporting,14,15 but they will not be discussed here. Similarly, other topics, such as quality of research questions and overinterpretation of results, will not be addressed in this editorial.16The NOS was developed to rate the quality of nonrandomized studies included in systematic reviews and has data to support its validity.12 Although the NOS was created for clinical research, it has been modified and used in systematic reviews of educational research.6,10 Examining one's own research for the presence or absence of specific items may be instructive (table 1).In 2 studies, the modified NOS was highly correlated with the MERSQI (table 2) and the BEME global rating scale (table 3).7,10 Of the 3 scales, the MERSQI may be most useful for researchers wishing to examine their work for methodologic rigor, as it includes a comprehensive list of review items and also has a growing body of validity evidence.5,7 Less evidence is available for the BEME global rating scale, which includes a modified version of Kirkpatrick's hierarchy17 of the outcomes of educational interventions. Kirkpatrick's hierarchy of levels is also included in the MERSQI scale, with higher points assigned to higher levels of outcomes. Kirkpatrick's hierarchy, also termed Kirkpatrick's pyramid (figure), is employed widely by education experts to characterize the level of outcomes in an educational intervention. Authors could enhance the quality of their papers by including a discussion of their work in relation to the BEME global or Kirkpatrick frameworks. To date, these discussions rarely occur in JGME submissions.The levels of Kirkpatrick's outcomes include (1) participation rates or learner satisfaction; (2) changes in attitudes, knowledge, and skills; (3) changes in behaviors; and (4) changes to the care system or patient outcomes. For example, in a study comparing an interactive web-based program with readings and lectures for teaching residents techniques for smoking cessation, potential outcomes could be classified as follows:In systematic reviews of education research, the majority of studies reported outcomes at Kirkpatrick levels 1 and 2.5,7 Although undoubtedly easier to study, achievement of outcomes at these levels may not translate into effective, sustained changes in behaviors or improved patient outcomes. In general, outcomes reported were more often subjective rather than objective. Of greater concern is that outcomes are entirely absent in many studies: 19% in one 2008 review.7 On average, the data analysis portion of reviewed papers received the highest quality ratings, while validation of assessment instruments received the lowest quality ratings.5,7,9,10Other areas of methodologic concern found in literature reviews include (1) predominance of single-site studies; (2) small studies that are underpowered to find a difference between intervention and comparison groups; (3) lack of a comparison group or lack of description of the intervention for the comparison group (eg, description of usual teaching); (4) inadequate description of multifactorial interventions; (5) overconfidence in randomization to eliminate the influence of confounding variables (ie, bias); and (6) overuse of the single-group pretest/posttest strategy to assess differences, with resulting potential overestimates of the magnitude of the effect of the intervention.16,18 In future issues of JGME, we will examine some of these issues in greater detail.JGME editors suggest that authors consider evaluating their planned and ongoing work with the above-described instruments, the MERSQI, NOS, and BEME global scale, and other quality scales developed for specific interventions, such as online teaching modules.19 In addition, authors should consider Kirkpatrick's hierarchy when formulating studies and considering outcome measures. These additional steps in reflection may produce a study and eventual manuscript that requires fewer revision cycles and is ultimately of greater value to consumers of medical education research.
- Research Article
1
- 10.1186/s12909-023-04383-1
- May 30, 2023
- BMC Medical Education
There are many parameters that could be used to evaluate the quality of scientific meetings such as publication rates of meeting abstracts as full-text articles after the meeting or scoring with validated quality scales/tools that evaluate individual papers, project proposals, or submitted abstracts. This study aimed to determine the full-text publication rates for abstracts presented at Turkish National Medical Education Congresses and Symposia and to assess the quality of given abstracts. s presented at national medical education congresses and symposia between 2010 and 2014 in Türkiye were evaluated. Initially, the abstracts were evaluated if they were published as full-text articles in international and national peer-reviewed journals following the meeting. Secondly, the quality of presented abstracts was assessed with the Medical Education Research Study Quality Instrument (MERSQI) scale. Overall publication rate for the abstracts was 11.3%. The publication rate of oral and poster presentations were 26.6% and 8.1%, respectively. Oral presentations had a statistically higher publication rate than poster presentations (p = .000). The mean MERSQI score for abstracts was 7.73 ± 2.59. The oral presentations had higher MERSQI mean scores than poster presentations (8.28 ± 2.46 vs. 7.61 ± 2.6; p = .032). Similarly, published abstracts had a significantly higher score compared to unpublished abstracts (10.07 ± 2.74 vs. 7.43 ± 2.41; p = .000). Interestingly, there was no statistical difference between the mean MERSQI scores of the published oral and poster presentations (9.33 ± 2.45 vs. 10.61 ± 2.72; p = .101). This study showed that the main factor for a meeting abstract to be published as a full-text article is the scientific quality of the study. The quality of presentations at annual medical education meetings in Türkiye were low compared with international meetings which did not improve over five years. An institutional policy that would set quality standards for medical education research and increase the awareness of researchers on the topic might help improve the design, execution, and reporting of such studies in Türkiye. The MERSQI could be a valuable tool to monitor the quality of submitted abstracts and to increase the awareness of novice researchers on high quality research.
- Research Article
259
- 10.1007/s11606-008-0664-3
- Jul 1, 2008
- Journal of General Internal Medicine
BackgroundDeficiencies in medical education research quality are widely acknowledged. Content, internal structure, and criterion validity evidence support the use of the Medical Education Research Study Quality Instrument (MERSQI) to measure education research quality, but predictive validity evidence has not been explored.ObjectiveTo describe the quality of manuscripts submitted to the 2008 Journal of General Internal Medicine (JGIM) medical education issue and determine whether MERSQI scores predict editorial decisions.Design and ParticipantsCross-sectional study of original, quantitative research studies submitted for publication.MeasurementsStudy quality measured by MERSQI scores (possible range 5–18).ResultsOf 131 submitted manuscripts, 100 met inclusion criteria. The mean (SD) total MERSQI score was 9.6 (2.6), range 5–15.5. Most studies used single-group cross-sectional (54%) or pre-post designs (32%), were conducted at one institution (78%), and reported satisfaction or opinion outcomes (56%). Few (36%) reported validity evidence for evaluation instruments. A one-point increase in MERSQI score was associated with editorial decisions to send manuscripts for peer review versus reject without review (OR 1.31, 95%CI 1.07–1.61, p = 0.009) and to invite revisions after review versus reject after review (OR 1.29, 95%CI 1.05–1.58, p = 0.02). MERSQI scores predicted final acceptance versus rejection (OR 1.32; 95% CI 1.10–1.58, p = 0.003). The mean total MERSQI score of accepted manuscripts was significantly higher than rejected manuscripts (10.7 [2.5] versus 9.0 [2.4], p = 0.003).ConclusionsMERSQI scores predicted editorial decisions and identified areas of methodological strengths and weaknesses in submitted manuscripts. Researchers, reviewers, and editors might use this instrument as a measure of methodological quality.Electronic supplementary materialThe online version of this article (doi:10.1007/s11606-008-0664-3) contains supplementary material, which is available to authorized users.
- Research Article
56
- 10.1186/s12909-023-04033-6
- Jan 25, 2023
- BMC Medical Education
BackgroundThe Medical Education Research Study Quality Instrument (MERSQI) is widely used to appraise the methodological quality of medical education studies. However, the MERSQI lacks some criteria which could facilitate better quality assessment. The objective of this study is to achieve consensus among experts on: (1) the MERSQI scoring system and the relative importance of each domain (2) modifications of the MERSQI.MethodA modified Delphi technique was used to achieve consensus among experts in the field of medical education. The initial item pool contained all items from MERSQI and items added in our previous published work. Each Delphi round comprised a questionnaire and, after the first iteration, an analysis and feedback report. We modified the quality instruments’ domains, items and sub-items and re-scored items/domains based on the Delphi panel feedback.ResultsA total of 12 experts agreed to participate and were sent the first and second-round questionnaires. First round: 12 returned of which 11 contained analysable responses; second-round: 10 returned analysable responses. We started with seven domains with an initial item pool of 12 items and 38 sub-items. No change in the number of domains or items resulted from the Delphi process; however, the number of sub-items increased from 38 to 43 across the two Delphi rounds. In Delphi-2: eight respondents gave ‘study design’ the highest weighting while ‘setting’ was given the lowest weighting by all respondents. There was no change in the domains’ average weighting score and ranks between rounds.ConclusionsThe final criteria list and the new domain weighting score of the Modified MERSQI (MMERSQI) was satisfactory to all respondents. We suggest that the MMERSQI, in building on the success of the MERSQI, may help further establish a reference standard of quality measures for many medical education studies.
- Research Article
2
- 10.1007/s00464-022-09104-1
- Feb 22, 2022
- Surgical endoscopy
Surgical endoscopy (SE), the official journal of the Society of American Gastrointestinal and Endoscopic Surgeons and the European Association for Endoscopic Surgery, is an important source of new evidence pertaining to surgical education in the field. However, qualitative deficiencies in medical education research have prompted medical education leaders to advocate for increased methodological rigor. The purpose of this study is to review the quality of education-focused research published through SE. A PubMed search examining all SE articles categorized as education-related research from 2010 to 2019 was conducted; studies not meeting inclusion criteria were excluded. Remaining publications were independently reviewed, classified, and scored by 7 raters using the medical education research study quality instrument (MERSQI). Intraclass correlation was calculated and data were examined with descriptive statistics. A total of 227 studies met inclusion criteria. There was no significant difference in number of publications by year (average 25.88 [SD 5.6]); 60% were conducted outside of the United States, and 47% (n = 106) were funded. The average MERSQI was 12.5 (SD 2). Most studies used two-group non-random (42%, n = 96) or post/cross-sectional designs (29%, n = 65). Thirty-six (16%) were randomized controlled trials. Multi-institutional studies comprised 24% (n = 54). Of the manuscripts, 96% (n = 217) reported at least one measure of validity evidence and 28% (n = 67) described three levels of validity evidence. Studies primarily reported changes in skills or knowledge (45%, n = 103) or satisfaction or general facts (44%, n = 99), while patient-related outcomes encompassed 3% (n = 6) of studies. ICC between raters was 0.93 (CI 0.90-0.93, p < 0.001). Based on publications to date, this journal's peer review process appears to facilitate the dissemination of education-related studies of moderate to good quality. However, there were uncovered deficits, ranging from validity evidence to study designs and level of outcomes. This journal's breadth of viewership offers a potential venue to advance education-related research.
- Conference Article
- 10.1370/afm.21.s1.4073
- Jan 1, 2023
<h3>Context:</h3> To augment our research curriculum, our three family medicine residency programs participated in the American Board of Family Medicine (ABFM) Journal Club Pilot. We implemented the Medical Education Research Study Quality Instrument (MERSQI) to score journal articles and identify potential curriculum gaps. <h3>Objective:</h3> 1) summarize the body of literature included in the ABFM Journal Club Pilot by scoring each article for methodological quality; 2) identify research curriculum strengths and areas for growth <h3>Study Design:</h3> A bibliometric analysis was conducted to examine the methodologic quality of research studies included in the ABFM Journal Club Pilot. Additionally, we surveyed residents to document their confidence critically appraising journal articles. <h3>Dataset:</h3> MERSQI quality scores were calculated for 40 studies published in 25 journals selected by the ABFM. <h3>Intervention/Instrument:</h3> The 10-item MERSQI was used to assess methodological quality across six domains: study design, sampling (number of institutions and response rate), type of data, validity (internal structure, content, and relationships to other variables), data analysis (appropriateness and complexity), and outcomes. Previous studies document strong validity evidence for the MERSQI. The resident survey included 12 Likert scale items measuring confidence appraising different elements of journal club articles (e.g. interpreting confidence intervals, statistical power, etc.). <h3>Results:</h3> MERSQI scores ranged from 13 to 18, with the average being 16.31 (higher scores indicate higher quality). A majority of articles (80%) implemented a randomized control trial. Most articles (82%) with a survey had a response rate of 75% or above. Most studies were multi-institutional (90%) and presented objective measurements (87.2%) as opposed to self-assessment data alone (12.8%). At baseline before implementing the journal club pilot, a majority of residents indicated they had none or minimal experience evaluating journal articles (n=22, 52.4%). <h3>Conclusions:</h3> On average, ABFM journal club articles had relatively high MERSQI scores compared to other bibliometric analyses. The MERSQI was a useful tool to identify gaps in our journal club curriculum. This information may guide the selection of future journal articles and refine the curriculum moving forward.
- Research Article
41
- 10.1007/s11606-015-3269-7
- Mar 27, 2015
- Journal of General Internal Medicine
Studies reveal that 44.5% of abstracts presented at national meetings are subsequently published in indexed journals, with lower rates for abstracts of medical education scholarship. We sought to determine whether the quality of medical education abstracts is associated with subsequent publication in indexed journals, and to compare the quality of medical education abstracts presented as scientific abstracts versus innovations in medical education (IME). Retrospective cohort study. Medical education abstracts presented at the Society of General Internal Medicine (SGIM) 2009 annual meeting. Publication rates were measured using database searches for full-text publications through December 2013. Quality was assessed using the validated Medical Education Research Study Quality Instrument (MERSQI). Overall, 64 (44%) medical education abstracts presented at the 2009 SGIM annual meeting were subsequently published in indexed medical journals. The MERSQI demonstrated good inter-rater reliability (intraclass correlation range, 0.77-1.00) for grading the quality of medical education abstracts. MERSQI scores were higher for published versus unpublished abstracts (9.59 vs. 8.81, p = 0.03). Abstracts with a MERSQI score of 10 or greater were more likely to be published (OR 3.18, 95% CI 1.47-6.89, p = 0.003). ). MERSQI scores were higher for scientific versus IME abstracts (9.88 vs. 8.31, p < 0.001). Publication rates were higher for scientific abstracts (42 [66%] vs. 37 [46%], p = 0.02) and oral presentations (15 [23%] vs. 6 [8%], p = 0.01). The publication rate of medical education abstracts presented at the 2009 SGIM annual meeting was similar to reported publication rates for biomedical research abstracts, but higher than publication rates reported for medical education abstracts. MERSQI scores were associated with higher abstract publication rates, suggesting that attention to measures of quality--such as sampling, instrument validity, and data analysis--may improve the likelihood that medical education abstracts will be published.
- Research Article
106
- 10.1016/j.surg.2015.03.007
- May 5, 2015
- Surgery
Systematic review of coaching to enhance surgeons' operative performance
- Research Article
7
- 10.1016/j.jsurg.2019.07.006
- Aug 2, 2019
- Journal of Surgical Education
Tracking Surgical Education Survey Research Through the APDS Listserv