Articles published on Journal Of The American Medical Association
Authors
Select Authors
Journals
Select Journals
Duration
Select Duration
2121 Search results
Sort by Recency
- New
- Research Article
- 10.1016/j.amjmed.2026.02.048
- Jul 1, 2026
- The American journal of medicine
- Carl A Andersen + 6 more
Update in Outpatient General Internal Medicine: Practice-Changing Evidence Published in 2025.
- New
- Research Article
- 10.1186/s12894-026-02234-x
- Jun 29, 2026
- BMC urology
- Chengyang Tian + 4 more
Benign prostatic hyperplasia (BPH) is a common urological condition among older men and can cause lower urinary tract symptoms (LUTS). As patients increasingly seek disease-related information through social media, evaluations of the quality, reliability, and transparency of patient-facing BPH content remain limited. Repeated cross-sectional searches were conducted on Bilibili and TikTok on December 1, 16, and 31, 2025. Basic video characteristics were extracted. Video reliability, quality, and transparency were assessed using the modified DISCERN (mDISCERN), Global Quality Score (GQS), and the Journal of the American Medical Association (JAMA) benchmark criteria, respectively. Exploratory supplementary assessments of content coverage and accessibility-related features were conducted using the content coverage checklist (CCC) and video accessibility checklist (VAC). Potentially misleading or harmful information signals were also descriptively assessed. Spearman correlation analysis and ordinal logistic regression were performed to examine associations between video characteristics and assessment scores. A total of 217 videos were included. Overall, the median (IQR) scores were 3.00 (2.00-3.00) for mDISCERN, 3.00 (2.00-3.00) for GQS, and 2.00 (2.00-2.00) for JAMA. In exploratory supplementary analyses, the median (IQR) scores were 3.00 (2.00-4.00) for CCC and 6.00 (5.00-6.00) for VAC. Potentially misleading or harmful information signals were observed in 29 videos (13.4%). Videos published by professional individuals and professional institutions had significantly higher mDISCERN, GQS, and JAMA scores than those published by non-professional individuals (all P < 0.001). Spearman correlation analysis and ordinal logistic regression indicated that engagement metrics were not independent predictors of assessment scores. Overall, BPH-related videos on social media showed moderate reliability, quality, and transparency, but consistently high-value content remained limited. A minority of videos also contained potentially misleading or harmful information signals. Videos from professional sources tended to provide higher-quality information, whereas user engagement was not a reliable indicator of content quality. These findings suggest that urologists and healthcare institutions may have an important role in providing more structured, understandable, and patient-oriented digital education for men seeking BPH information online.
- New
- Research Article
- 10.1186/s41182-026-01006-5
- Jun 24, 2026
- Tropical medicine and health
- Jiaye Du + 4 more
Snakebite envenomation is a severe medical emergency and major global health threat. As short-video platforms increasingly shape public access to health information, accurate and reliable digital education is essential for snakebite prevention and first-aid decision-making. Although the quality of medical videos on short-video platforms has been examined for several diseases, little is known about the quality, reliability, and content characteristics of snakebite envenomation-related videos on major Chinese platforms. This study aimed to systematically assess the quality, reliability, and content characteristics of snakebite envenomation-related videos on TikTok (Douyin) and Bilibili, and to explore strategies for improving the quality of online health information. On February 17, 2026, we searched TikTok (Douyin) and Bilibili using the keywords "snakebite envenomation" and "snake envenoming". After videos with content irrelevant to the study keywords, commercial advertisements, duplicate videos, and videos uploaded within the preceding week were excluded according to predefined criteria, 220 eligible videos were included in the final analysis. For each video, we extracted duration, audience engagement metrics, uploader type, and other basic characteristics. Video quality was independently evaluated in a double-blind manner using the Global Quality Score (GQS), modified DISCERN (mDISCERN), and Journal of the American Medical Association (JAMA) benchmarks. Group differences were analyzed using Mann-Whitney U and Kruskal-Wallis tests, and correlations among variables were assessed using Spearman's rank correlation. Of the 300 initially screened videos, 220 were included in the final analysis, comprising 114(51.82%) videos from TikTok (Douyin) and 106(48.18%) from Bilibili. Most videos were presented in Mandarin Chinese. Etiology, clinical manifestations, treatment, and diagnosis were frequently covered, whereas epidemiology and prevention were less commonly addressed. Bilibili videos were significantly longer than TikTok (Douyin) videos 311.00 (150.75, 514.25) vs 101.00 (58.50, 171.00) seconds (P < 0.001), whereas TikTok (Douyin) videos had significantly more shares 59.00 (10.25, 717.75) vs 34.00 (3.00, 282.25) (P = 0.049). The overall median GQS, mDISCERN, and JAMA scores were 3.00 (3.00, 4.00), 3.00 (3.00, 3.00), and 1.00 (1.00, 2.00), respectively. No significant platform-based differences were observed in GQS or mDISCERN scores, but JAMA scores differed significantly between platforms (P = 0.020). Videos uploaded by healthcare professionals and nonprofit organizations had significantly higher GQS scores than those uploaded by individual users 4.00 (4.00, 5.00) and 4.00 (3.00, 4.00) vs 3.00 (2.00, 3.00), respectively; (P < 0.001), as well as higher JAMA scores both 2.00 (2.00, 2.00) vs 1.00 (1.00, 1.00) (P < 0.001). Engagement metrics were strongly interrelated (all P < 0.001), but showed limited associations with quality scores. GQS and JAMA correlated only with shares (P = 0.015 and P = 0.045, respectively), whereas mDISCERN was positively correlated with likes, collections, comments, and shares (P = 0.042, P = 0.025, P = 0.011, and P = 0.004, respectively). Snakebite envenomation-related videos on TikTok (Douyin) and Bilibili provided moderately useful health information, but showed important limitations in preventive content coverage and information transparency. Videos uploaded by healthcare professionals and nonprofit organizations were generally of higher quality and transparency than those uploaded by individual users, while audience engagement was not consistently aligned with content quality. These findings highlight the need for greater professional involvement, clearer source disclosure, and more standardized prevention- and first-aid-oriented content to improve the reliability and public health value of snakebite-related information on short-video platforms.
- New
- Research Article
- 10.1245/s10434-026-19996-1
- Jun 22, 2026
- Annals of surgical oncology
- Sedat Carkit + 1 more
Esophageal cancer is associated with substantial morbidity and mortality worldwide and is frequently diagnosed at advanced stages, leading patients and their relatives to seek health-related information beyond traditional clinical encounters. In recent years, YouTube has become a popular source of medical information. Nevertheless, questions persist regarding the accuracy, credibility, and overall reliability of the content available on the platform. This cross-sectional study evaluated publicly available YouTube videos related to esophageal cancer. Data were collected on December 3, 2025, using a browser without a user login to minimize algorithm-driven bias. Viewer engagement metrics (views, likes, and comments), source categories, and country of origin were recorded for each video. Content quality and reliability were assessed using the DISCERN instrument, Journal of the American Medical Association (JAMA) benchmark criteria, and Global Quality Score (GQS). Non-parametric statistical analyses were used to compare quality outcomes across source categories and evaluate the correlations between the engagement metrics and quality scores. A total of 78 videos met the inclusion criteria, most of which originated in the USA (83.3%). Health-related channels constituted the largest source category (35.9%), followed by patient experience-based videos (23.1%), and private institutions (20.5%). Viewer engagement metrics (views, likes, and comments) did not differ significantly among source types (p > 0.05). In contrast, the content quality varied substantially. Videos produced by public institutions achieved the highest DISCERN, JAMA, and GQS values, whereas patient-experience-based videos demonstrated significantly lower quality and reliability (p < 0.001). Engagement metrics were strongly intercorrelated but showed no association with quality scores. YouTube videos related to esophageal cancer frequently exhibit moderate informational quality, and popularity metrics do not reflect content reliability. Source credibility plays a critical role in determining video quality, underscoring the need for greater involvement of healthcare professionals and public institutions in digital health content production.
- New
- Research Article
- 10.1097/md.0000000000049361
- Jun 19, 2026
- Medicine
- Honggang Li + 2 more
TikTok and similar short-form video services are now widely adopted as prominent platforms for circulating health knowledge. However, the quality of content on keloid remains unclear. Keloid, a pathological scar with significant physical and psychological impact, necessitates accurate public education. This cross-sectional study analyzed 122 keloid-related TikTok videos on February 9, 2026, using an unlogged search to minimize bias. Video quality was assessed with 3 scoring systems: the global quality score, the modified DISCERN, and the benchmark criteria of the Journal of the American Medical Association, and uploader categories, content themes, and engagement metrics were analyzed. Videos featured a median length of 65.5 seconds and strong user engagement, but holistic quality was modest (median global quality score = 3.0, modified DISCERN = 2.0, Journal of the American Medical Association = 3.0). Content predominantly covered treatment (86.1%) and clinical manifestations (52.5%), whereas etiology, diagnosis, and recurrence were underrepresented. Videos from plastic surgeons and healthcare professionals had significantly higher quality scores than those from individual users (P < .05). There was no relationship found between engagement metrics and quality. In conclusion, keloid-related TikTok videos achieve wide reach but have limited informational quality, emphasizing the need for enhanced professional involvement and more comprehensive content to improve educational value.
- New
- Research Article
- 10.1097/md.0000000000049362
- Jun 19, 2026
- Medicine
- Ziwen Wang + 1 more
With the rapid growth of short video platforms, TikTok and Bilibili are becoming important channels for the public to access health information. Given that Asia is underrepresented in English-language medical literature, this study focuses on Mandarin Chinese videos to provide region-specific insights. This study aims to evaluate the content, quality, reliability, and transparency of sarcopenia-related videos on TikTok and Bilibili. Using the keyword “sarcopenia,” a search was conducted on TikTok and Bilibili to retrieve the top 150 videos based on the comprehensive ranking. Video duration, engagement metrics, uploader type, and content information were extracted. Videos were evaluated using the Global Quality Score (GQS), modified DISCERN (mDISCERN), Journal of the American Medical Association (JAMA) benchmark criteria, and content completeness scores. Group comparisons were performed using Mann–Whitney U and Kruskal–Wallis tests, and correlation analysis was performed using Spearman correlation. A total of 188 videos were included. The content primarily focused on symptoms and treatment, with less coverage on diagnosis and prognosis. The median GQS score was 2.00 (interquartile range [IQR]: 2.00–3.00), the median mDISCERN score was 2.00 (IQR: 1.75–3.00), the median JAMA score was 1.00 (IQR: 1.00–1.00), and the median content completeness score was 6.00 (IQR: 3.00–9.00). Compared with nonprofessional individuals and nonprofessional organizations, videos uploaded by professional individuals achieved higher scores in GQS, mDISCERN, JAMA, and content completeness (P < .05). No significant correlations were found between engagement metrics and mDISCERN scores (P > .05). The overall quality and reliability of sarcopenia-related videos on TikTok and Bilibili are suboptimal. Videos uploaded by professional individuals demonstrated higher quality and reliability. Future efforts should strengthen platform content oversight, encourage more professional individuals to contribute to health education, and promote the dissemination of high-quality videos.
- New
- Research Article
- 10.1097/md.0000000000049340
- Jun 19, 2026
- Medicine
- Feng Cai + 2 more
Temporomandibular disorders (TMD) are common conditions that may substantially impair quality of life, yet the quality of health information available on short-video platforms remains unclear. This study evaluated the coverage of TMD-related content, the educational quality, and the reliability of TMD-related Chinese videos on TikTok and Bilibili. On October 5, 2025, TikTok and Bilibili were searched using the Chinese keyword “颞下颌关节紊乱病.” After screening, 200 eligible videos were included, comprising 100 from TikTok and 100 from Bilibili. Video characteristics, uploader categories, engagement indicators, and content coverage were recorded. Video quality and reliability were assessed using the Global Quality Scale (GQS), modified DISCERN (mDISCERN), and Journal of the American Medical Association (JAMA) benchmark criteria. The median video duration was 136.50 seconds (Q1, 66.00; Q3, 235.25). The median GQS, mDISCERN, and JAMA scores were 3.00 (Q1, 2.00; Q3, 3.00), 3.00 (Q1, 2.00; Q3, 4.00), and 2.00 (Q1, 2.00; Q3, 3.00), respectively, indicating overall moderate quality and reliability. Diagnosis, symptoms, and treatment were the most frequently covered domains, whereas epidemiology was the least well covered; only 9.0% of videos provided a complete explanation of epidemiology, and 70.0% did not mention it at all. Bilibili videos were significantly longer than TikTok videos (median, 164.00 vs 81.50 seconds, P < .001), whereas TikTok videos showed significantly higher engagement in likes, comments, shares, and saves (all P < .001). No significant differences were observed between the 2 platforms in GQS, mDISCERN, or JAMA scores. Videos uploaded by specialized healthcare professionals had significantly higher GQS, mDISCERN, and JAMA scores than videos uploaded by other sources (all P < .001). Engagement indicators were strongly correlated with one another but did not reflect better informational quality. TMD-related videos on TikTok and Bilibili attracted substantial public attention; however, their overall educational quality and reliability were only moderate. Greater involvement of specialized healthcare professionals, clearer source disclosure, and more balanced topic coverage may help improve the quality of TMD-related health communication on short-video platforms.
- New
- Research Article
- 10.1038/s41598-026-57615-x
- Jun 18, 2026
- Scientific reports
- Jiajun Liu + 7 more
Short-form video platforms are increasingly used to obtain information about chronic obstructive pulmonary disease (COPD), but the quality and reliability of COPD-related content across Chinese platforms remain unclear. We evaluated 228 COPD-related videos from Douyin, Kwai, and Bilibili using the Journal of the American Medical Association (JAMA) benchmark criteria, Global Quality Scale, modified DISCERN (mDISCERN), and Patient Education Materials Assessment Tool. Video characteristics, creator identity, verification status, content theme, presentation format, and visible engagement metrics were also analyzed. Video quality differed significantly across platforms. Douyin showed the highest transparency, reliability, and overall quality scores. Kwai showed the highest visible engagement but had lower overall quality and actionability scores. Bilibili had the longest videos and the highest understandability scores. Organization verified accounts generally achieved higher quality scores than individual verified or unverified accounts, although this finding should be interpreted cautiously because of the small number of such accounts. Visible engagement metrics, including likes, comments, saves, and shares, were not significantly correlated with medical quality scores. Together, these findings suggest that visible popularity is a poor proxy for medical quality. COPD-related digital health communication may benefit from more accessible evidence-based content, clearer source identification, and closer oversight of high-risk therapeutic claims.
- New
- Research Article
- 10.2196/93054
- Jun 18, 2026
- JMIR Medical Informatics
- Fulin Pu + 3 more
BackgroundAlthough large language models (LLMs) show potential for patient education, their accuracy, usability, and comprehensibility lack validation in high-risk pediatric anesthesia. Rigorous evaluation is therefore essential prior to widespread clinical use in perioperative parental anesthesia education.ObjectiveThis study aims to evaluate the accuracy, reliability, and readability of responses generated by 5 LLMs to parental inquiries regarding pediatric anesthesia, and to assess their suitability for clinical use in perioperative caregiver education.MethodsTwo expert anesthesiologists identified 33 parental questions on pediatric anesthesia by screening authoritative resources and Google Trends. On December 14, 2025, these questions were submitted to 5 LLMs (DeepSeek-V3.2, ChatGPT-5, Gemini 2.5 Flash, Copilot, and Perplexity) via official web interfaces with default settings and zero-shot prompting, with each query in a separate conversation. Responses were standardized for blinded assessment. Two pediatric anesthesiologists with ≥10 years of clinical experience independently evaluated accuracy and reliability using the 4-point Likert accuracy scale, DISCERN, Ensuring Quality Information for Patients (EQIP), Journal of the American Medical Association (JAMA) benchmark, and Global Quality Score (GQS). After text preprocessing, readability was evaluated using 6 algorithms (Automated Readability Index [ARI], Flesch Reading Ease Score [FRES], Gunning Fog Index [GFI], Flesch-Kincaid Grade Level [FKGL], Coleman-Liau Index [CL], and the Simple Measure of Gobbledygook [SMOG]) via an online calculator. Interrater reliability was analyzed using the intraclass correlation coefficient (ICC); differences across models were assessed with the Kruskal-Wallis H test; and deviations from the sixth-grade benchmark were evaluated using 1-sample Wilcoxon signed-rank tests (P<.05 considered significant).ResultsAll 5 LLMs demonstrated high clinical accuracy (>90%; P=.12), with Gemini reaching 100%. Nevertheless, safety risks and content hallucinations were still observed. Excluding Gemini and Copilot, the remaining 3 models (ChatGPT, DeepSeek, and Perplexity) each produced unsafe content in 3.03% (n=1) of the 33 queries. Hallucinations were detected in all models except Gemini, with DeepSeek and Perplexity showing the highest hallucination rate (3/33, 9.09%). Furthermore, Perplexity showed superior reliability on DISCERN (median 41; P<.05), yet no model achieved a “good” rating. Gemini achieved the highest EQIP (median 66.67%; P<.05) despite lower GQS (median 3). Transparency was universally poor (JAMA median ≤1), with DeepSeek and ChatGPT showing a “floor effect.” ChatGPT had superior readability, but all models exceeded the recommended 6-grade complexity level.ConclusionsIn this study, 5 LLMs generally provided clinically accurate information when responding to parental questions about pediatric anesthesia. However, limitations were also identified, including hallucinated content, safety-related deficiencies, limited source transparency, and readability levels exceeding recommended standards. Therefore, LLM-generated information should be interpreted with caution and should not replace clinician guidance.
- New
- Research Article
- 10.1371/journal.pone.0350430
- Jun 17, 2026
- PLOS One
- Sankirth Madabhushi + 5 more
BackgroundThere is a concerning trend of misinformation of healthcare related content on social media. Recent studies have examined themes and narratives about Crohn’s disease but have not quantitatively assessed the accuracy and quality of content on Instagram Reels. Our aim was to assess the quality and accuracy of Instagram Reels about Crohn’s disease and examine differences in content by type of creator, from medical professionals to lay individuals.MethodsSeventy-eight top-viewed English-language Instagram Reels tagged with “#crohns” were evaluated. Videos were categorized by creator and content type. Two reviewers evaluated each video for accuracy and quality using an adapted harm/benefit score and the Journal of the American Medical Association (JAMA) benchmark criteria, respectively.ResultsSeventeen percent of videos were created by medical professionals and 83% by non-medical users. Educational content was significantly more common among medical professionals than other content creators (62% vs 23%; P = 0.005). No significant correlation was found between engagement metrics and either JAMA or harm/benefit scores. Medical professionals had significantly higher JAMA scores than non-medical users (2.5 vs 2, P < 0.001), but there was no significant difference in harm/benefit scores between groups (0 vs 0, P = 0.9601). Videos offering medical advice had the lowest median harm/benefit score (−1), with frequent misinformation noted. Forty-two percent of harmful videos were created by medical professionals.ConclusionsThe average Instagram Reel about Crohn’s disease was of moderate quality and neutral impact. Accuracy or quality was unrelated to video popularity. While videos by medical professionals had higher JAMA scores, this did not correspond to greater accuracy. Medical advice videos by medical professionals were not more accurate than those by non-medical creators, and multiple harmful videos were created by medical professionals, underscoring the need for critical evaluation of Crohn’s disease-related social media content.
- New
- Research Article
- 10.1055/s-0046-1820460
- Jun 16, 2026
- Revista Brasileira de Ortopedia
- Vinicius Borges Alencar + 5 more
ObjectiveArtificial intelligence (AI) tools based on natural language, such as ChatGPT 4.1 mini (OpenAI Group PBC) and Gemini 2.5 Flash (Alphabet Inc.), are used by patients as sources of medical information. The current study aimed to evaluate and compare the quality and readability of responses provided by these AIs, in Brazilian Portuguese, regarding rotator cuff surgery.MethodsThe present cross-sectional, descriptive, and comparative study followed qualitative and quantitative approaches. A total of 24 frequently-asked patient questions were used, classified according to Rothwell. Each question was entered individually into both platforms, and only the first response was considered. The quality assessment used the DISCERN instrument, developed by the University of Oxford and the British Library, and the Journal of the American Medical Association (JAMA) benchmark criteria. Readability was estimated using Análise de Legibilidade Textual (ALT, “Text Readibility Anallysis”, in Portuguese) software, validated for Brazilian Portuguese. The statistical analyses included the Wilcoxon and Friedman tests, repeated-measures analysis of variance (ANOVA), and the Conover post-hoc test with Bonferroni correction.ResultsChatGPT achieved a mean DISCERN score of 58.7 ± 4.0, and Gemini, 56.3 ± 3.5, with no significant difference (p = 0.174), but with a maximum effect size (rank-biserial correlation [rrb] = 1.0). Both models showed a mean readability corresponding to 13.3 years of schooling (p = 1.000). No response met the JAMA benchmark criteria. Value-based questions achieved the highest quality scores, whereas policy-related questions were the most complex in terms of readability. The correlation between quality and readability was moderate (ρ = 0.73;p = 0.099).ConclusionChatGPT 4.1 mini and Gemini 2.5 Flash do not yet provide adequate medical information in Brazilian Portuguese regarding editorial reliability, quality, and textual accessibility for the general public.
- New
- Research Article
- 10.1055/s-0046-1820461
- Jun 16, 2026
- Revista Brasileira de Ortopedia
- Vinicius Borges Alencar + 5 more
ResumoObjetivoFerramentas de inteligência artificial (IA) baseadas em linguagem natural, como ChatGPT-4.1 mini (OpenAI Group PBC) e Gemini 2.5 Flash (Alphabet Inc.), são utilizadas por pacientes como fonte de informação médica. Este estudo avaliou e comparou a qualidade e a legibilidade das respostas fornecidas por essas IAs, em português brasileiro, sobre cirurgia do manguito rotador.MétodosEstudo transversal, descritivo e comparativo, com abordagem qualiquantitativa. Foram utilizadas 24 perguntas frequentes de pacientes, classificadas segundo Rothwell. Cada pergunta foi inserida individualmente nas plataformas dos dois modelos, sendo considerada apenas a primeira resposta. A qualidade foi avaliada por meio do instrumento DISCERN, desenvolvido pela University of Oxford e pela British Library, e dos critérios editoriais da Journal of the American Medical Association (JAMA). A legibilidade foi estimada com o programa Análise de Legibilidade Textual (ALT), validado para o português brasileiro. As análises estatísticas incluíram os testes de Wilcoxon, Friedman, análise de variância ([analysis of variance, ANOVA, em inglês] para medidas repetidas) epost hocde Conover com correção de Bonferroni.ResultadosO ChatGPT obteve escore médio DISCERN de 58,7 ± 4,0, e o Gemini, 56,3 ± 3,5, sem diferença significativa (p = 0,174), mas com efeito máximo (rank-biserial correlation[rrb, em inglês] = 1,0). Ambos os modelos apresentaram legibilidade média correspondente a 13,3 anos de escolaridade (p = 1,000). Nenhuma resposta atendeu aos critérios editoriais da JAMA. Perguntas relacionadas a valores obtiveram os maiores escores de qualidade, ao passo que as perguntas sobre política foram as mais complexas em termos de leitura. A correlação entre qualidade e legibilidade foi moderada (ρ = 0,73;p = 0,099).ConclusãoChatGPT–4.1 mini e Gemini 2.5 Flash ainda não oferecem informação médica, em português brasileiro, adequada quanto à confiabilidade editorial, qualidade e acessibilidade textual para o público leigo.
- Research Article
- 10.1097/md.0000000000049308
- Jun 12, 2026
- Medicine
- Min Wang + 2 more
Diabetic nephropathy (DN) is a common, life-threatening complication of diabetes, contributing to the global disease burden. With the advent of video platforms, health information is being more widely disseminated. However, the quality of such content varies widely, which may influence the public’s perception. This study aimed to evaluate the upload sources, content, and characteristics of DN-related videos on TikTok and Bilibili and to explore descriptive associations between video quality scores and selected video characteristics. This cross-sectional content analysis included 166 DN-related videos. Video quality was assessed using the Global Quality Scale (GQS), modified DISCERN (mDISCERN), and Journal of the American Medical Association (JAMA) benchmark criteria. Descriptive subgroup and correlation analyses were performed to examine cross-sectional associations between video quality scores and selected video attributes. No multivariable adjustment was performed. In unadjusted cross-sectional comparisons, TikTok videos showed higher observed engagement counts at the time of data collection than Bilibili videos, whereas no statistically significant differences were observed in video duration or quality indicators after correction for multiple comparisons. In unadjusted descriptive subgroup comparisons, videos uploaded by experts showed more favorable results in selected quality-related measures, particularly GQS and JAMA, than videos uploaded by individual users. No clear association was observed between video quality and snapshot engagement metrics recorded at the time of retrieval. This study identified descriptive differences in the presentation and dissemination patterns of DN-related health information across TikTok and Bilibili. Because the analyses were observational, cross-sectional, and unadjusted for potential confounders such as video length and content type, the observed differences between platforms and uploader types should be interpreted as descriptive associations only rather than independent effects.
- Research Article
- 10.1007/s00210-026-05560-x
- Jun 12, 2026
- Naunyn-Schmiedeberg's archives of pharmacology
- Ruogu Nie + 5 more
This study is a cross-sectional analysis aimed at systematically evaluating the information quality and reliability of popular science content related to drug-induced liver injury (DILI) on the two major Chinese video platforms, TikTok (Douyin) and BiliBili, and analyzing its content characteristics. On December 20, 2025, searches were conducted in the Chinese versions of the TikTok (Douyin) and BiliBili mobile applications using "" as the single keyword. Exclude content that does not meet the requirements based on the exclusion criteria, and ultimately retain the top 100 videos from each platform that meet the standards for analysis (N = 200). Two trained reviewers independently performed blinded assessments using the Global Quality Score (GQS), Journal of the American Medical Association (JAMA) benchmark criteria, and the modified DISCERN (mDISCERN) tool. Non-parametric tests were used to compare differences between groups, and Spearman correlation analysis was employed to explore the relationship between video characteristics and quality scores. User engagement metrics (likes, favorites, shares) for TikTok (Douyin) videos were significantly higher than those for BiliBili (p < 0.001). In terms of information quality, TikTok (Douyin) videos scored significantly higher than BiliBili on the GQS, JAMA, and mDISCERN scales (p < 0.001). There were differences in quality across content types: videos on "medication knowledge" received the highest mDISCERN reliability scores and "disease knowledge" videos scored higher in GQS practicality. Correlation analysis showed a weak positive correlation between user engagement metrics and mDISCERN scores. Among the 107 videos mentioning liver injury-related drugs, chemical drugs (antibacterial, chemotherapeutic, anti-tuberculosis drugs, and anti-inflammatory drugs) and traditional Chinese medicines (such as He Shou Wu) were mentioned most frequently. In this cross-sectional sample, TikTok (Douyin) videos demonstrated higher quality scores and user engagement than those on BiliBili, while professionals outperformed general users only on the JAMA criteria. Although some videos mentioned medications associated with liver injury, the information was generally oversimplified and biased toward trending topics. Hence, active information seekers should critically appraise the scientific soundness of medical short videos on platforms like TikTok and BiliBili before making healthcare decisions.
- Research Article
- 10.1016/j.artd.2026.102007
- Jun 1, 2026
- Arthroplasty today
- Trace Clark + 5 more
Scroll, Learn, Replace: Evaluating TikTok's Role in Patient Education in Total Knee Arthroplasty.
- Research Article
- 10.1177/10926429261435688
- Jun 1, 2026
- Journal of laparoendoscopic & advanced surgical techniques. Part A
- Alp Omer Canturk + 5 more
Video-based learning is a central tool in minimally invasive surgical training; however, the educational/reporting quality and reliability of online content must be evaluated using objective criteria. This study aimed to compare the educational quality of transanal total mesorectal excision (TaTME) videos published on YouTube and WebSurg and to test the validity of the TaTME-specific StepScore scale. A cross-sectional content analysis was performed across platforms (total n = 30; YouTube = 15, WebSurg = 15). Videos were scored using Laparoscopic Surgery Video Educational Guidelines (LAP-VEGaS) (0-18), Journal of the American Medical Association (JAMA) (0-4), modified DISCERN (mDISCERN) (5-25), and total mesorectal excision (TME) StepScore (14 steps, 0-28). Correlations were examined using Spearman's ρ with false discovery rate adjustment. Logistic regression and ROC/AUC were used to predict LAP-VEGaS ≥ 11 adequacy; the optimal threshold was determined using the Youden index. Total scores and LAP-VEGaS ≥ 11 rates were similar across platforms (YouTube 66.7%; WebSurg 60.0%). StepScore showed a strong correlation with LAPVEGaS and mDISCERN and a moderate correlation with JAMA. Each +1 point increase in StepScore increased the odds of a LAP-VEGaS score ≥ 11. According to the Youden analysis, a StepScore ≥ 15 was found to be the best threshold. Popular TaTME videos on YouTube and WebSurg appear similar in terms of educational/reporting quality. The procedure-specific StepScore is consistent with general quality measures and can predict LAP-VEGaS ≥ 11 adequacy with a practical ≥15 point target. Using StepScore for video assessment and as a step-by-step instructional checklist may contribute to improving TaTME training standards.
- Research Article
- 10.1002/ovs2.70071
- Jun 1, 2026
- Optometry and vision science : official publication of the American Academy of Optometry
- Junran Li + 8 more
To systematically evaluate the quality of eye disease videos on TikTok, WeChat, and rednote, explore links between engagement and quality, and offer evidence-based guidance for ophthalmic health communication. The top 100 videos retrieved using the keywords "cataract," "glaucoma," and "high myopia" were screened on TikTok, WeChat, and rednote on 3 October 2025. Two reviewers independently assessed video quality using Journal of the American Medical Association (JAMA), the global quality score (GQS), modified DISCERN, and the Patient Education Materials Assessment Tool (PEMAT). Group differences were analyzed using Kruskal-Wallis and χ2/Fisher exact tests, and adjusted associations were examined using Poisson regression with robust standard errors. A total of 827 eligible videos were analyzed. Most videos were uploaded by physicians and focused on disease knowledge. Across TikTok, WeChat, and rednote, video characteristics, engagement, source, content, presentation form, and quality scores differed significantly. In adjusted analyses, compared with TikTok, WeChat videos had lower likes and comments, whereas rednote videos had lower engagement across all four outcomes. High-myopia videos showed higher engagement across all outcomes, while glaucoma videos showed higher collections and shares. Hospital-uploaded videos were associated with lower engagement, whereas news agency videos were associated with higher engagement. Personal experience videos were associated with higher comments and collections. Higher JAMA scores were consistently associated with lower engagement, whereas modified DISCERN and PEMAT actionability showed inverse associations only for selected outcomes. This study represents the first large-scale cross-sectional evaluation of science communication on potentially blinding eye diseases across major Chinese short-video platforms. High engagement does not equate to high quality; in fact, engagement metrics were significantly negatively correlated with reliability, scientific accuracy, and understandability. Clinicians should uphold scientific rigor and use accessible and friendly language to improve public eye health literacy.
- Research Article
- 10.1097/md.0000000000049080
- May 29, 2026
- Medicine
- Nengying He + 5 more
Video platforms are major sources of health information; however, the accuracy of musculoskeletal content is uncertain. Triangular fibrocartilage complex (TFCC) injury is common, and many patients seek guidance online. The present study aims to evaluate TFCC-related videos on Bilibili and TikTok to measure dissemination, quality, and reliability, and to identify features associated with higher-quality content. This study conducted a systematic screening and assessment of videos related to the TFCC on the Bilibili and TikTok platforms. Video characteristics were collected, and quality was assessed by Global Quality Score, modified DISCERN, and Journal of the American Medical Association (JAMA) benchmark criteria. A total of 211 TFCC-related videos were analyzed. TikTok exhibited higher dissemination metrics (likes, comments, collections, and shares; P < .001) but a shorter duration than Bilibili. TikTok also hosted a higher proportion of medical professionals (51% vs 27%) and demonstrated significantly higher JAMA benchmark criteria scores (P = .04). Videos from orthopedic and rehabilitation specialists achieved superior Global Quality Score, modified DISCERN, and JAMA scores (P < .001) compared to nonmedical uploaders. Video length correlated positively with quality on Bilibili (r = 0.39, P < .05), whereas engagement metrics did not correlate with information quality. Video quality and reliability of TFCC-related content varied by platform and creator, although the overall quality remained suboptimal. TikTok videos achieved broader reach and higher average quality, whereas medical professionals produced the most reliable content. Uploaders should ensure accuracy, originality, and clarity, while platforms should refine algorithms to highlight evidence-based videos.
- Research Article
- 10.21037/jtd-2026-0582
- May 27, 2026
- Journal of Thoracic Disease
- Yuyu Dai + 3 more
BackgroundChronic obstructive pulmonary disease (COPD) continues to pose a significant global health burden, while social media platforms such as TikTok and Bilibili have become important sources of public health information. However, the quality and reliability of video content on these platforms remain unclear. Therefore, this study aimed to evaluate the quality and reliability of COPD-related short videos on TikTok and Bilibili in China.MethodsA cross-sectional study was conducted on 275 COPD-related short videos on TikTok and Bilibili. Video characteristics, uploader types, content themes, and presentation formats were extracted. Video quality and reliability were evaluated using the Global Quality Score (GQS), the modified Decision-making Information Support Criteria for Evaluating the Reliability of Non-randomised Studies (mDISCERN) tool, the Journal of the American Medical Association (JAMA) benchmark criteria, and the Video Information and Quality Index (VIQI). Correlation analysis was further performed among video metrics and quality scores.ResultsCompared to Bilibili, TikTok was more popular, although the length of the videos on TikTok was shorter than that of the videos on Bilibili (P<0.001). Videos on Bilibili had significantly higher GQS scores (P<0.001) and mDISCERN scores (P=0.01) than those on TikTok. Videos from medical practitioners and science communicators generally exhibited higher quality and reliability compared to those from general users across most assessment tools. Medical practitioners scored higher on the JAMA criteria compared to science communicators, with no significant differences observed in the other assessment tools. Negative correlations were found between GQS scores and engagement metrics (likes, comments, and shares), while positive correlations were found between VIQI scores and all four engagement metrics.ConclusionsCOPD-related videos on TikTok and Bilibili were deficient in quality and reliability. Enhancing the educational value of health information on both platforms requires greater involvement from medical practitioners and science communicators, alongside improved platform-level content regulation.
- Research Article
- 10.1186/s12876-026-04960-w
- May 25, 2026
- BMC gastroenterology
- Xingyao Lu + 2 more
Choledocholithiasis, a disease with a rising incidence and potential for severe complications, has prompted many to seek health information online. TikTok and Bilibili have emerged as key platforms for disseminating such information. This study evaluates the quality and reliability of short videos on choledocholithiasis on these platforms. This study analyzed the top 100 choledocholithiasis-related videos from TikTok and Bilibili. The Global Quality Score (GQS), modified DISCERN (mDISCERN) tool, and the Journal of the American Medical Association (JAMA) criteria were employed to assess video quality. Cohen's Kappa coefficient is used to assess inter-rater agreement.Group comparisons were conducted using Mann-Whitney U and Kruskal-Wallis H tests, while Spearman's correlation was utilized for correlation analysis. A total of 170 videos were included, predominantly uploaded by hepatobiliary surgeons. The content of these videos mainly focuses on treatment (81.76%), with limited coverage of etiology and diagnosis. The overall video quality was mediocre, with median scores of 3 for GQS (IQR: 3.00-4.00), 2 for mDISCERN (IQR: 2.00-2.00), and 2 for JAMA (IQR: 2.00-2.00). Videos from hepatobiliary surgeons generally exhibited superior quality. Notably, video quality showed no correlation with engagement metrics. The content of choledocholithiasis-related videos is structurally deficient. The videos' quality, reliability, and transparency are relatively poor, though those uploaded by hepatobiliary surgeons stand out for their superior quality and reliability. Importantly, video quality is independent of engagement metrics.