Supplements

Generative AI and large language models show promise as supervised decision-support tools in sports medicine, but current evidence does not support unsupervised clinical use.

A scoping review of 32 studies found that LLM accuracy varies substantially across sports medicine domains, with content validity ratios for sleep recommendations ranging from 0.33 (GPT-3.5) to 0.67 (GPT-4). Hallucination risk was rated critical for general-purpose chatbots but reduced in retrieval-augmented systems. The review concludes that while GenAI has potential, it is not yet safe for unsupervised use.

Last updated: Aug 17, 2026โ€ข0 RCTsโ€ข๐Ÿ“– Read as article โ†’

Evidence Score

Evidence Score32/100
Human RCTโ˜†โ˜†โ˜†โ˜†โ˜†
Meta-analysisโ˜†โ˜†โ˜†โ˜†โ˜†
Mechanismโ˜…โ˜…โ˜…โ˜…โ˜…
Safetyโ˜…โ˜…โ˜…โ˜…โ˜†
Confidencelow

Study Evidence

Study 1. Generative artificial intelligence and large language models in sports medicine: a scoping review of applications, accuracy, and ethical implications.

observational

Dergaa I, Dergaa MA, Razzak M, Ceylan Hİ, Stefanica V, Muntean RI, Guelmami N ยท Frontiers in public health (2026)

Participants: N/A
Duration: N/A
Intervention: Evaluation of GenAI/LLM applications across seven sports medicine domains (training, nutrition, rehabilitation, mental health, clinical decision support, academic writing, ethics/governance)
Outcome: Accuracy, hallucination risk, and ethical/governance concerns
Effect Size: N/A
Population: Sports medicine practitioners, athletes, and coaches (studies included in scoping review)

Result:

Mechanism Graph

LLMs generate responses based on training data
โ†“
Responses are evaluated for content validity against expert guidelines
โ†“
Hallucination risk assessed for different system types
โ†“
General-purpose chatbots show critical hallucination risk, retrieval-augmented systems reduce risk
โ†“
Result: Unsupervised clinical use not supported

Limitations

  • โš Scoping review, not a systematic review with meta-analysis
  • โš Only 32 studies included, with limited validation studies
  • โš No sport-specific validation datasets available
  • โš Rapidly evolving technology may make findings outdated

Frequently Asked Questions

What is the accuracy of GPT-4 for sleep recommendations in sports medicine?โ–ผ

GPT-4 achieved a content validity ratio of 0.67 for sleep recommendations, which is considered acceptable, while GPT-3.5 had a lower CVR of 0.33.

Are large language models safe for unsupervised use in sports medicine?โ–ผ

No, the review found that current evidence does not support unsupervised clinical use due to hallucination risks and lack of sport-specific validation.

What are the main applications of GenAI in sports medicine?โ–ผ

Applications include training prescription, nutrition, rehabilitation, mental health support, injury prevention, clinical decision support, and academic writing.

What is the hallucination risk of general-purpose chatbots compared to retrieval-augmented systems?โ–ผ

General-purpose chatbots have a critical hallucination risk, while retrieval-augmented systems substantially reduce this risk.

Products

Affiliate links coming soon. We only recommend products that match the doses and forms used in the cited research.

References

  1. 1.Dergaa I, Dergaa MA, Razzak M, Ceylan Hİ, Stefanica V, Muntean RI, Guelmami N. "Generative artificial intelligence and large language models in sports medicine: a scoping review of applications, accuracy, and ethical implications.." Frontiers in public health, 2026. PMID: 42591395 DOI: 10.3389/fpubh.2026.1843535
Disclaimer: This content is for educational purposes only and is not medical advice. Evidence scores reflect the quality and quantity of available research, not clinical recommendations. Always consult a healthcare professional before starting any supplement or intervention.