Expert evaluation improves clarity, safety, and athlete-specific relevance of GPT-4-generated sleep and jet lag guidance
In a two-round Delphi process with 17 sleep experts, GPT-4-generated FAQ responses on sleep and jet lag were evaluated and revised. Consensus increased from 75% (15/20 items) in round 1 to 90% (18/20 items) in round 2, with qualitative feedback indicating improved clarity, accuracy, and athlete-specific relevance after expert revision.
This article is automatically generated from the structured evidence profile behind the claim above. Scores reflect the quality and quantity of available research, not clinical advice.
Expert evaluation improves clarity, safety, and athlete-specific relevance of GPT-4-generated sleep and jet lag guidance The current body of evidence comprises 1 study. EvidenceHub rates the overall confidence at 32/100 (low).
The Claim
Expert evaluation improves clarity, safety, and athlete-specific relevance of GPT-4-generated sleep and jet lag guidance
This conclusion is most relevant to: 17 international sleep and circadian experts from the Athlete Travel & Sleep Interest Group (ATSIG).
What the Research Shows
The conclusion draws on 1 linked study. Highlights from the cited literature:
- ▸From GPT-4 to Expert-Endorsed Athlete Guidance: A Delphi Consensus on Sleep and Jet Lag. (Sports medicine (Auckland, N.Z.), 2026) —
How It Works
The proposed biological pathway:
- ▸GPT-4 generated 20 FAQ responses on sleep and jet lag
- ▸Experts rated items for appropriateness and provided qualitative feedback
- ▸Items revised using inductive thematic coding after round 1
- ▸Consensus increased from 75% to 90% after two rounds
Who Might Benefit
Evidence fit by population:
- ▸17 international sleep and circadian experts from the Athlete Travel & Sleep Interest Group (ATSIG)
Recommended Dose
N/A
Limitations & Caveats
Important context when interpreting this evidence:
- ▸Item-level statistical comparisons did not remain significant after Bonferroni correction for multiple testing
- ▸GPT-4-derived content should not be used as standalone guidance; final outputs are consensus-based rather than definitive proof of correctness
Frequently Asked Questions
Can GPT-4-generated health advice be used directly for elite athletes?▼
No, GPT-4-derived content may serve as useful preliminary material for expert discussion but should not be used as standalone guidance for elite athletes.
What were the main concerns with GPT-4-generated sleep advice?▼
Common concerns included imprecise or misleading content (55%), lack of athlete-specific relevance (30%), and outdated evidence (11%).
How effective was expert evaluation in improving AI-generated guidance?▼
Expert evaluation improved clarity, safety, and athlete-specific relevance, with consensus increasing from 75% to 90% of items after two Delphi rounds.
Which topics did not reach consensus after expert review?▼
Sleep Q6 (sleep and injury risk) narrowly missed the threshold with 64.7% agreement, and Jet Lag Q9 (melatonin and sleep aids) remained below the 70% consensus threshold.
References
- 1.Vitale J, McCall A, Halson S, van Rensburg DCJ. “From GPT-4 to Expert-Endorsed Athlete Guidance: A Delphi Consensus on Sleep and Jet Lag..” Sports medicine (Auckland, N.Z.), 2026. PMID: 42470603 DOI: 10.1007/s40279-026-02484-7