Gluten-free diet plans generated by ChatGPT and Gemini differ in nutrient adequacy, while diet quality and carbon footprint are comparable
Nutrition, cilt.150, 2026 (SCI-Expanded, Scopus)
- Yayın Türü: Makale / Tam Makale
- Cilt numarası: 150
- Basım Tarihi: 2026
- Doi Numarası: 10.1016/j.nut.2026.113323
- Dergi Adı: Nutrition
- Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, EMBASE, MEDLINE, Natural Science Collection (ProQuest), Biological Science Database (ProQuest), Health Research Premium Collection (ProQuest), Pharma Collection (ProQuest)
- Anahtar Kelimeler: Carbon footprint, Diet quality, Gluten-free diet, Large language models, Nutrient adequacy
- Ankara Üniversitesi Adresli: Evet
Özet
Background Large language models are used for meal planning, but their suitability for restrictive diets remains unclear. We hypothesized that ChatGPT would generate gluten-free (GF) menus with lower energy and micronutrient adequacy than Gemini. Methods In this exploratory prompt-based comparison, ChatGPT (GPT-4o) and Gemini (2.5 Flash) each generated six daily 2000 kcal GF menus for a hypothetical 30-year-old woman with celiac disease using a standardized prompt. Generation was repeated after 1 wk for test–retest and sensitivity analyses (12 menus/model). The nutrient composition was analyzed using BeBiS v9.0. Micronutrient adequacy was summarized using nutrient adequacy ratios (NARs; truncated at 1.00) and the mean adequacy ratio. Diet quality was assessed with the Diet Quality Index-International, and carbon footprint was estimated using sustainability assessment of food and diets factors. Between-model differences were assessed using t tests or exact Mann–Whitney U tests with multiplicity adjustment; short-term output variability was characterized using intraclass correlation coefficients and Bland–Altman plots. Results In the initial dataset, ChatGPT menus provided less energy than Gemini menus (−379.8 kcal/d; P adj = 0.002) and showed lower mean adequacy ratio (−0.05; P adj = 0.002), mainly reflecting calcium and iron shortfalls. Diet Quality Index-International scores were similar between models, and absolute carbon footprint was numerically lower in ChatGPT menus but not statistically significant ( P adj = 0.089). 1-wk output variability was observed across outcomes. Conclusions In this exploratory prompt-based study, ChatGPT and Gemini generated GF menus with different nutrient adequacy profiles under the tested conditions, whereas diet quality and carbon footprint were comparable. Given the limited prompt sample and observed output variability, these findings should be interpreted as hypothesis-generating. AI-generated GF menu outputs require standardized prompting, independent nutrient verification, GF safety assessment, and dietitian review before clinical use.