Skip to content
Open access

An Analysis of the Nutritional Value of Diets Generated by Large Language Models Under Average User Conditions—3(LM)Diet Study

Jul 2026 · Nutrients · Vol 18, pp. 2363 · 0 citations · 60 references
Medicine

TL;DR

Unsupervised use of LLM-generated meal plans by non-expert users may produce nutritionally inaccurate diets and should not replace professional dietary counselling, according to the authors.

Abstract

Background/Objectives: AI tools are becoming increasingly common, including in areas related to health and diet. However, there are no studies indicating the nutritional value of diets generated by large language models (LLMs). The aim of this study was to analyse the nutritional value of diets generated by three major language models. Methods: Three LLMs were used in the study: ChatGPT, Gemini and Copilot. Using each of these, 35-day meal plans were generated for three energy variants: 2000 kcal, 2500 kcal and 3000 kcal. The meal plans were then analysed using dietary software. Results: All LLMs generated meal plans with an insufficient energy value (p < 0.001, one-sample t-test) by 284–546 kcal, depending on the model and the targeted energy content. The meal plans were characterised by an excessively high proportion of energy from protein and an excessively low proportion of energy from fat. Carbohydrates were planned in the correct amounts. ChatGPT and Copilot generated meal plans with a generally adequate vitamin and mineral content. Gemini generated meal plans that were often deficient, particularly for diets with lower energy values. Conclusions: Diets generated by LLMs have many shortcomings. Unsupervised use of LLM-generated meal plans by non-expert users may produce nutritionally inaccurate diets and should not replace professional dietary counselling.

Read PDF

Similar papers

Review Jul 2026

Evaluation of Energy and Nutrient Estimates from Large Language Models Using Text-Based Queries.

BACKGROUND Large language models (LLMs) have emerged as promising tools for estimating energy and nutrient values, yet most existing evaluations focus on image-based queries rather than text. Few studies compare LLM estimates with reference databases commonly used in nutrition research. OBJECTIVE To examine agreement between LLM estimates and a research food composition database for energy and nutrients, and to determine if agreement varies by food group. METHODS We conducted a cross-sectional analysis of energy and nutrient estimates for frequently consumed food items in the United States (US). Food items were entered as text prompts into four LLMs (ChatGPT 5.2, Claude Opus 4.5, Gemini 3 Pro Preview, and Llama 4 Maverick), which provided energy and nutrient estimates. Corresponding foods were matched to the Nutrition Coordinating Center (NCC) Food and Nutrient Database, and agreement between LLM and database values was assessed using intraclass correlation coefficients (ICCs) and Bland-Altman analyses. Agreement was also evaluated within the three most frequently consumed food groups. RESULTS Agreement was high for energy and macronutrients for all LLMs. Variability was observed for several micronutrients, particularly vitamin D, folate, and iron. Claude Opus 4.5 showed consistently high agreement, with no nutrients classified as poor. Other LLMs exhibited poor agreement for at least one micronutrient. Certain food categories, including condiments and mixed dishes, contributed disproportionately to variability. However, agreement remained high within the most frequently consumed broader food groups. CONCLUSIONS LLMs show promise for estimating energy and macronutrients. However, performance for micronutrients requires further improvement and may affect overall dietary assessment.

Razi Lawabni, Olivia Solano, Leo L. Kampen et al. · 0 citations
Aug 2026

Evaluating LLM Accuracy in Predicting Peruvian Meal Nutrition.

BACKGROUND Artificial intelligence applications have been developed to predict the nutrient content of meals. However, none have been evaluated in the context of Peruvian cuisine, characterized by diverse ingredients and recipes. We assessed whether large language models (LLMs) could predict the nutritional content of Peruvian meals. METHODS Using a dataset of 510 unique lunch images extracted from a Peruvian cookbook, we compared nutrient values from recipe data against predictions generated by LLMs (Gemma-3 4B, 12B, and 27B). The LLMs were given the meal name and a photograph and prompted to produce narrative descriptions of the meal. Using the descriptions, the same LLMs were prompted to estimate six nutrients: energy (kcal/serving), protein (g/serving), carbohydrates (g/serving), iron (mg/serving), vitamin A (μg/serving), and zinc (mg/serving). Agreement proportions and errors metrics were calculated against the values from the recipe book. RESULTS The 27B LLM achieved the highest agreement proportions across most nutrients-calories (45%), carbohydrates (31%), iron (15%), vitamin A (19%), and zinc (31%)-while the 12B model performed best for protein (70% agreement). The 27B model yielded the lowest mean absolute error (MAE) for calories (108 kcal), carbohydrates (26 g), iron (4 mg), and zinc (1 mg). The 12B LLM had the lowest MAE for protein (6 g) and vitamin A (667 μg). The 4B LLM showed the poorest performance across metrics. CONCLUSIONS LLMs can generate estimates of nutrient content from narrative descriptions of Peruvian meals, but current performance levels fall short of the precision required for clinical deployment or commercial consumer-facing applications.

R. M. Carrillo-Larco, Mariano Gallo Ruelas, Mika Matsuzaki et al. · 0 citations

The role of artificial intelligence in developing meal plans in outpatient dietetics: A feasibility and proof of concept study.

BACKGROUND Personalized meal planning by registered dietitian nutritionists (RDNs) is time-intensive. Large language models (LLMs) may automate drafting meal plans, but their nutritional accuracy in clinical practice is uncertain. METHODS In this proof-of-concept study, five outpatient RDNs and four LLMs (Gemini, CoPilot, ChatGPT 4.0, and customized ChatGPT 4.0) each generated 3-day meal plans for five validated clinical scenarios. Effectiveness was defined as accuracy in meeting pre-specified energy, protein, carbohydrate, fat, and sodium targets. Time to create plans and RDN comfort (self-rated confidence in nutritional accuracy and clinical appropriateness on 1-5 Likert scale) were recorded. Three independent RDNs, blinded to source, analyzed nutrient content using Nutritionist Pro. Group differences were assessed with t-test and ANOVA. RESULTS All LLMs and RDNs produced feasible meal plans. LLMs generated meal plans in under 1 min, whereas RDNs required a mean of 44 min per scenario. RDNs reported comfort levels ranging from 3.8 to 4.8. Across most scenarios, LLM plans delivered a smaller proportion of requested energy than RDN plans, which more consistently approached energy targets. Both groups performed similarly for the Mediterranean diet scenario. Overall, protein accuracy did not differ. However, in chronic kidney disease, LLMs undershot the guideline-based protein target, while RDNs tended to modestly exceed it. Accuracy for low-carbohydrate, fat, and sodium diets was comparable. CONCLUSION LLMs can rapidly generate clinically plausible meal plans but are less reliable than RDNs in achieving prescribed energy and selected macronutrient goals. Prompt precision is essential for nutrient-specific targets. A hybrid model in which RDNs refine LLM-generated drafts may leverage efficiency without sacrificing clinical accuracy.

M. Mundi, Osman Mohamed Elfadil, Danielle P. Johnson et al. · 0 citations
Open access Aug 2026

Exploratory benchmarking of AI-generated diet plans for inherited protein metabolism disorders: a simulation-based evaluation of nutritional accuracy and clinical safety

Artificial intelligence (AI)-based large language models (LLMs) are increasingly used to support nutrition-related decision-making; however, their ability to generate clinically appropriate dietary plans for inherited protein metabolism disorders remains largely unexplored. This study aimed to perform an exploratory simulation-based benchmarking analysis of AI-generated dietary plans for phenylketonuria (PKU), maple syrup urine disease (MSUD), and propionic acidemia (PPA) using disease-specific metabolic nutrition guidelines. Standardized pediatric case scenarios were developed for PKU, MSUD, and PPA. Using identical English-language prompts, ChatGPT-5.3 Pro and Gemini 3 Pro Advanced each generated 3-day dietary plans. Nutrient composition was analyzed using the BeBiS Nutrition Information System and evaluated against Dietary Reference Intakes (DRIs). Disease-specific nutritional targets, including amino acid intake, protein distribution, and energy provision, were benchmarked against recommendations from Genetic Metabolic Dietitians International (GMDI). Nutritional characteristics of the dietary plans generated by the two AI models were compared using exploratory statistical analyses. Both LLMs generated structured dietary plans with generally acceptable overall nutritional characteristics; however, clinically relevant deviations from disease-specific nutritional targets were identified across all three disorders. In the PKU case, both models achieved the recommended phenylalanine range, but neither simultaneously met protein and tyrosine recommendations. In the MSUD case, differences were primarily related to energy provision and branched-chain amino acid targets, while in the PPA case neither model achieved the recommended balance between intact protein and total protein. These findings demonstrated that conventional measures of nutritional adequacy alone were insufficient to determine the clinical appropriateness of AI-generated dietary plans for inherited protein metabolism disorders. General-purpose LLMs can generate structured dietary plans for inherited protein metabolism disorders; however, disease-specific metabolic targets are not consistently achieved. Evaluation of AI-generated dietary plans should therefore extend beyond conventional nutritional assessment and incorporate disease-specific benchmarking against established metabolic nutrition guidelines. This study provides a disease-specific benchmarking framework for evaluating AI-generated dietary plans in inherited protein metabolism disorders.

Taha Gökmen Ülger, E. Adıgüzel · 0 citations
Open access Jul 2026

What are the best sources of protein? Introducing the BPI Score

The question, “What are the best sources of protein?” is complicated by the significant hidden health and environmental costs associated with many popular options. Existing food rating systems, such as nutritional traffic lights and carbon footprint labels, are valuable but often unidimensional. This narrow focus can conceal critical trade-offs and mislead consumers by failing to present a holistic picture. This perspective paper introduces the BPI Score, a new, multidimensional rating system designed to provide a more comprehensive answer. The goal is not necessarily to create another front-of-package label, but to foster a new literacy that empowers consumers, policymakers, and organizations to make more informed decisions. The initial version (V1) of the BPI Score was used to evaluate 20 high-protein products across two primary categories: people and planet. The Planet Score assesses environmental impact using data on GHG emissions, water pollution, and resource use. The People Score considers health factors like sodium and saturated fat content alongside accessibility metrics such as affordability and availability. A clear disparity emerged between plant-based and animal-based proteins. Tofu emerged as the highest-scoring product, while cheddar cheese ranked last. The top quintile of products consisted solely of plant-based proteins, while the bottom quintile was composed entirely of animal-based proteins. While acknowledging the limitations of this first iteration, the score provides a robust foundation for a more nuanced conversation. Ultimately, beyond the spreadsheets and scores, lies a fundamental reckoning: our protein choice is a referendum on the compassion we are willing to show to the planet and to future generations.

C. MacDonald · 0 citations
Open access Jul 2026

Who is at high risk of poor diet quality? A precision public health nutrition approach using machine learning.

BACKGROUND Suboptimal diet quality contributes to poor health. Numerous factors at all levels of the socioecological model intersect to influence individuals' diet quality. A precision public health nutrition framework can be used to understand how these factors jointly shape diet quality across several subgroups in the population. This is, however, very challenging to operationalize. OBJECTIVE The purpose of this study was to assess whether machine learning can be leveraged to operationalize a precision public health nutrition approach by more precisely identifying subgroups with the highest proportion of adults with poor diet quality and the most important predictors of diet quality. METHODS We conducted a secondary analysis of cross-sectional data from the 2018 and 2019 International Food Policy Study in Canada (n=5,093). A total of 42 candidate predictors within four domains (sociodemographic characteristics and socioeconomic position; food policies and environments; food literacy; health-related practices and indicators) were used. The Healthy Eating Index-2015 (HEI-2015) was used to assess diet quality; tertiles of HEI-2015 scores were defined as the outcome. Conditional inference tree (CIT) and conditional random forest (CRF) analyses were conducted. RESULTS The CIT partitioned the sample into seven subgroups with different proportions of adults with lower, moderate or higher diet quality using six predictors. The probability of lower diet quality ranged from 16.9% to 49.9% across subgroups. The subgroup with the highest proportion of adults with lower diet quality was characterised by individuals confident in using ≤4 cooking techniques. Based on the CRF model and conditional variable importance, the five most important predictors of diet quality in were: frequency of food label use, confidence in using cooking techniques, perceived general health, household food insecurity status, and health literacy. CONCLUSIONS This study provides evidence that machine learning approaches can be leveraged to operationalize a precision public health nutrition approach.

R. Blanchet, N. Doan, Sara Nejatinamini et al. · 0 citations