NutriMLLM: 식품 이미지 기반 미량 영양소 분석을 위한 다중 모드 대규모 언어 모델
NutriMLLM: Multimodal Large Language Models for Dietary Micronutrient Analysis
음식 이미지를 활용한 체계적인 식이 미량 영양소 추정은 임상 영양 관리를 개선할 수 있지만, 이러한 모델 훈련에는 다양한 음식과 완전한 영양 프로필을 연결하는 대규모 다중 모드 데이터 세트가 필요합니다. 본 연구에서는 기존의 다중 모드 대규모 언어 모델(MLLM), 심지어 선도적인 독점 모델조차도 이 작업에 신뢰성이 없음을 보여줍니다. 5개의 모델 패밀리와 4개의 독립적인 평가 벤치마크(ASA24, SNAPMe, FNDDS 및 NutriBench)를 통해 모델이 종종 답변을 회피하거나 통계적으로 타당하지 않은 값을 반환하는 것을 확인했습니다. 이러한 격차를 비용이 많이 드는 전문가 주석 없이 해결하기 위해, 우리는 10년간 수집된 인구 규모의 24시간 식이 회고 데이터를 활용하여 텍스트-이미지 생성의 구조화된 프롬프트로 변환했습니다. 이 파이프라인을 통해 약 110만 개의 이미지-설명-영양소 3중항으로 구성된 합성 데이터 코퍼스를 생성했으며, 각 항목은 생성된 음식 이미지와 65가지 영양소를 모두 포함하는 라벨을 연결합니다. 현재까지 알려진 바로는, 이는 포괄적인 미량 영양소 주석이 포함된 가장 큰 규모의 합성 식품-이미지 코퍼스이며, 출판 후 공개될 예정입니다. Qwen3-VL (2B/4B/8B/30B) 및 GLM-4.6V-Flash 모델을 이 데이터 코퍼스로 미세 조정하여 NutriMLLM이라는, 체계적인 식이 미량 영양소 추정에 특화된 최초의 비전-언어 모델 패밀리를 개발했습니다. 우리는 회피율, 환각 현상, 전반적인 사용성 및 개별 영양소에 대한 수치 정확도를 별도로 측정하는 4가지 구성 요소 프레임워크를 사용하여 이러한 모델을 평가했습니다. 실제 음식 이미지에서 모든 NutriMLLM 변형은 65가지 영양소 모두에 대해 거의 완벽한 커버리지를 달성했으며, 가장 큰 변형은 대부분의 영양소에서 독점적인 기준 모델(GPT-5, Gemini 3 및 Claude Sonnet 4.5)과 동등하거나 그 이상의 정확도를 보였습니다. 이러한 결과는 식이 회고 데이터를 활용한 합성 데이터 기반 지도 학습이 이미지 기반의 포괄적인 미량 영양소 추정을 실현 가능한 엔지니어링 문제로 만들 수 있으며, 식이 평가, 개인 맞춤형 영양 지침 및 인구 규모의 미량 영양소 감시를 지원할 수 있음을 보여줍니다.
Comprehensive estimation of dietary micronutrients from food images could improve clinical nutrition care, but training such models requires large multimodal datasets linking diverse foods to complete nutrient profiles. We first show that existing multimodal large language models (MLLMs), including leading proprietary models, are unreliable for this task. Across five model families and four independent evaluation benchmarks (ASA24, SNAPMe, FNDDS, and NutriBench), models frequently abstained or returned statistically implausible values. To address this gap without costly expert annotation, we repurposed a decade of population-scale 24-hour dietary recalls as structured prompts for text-to-image generation. This pipeline produced a synthetic corpus of about 1.1 million image-description-nutrient triplets, each pairing a generated food image with a complete 65-nutrient label. To our knowledge, this is the largest synthetic food-image corpus with comprehensive micronutrient annotation planned for public release upon publication. Fine-tuning Qwen3-VL (2B/4B/8B/30B) and GLM-4.6V-Flash on this corpus yielded NutriMLLM, the first family of vision-language models specialized for comprehensive dietary micronutrient estimation. We evaluate these models with a four-component framework that separately measures abstention, hallucination, overall usability, and per-nutrient numerical accuracy. On real food images, every NutriMLLM variant achieved near-complete coverage across all 65 nutrients, and the largest variant matched or exceeded proprietary baselines (GPT-5, Gemini 3, and Claude Sonnet 4.5) in accuracy on most nutrients. These results show that recall-driven synthetic supervision can make image-based comprehensive micronutrient estimation a tractable engineering problem and support dietary assessment, personalized nutrition guidance, and population-scale micronutrient surveillance.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.