OmniFood-Bench: 영양 추론 및 맞춤형 건강 상담을 위한 VLM 평가
OmniFood-Bench: Evaluating VLMs for Nutrient Reasoning and Personalized Health Advice
대규모 시각-언어 모델(VLM)이 핵심 인프라에 빠르게 통합되면서 개인 맞춤형 의료 및 식단 관리에 혁신을 가져올 것으로 예상됩니다. 그러나 식품 시스템 분야에서 자율 에이전트는 독특하고 지속적인 과제인 '시각적 외관과 본질적인 영양 구성 간의 체계적인 정보 비대칭'에 직면합니다. 기존 벤치마크는 주로 음식 카테고리 인식과 같은 세분화되지 않은 분류 작업에 중점을 두며, 실제 식단 관리에 필요한 복잡한 추론 과정을 평가하는 데 부족합니다. 구체적으로, 숨겨진 재료를 파악하여 물리적 양을 추정하고, 최종적으로 안전이 중요한 의료 조언을 제공하는 능력을 평가하지 못합니다. 본 논문에서는 MM-Food-100K 데이터 세트를 기반으로 구축된 포괄적인 벤치마크인 OmniFood-Bench를 소개합니다. 이전 연구와 달리, OmniFood-Bench는 VLM의 세 가지 단계적 역량을 평가합니다. 즉, 기본적인 인지 능력(재료 및 조리 방법), 정량적 추론 능력(섭취량 및 영양 프로파일링) 및 안전이 중요한 상담 능력(질병별 권장 사항)입니다. gpt-5.1, gemini-3-flash, qwen3-vl-8B를 포함한 최첨단 VLM 6개를 평가했습니다. 광범위한 실험 결과는 놀라운 '의미-물리적 간극'을 보여주었습니다. 모델은 요리의 이름을 정확하게 인식하는 데 있어 거의 인간 수준의 정확도를 보이지만, 질량 추정에서는 심각한 오류를 범하고, 고위험 당뇨 환자에게 안전하지 않은 조언을 제공하는 경우가 빈번합니다. 본 연구는 공중 보건 분야에 배치되는 자율 에이전트의 신뢰성에 대한 엄격한 기준을 제시합니다. 코드 및 데이터 세트는 다음 위치에서 사용할 수 있습니다: https://anonymous.4open.science/r/OmniFood-Bench-7D0B
The rapid integration of Large Vision-Language Models (VLMs) into critical infrastructure promises to revolutionize personalized healthcare and dietary management. However, in the domain of food systems, autonomous agents face a unique and persistent challenge: the "Systemic Information Asymmetry" between visual appearance and intrinsic nutritional composition. Existing benchmarks primarily focus on coarse-grained classification tasks, such as food category recognition, which fail to evaluate the intricate reasoning chain required for real-world dietary management -- specifically, the ability to traverse from identifying hidden ingredients to estimating physical mass, and finally synthesizing safety-critical medical advice. In this paper, we introduce OmniFood-Bench, a comprehensive benchmark constructed from the MM-Food-100K dataset. Unlike previous works, OmniFood-Bench evaluates VLMs across three progressive capabilities: Basic Perception (Ingredients & Cooking Methods), Quantitative Reasoning (Portion Size & Nutritional Profiling), and Safety-Critical Advisory (Disease-Specific Recommendations). We evaluate six state-of-the-art VLMs, including gpt-5.1, gemini-3-flash, and qwen3-vl-8B. Our extensive experiments reveal a startling "Semantic-Physical Gap": while models achieve near-human accuracy in naming dishes, they exhibit catastrophic failure in mass estimation and frequently hallucinate benign advice for high-risk diabetic profiles. This work establishes a rigorous standard for trustworthiness in autonomous agents deployed for public health. The code and datasets are available in: https://anonymous.4open.science/r/OmniFood-Bench-7D0B
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.