2603.16934v1 Mar 14, 2026 cs.CV

AgriChat: 농업 이미지 이해를 위한 다중 모드 대규모 언어 모델

AgriChat: A Multimodal Large Language Model for Agriculture Image Understanding

Abderrahmene Boudiaf
Abderrahmene Boudiaf
Citations: 39
h-index: 4
Irfan Hussain
Irfan Hussain
Citations: 63
h-index: 3
Sajid Javed
Sajid Javed
Citations: 403
h-index: 10

현재 다중 모드 대규모 언어 모델(MLLM)을 농업 분야에 적용하는 데는 중요한 어려움이 있습니다. 기존 연구에서는 안정적인 모델 개발 및 평가를 위한 대규모 농업 데이터셋이 부족하며, 최첨단 모델은 다양한 분류에 대한 추론이 가능하도록 검증된 전문 지식이 부족합니다. 이러한 문제점을 해결하기 위해, 우리는 Vision-to-Verified-Knowledge (V2VK) 파이프라인을 제안합니다. V2VK는 시각적 캡션을 웹 기반 과학 검색과 통합하는 새로운 생성형 AI 기반 주석 프레임워크로, AgriMM 벤치마크를 자동으로 생성하여, 검증된 식물 병리학 문헌을 기반으로 학습 데이터를 제공함으로써 허구적인 오류를 효과적으로 제거합니다. AgriMM 벤치마크는 3,000개 이상의 농업 분류와 607,000개 이상의 시각적 질문-답변(VQA) 데이터로 구성되어 있으며, 미세한 식물 종 식별, 식물 질병 증상 인식, 작물 수 세기, 숙성도 평가 등 다양한 작업을 포함합니다. 이러한 검증 가능한 데이터를 활용하여, 우리는 광범위한 농업 분류에 대한 지식을 제공하고, 상세한 농업 평가와 함께 광범위한 설명을 제공하는 특화된 MLLM인 AgriChat을 제시합니다. 다양한 작업, 데이터셋 및 평가 조건에서의 광범위한 평가는 현재 농업 MLLM의 기능과 한계를 보여주며, AgriChat이 다른 오픈 소스 모델보다 우수한 성능을 보임을 입증합니다. 이러한 결과는 시각적 세부 정보를 유지하고 웹에서 검증된 지식을 결합하는 것이 안정적이고 신뢰할 수 있는 농업 AI 개발을 위한 효과적인 방법임을 확인합니다. 코드와 데이터셋은 https://github.com/boudiafA/AgriChat 에서 공개적으로 이용할 수 있습니다.

Original Abstract

The deployment of Multimodal Large Language Models (MLLMs) in agriculture is currently stalled by a critical trade-off: the existing literature lacks the large-scale agricultural datasets required for robust model development and evaluation, while current state-of-the-art models lack the verified domain expertise necessary to reason across diverse taxonomies. To address these challenges, we propose the Vision-to-Verified-Knowledge (V2VK) pipeline, a novel generative AI-driven annotation framework that integrates visual captioning with web-augmented scientific retrieval to autonomously generate the AgriMM benchmark, effectively eliminating biological hallucinations by grounding training data in verified phytopathological literature. The AgriMM benchmark contains over 3,000 agricultural classes and more than 607k VQAs spanning multiple tasks, including fine-grained plant species identification, plant disease symptom recognition, crop counting, and ripeness assessment. Leveraging this verifiable data, we present AgriChat, a specialized MLLM that presents broad knowledge across thousands of agricultural classes and provides detailed agricultural assessments with extensive explanations. Extensive evaluation across diverse tasks, datasets, and evaluation conditions reveals both the capabilities and limitations of current agricultural MLLMs, while demonstrating AgriChat's superior performance over other open-source models, including internal and external benchmarks. The results validate that preserving visual detail combined with web-verified knowledge constitutes a reliable pathway toward robust and trustworthy agricultural AI. The code and dataset are publicly available at https://github.com/boudiafA/AgriChat .

1 Citations
0 Influential
28.4657359028 Altmetric
6.9 Score
Original PDF
1

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!