초기 아동 교육 분야의 일상 활동 이미지 설명: 벤치마크 및 알고리즘
Captioning Daily Activity Images in Early Childhood Education: Benchmark and Algorithm
초기 아동 교육(ECE) 분야의 이미지 설명을 통해 활동 이해 및 교육 평가를 자동화하는 것은 매우 중요합니다. 그러나 기존 방법은 다음과 같은 두 가지 주요 어려움에 직면합니다. 첫째, 대규모의 특정 분야 데이터셋 부족은 모델이 ECE 시나리오에서 나타나는 미세한 의미적 개념을 정확하게 파악하는 능력을 제한하여, 일반적이고 부정확한 설명을 생성하게 만듭니다. 둘째, 기존의 훈련 방식은 전문적인 객체 설명 능력을 향상시키는 데 한계를 보입니다. 지도 학습은 흔히 사용되는 표현을 선호하는 경향이 있으며, 강화 학습은 어려운 샘플에서 불안정한 최적화 문제를 겪을 수 있습니다. 이러한 한계를 해결하기 위해, 우리는 256,121개의 실제 이미지를 포함하는 대규모 ECE 일상 활동 이미지 설명 벤치마크인 ECAC을 제안합니다. ECAC은 전문가 수준의 설명과 미세한 라벨로 주석이 달려 있으며, 전문적인 객체 명칭 정확도를 명시적으로 측정하기 위한 도메인 지향적인 평가 프로토콜인 Teaching Toy Recognition Score (TTS)를 함께 제공합니다. 또한, 우리는 강화 학습과 지도 학습 미세 조정을 동적으로 번갈아 수행하는 하이브리드 훈련 프레임워크인 RSRS (Reward-Conditional Switch of Reinforcement Learning and Supervised Fine-Tuning)를 제안합니다. RSRS는 보상이 없는 어려운 샘플을 지도 학습 미세 조정으로 리라우팅하여, 장점 붕괴를 효과적으로 완화하고 미세한 인식에 대한 안정적인 최적화를 가능하게 합니다. ECAC과 RSRS를 활용하여, 우리는 도메인에 특화된 멀티모달 대규모 언어 모델인 KinderMM-Cap-3B를 개발했습니다. 광범위한 실험 결과, 우리 모델은 TTS에서 51.06의 성능을 달성하여 최첨단 모델보다 훨씬 뛰어난 성능을 보였으며, 우수한 설명 품질을 유지하여 특정 교육 분야에 적용될 가능성을 보여줍니다.
Image captioning for Early Childhood Education (ECE) is essential for automated activity understanding and educational assessment. However, existing methods face two key challenges. First, the lack of large-scale, domain-specific datasets limits the model's ability to capture fine-grained semantic concepts unique to ECE scenarios, resulting in generic and imprecise descriptions. Second, conventional training paradigms exhibit limitations in enhancing professional object description capability, as supervised learning tends to favor high-frequency expressions, while reinforcement learning may suffer from unstable optimization on difficult samples. To address these limitations, we introduce ECAC, a large-scale benchmark for ECE daily activity image captioning, comprising 256,121 real-world images annotated with expert-level captions and fine-grained labels. ECAC is further equipped with a domain-oriented evaluation protocol, the Teaching Toy Recognition Score (TTS), to explicitly measure professional object naming accuracy. Furthermore, we propose RSRS (Reward-Conditional Switch of Reinforcement Learning and Supervised Fine-Tuning), a hybrid training framework that dynamically alternates between RL and supervised optimization. By rerouting hard samples with zero rewards to supervised fine-tuning, RSRS effectively mitigates advantage collapse and enables stable optimization for fine-grained recognition. Leveraging ECAC and RSRS, we develop KinderMM-Cap-3B, a domain-adapted multimodal large language model. Extensive experiments demonstrate that our model achieves a TTS of 51.06, substantially outperforming state-of-the-art baselines while maintaining superior caption quality, highlighting its potential for specialized educational applications.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.