Banana100: 나노 바나나 프로를 활용한 100회 반복 이미지 복제를 통해 NR-IQA 지표를 혁신적으로 개선
Banana100: Breaking NR-IQA Metrics by 100 Iterative Image Replications with Nano Banana Pro
다중 모드 에이전트 시스템의 다단계, 반복적인 이미지 편집 기능은 디지털 콘텐츠 제작 방식을 변화시켰습니다. 최신 이미지 편집 모델은 지시 사항을 충실히 따르고 고품질 이미지를 생성하지만, 다중 단계 편집 과정에서 이미지 품질이 반복적으로 저하되는 중요한 문제점이 존재합니다. 이미지가 반복적으로 편집될수록 미세한 오류들이 누적되어 가시적인 노이즈가 심각하게 증가하고, 간단한 편집 지시 사항을 따르지 못하게 됩니다. 이러한 문제점을 체계적으로 연구하기 위해, 다양한 질감과 이미지 내용을 포함하는 28,000개의 손상된 이미지로 구성된 종합 데이터셋인 Banana100을 소개합니다. 놀랍게도, 이미지 품질 평가자들은 이러한 품질 저하를 감지하지 못합니다. 21개의 인기 있는 참조 이미지 품질 평가(NR-IQA) 지표 중 어느 것도 심하게 손상된 이미지에 대해 깨끗한 이미지보다 낮은 점수를 일관되게 부여하지 않습니다. 생성 모델과 평가 모델 모두의 이러한 실패는 다중 단계 편집으로 생성된 저품질 합성 데이터가 품질 필터를 통과하지 못할 경우, 미래 모델 훈련의 안정성과 배포된 에이전트 시스템의 안전을 위협할 수 있습니다. 보다 강력한 모델 개발을 촉진하고 다중 모드 에이전트 시스템의 취약성을 완화하기 위해 전체 코드와 데이터를 공개합니다.
The multi-step, iterative image editing capabilities of multi-modal agentic systems have transformed digital content creation. Although latest image editing models faithfully follow instructions and generate high-quality images in single-turn edits, we identify a critical weakness in multi-turn editing, which is the iterative degradation of image quality. As images are repeatedly edited, minor artifacts accumulate, rapidly leading to a severe accumulation of visible noise and a failure to follow simple editing instructions. To systematically study these failures, we introduce Banana100, a comprehensive dataset of 28,000 degraded images generated through 100 iterative editing steps, including diverse textures and image content. Alarmingly, image quality evaluators fail to detect the degradation. Among 21 popular no-reference image quality assessment (NR-IQA) metrics, none of them consistently assign lower scores to heavily degraded images than to clean ones. The dual failures of generators and evaluators may threaten the stability of future model training and the safety of deployed agentic systems, if the low-quality synthetic data generated by multi-turn edits escape quality filters. We release the full code and data to facilitate the development of more robust models, helping to mitigate the fragility of multi-modal agentic systems.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.