내가 무엇을 수정할 수 있을까? 개방형 전략 발견과 감정 편집 가능성 연구
What Can I Edit? Open-Ended Strategy Discovery and the Emotion Editability Landscape
감정적인 이미지 편집은 단순히 감정을 나타내는 필터를 적용하거나 미리 정의된 시각적 요소를 변경하는 것 이상을 요구합니다. 효과적인 편집은 특정 이미지가 어떤 방식으로 목표 감정을 표현할 수 있는지 파악해야 합니다. 기존의 감정 기반 이미지 조작 방법, 특히 최근의 지능형 변형 방식들은 대부분 미리 정의된 요소 분류 체계, 지식 라이브러리 또는 일반적인 편집 템플릿에 기반한 제한된 전략 공간 내에서 작동하며, 따라서 이미지별 특이성이나 맥락적 요소를 고려한 전략을 종종 놓치게 됩니다. 본 연구에서는 EmoScope라는 다중 에이전트 프레임워크를 소개합니다. EmoScope는 작업을 "어떻게 편집해야 할까?"가 아닌 "무엇을 수정할 수 있을까?"로 재구성합니다. EmoScope는 먼저 감정에 기반한 가능성 추론을 통해 이미지별 수정 가능한 공간을 발견하고, 콘텐츠 일관성과 감정적 표현 사이의 균형을 유지하기 위해 의미적 계층 구조(앵커, 변수, 맥락)를 사용하며, 편집 실행 및 검증 과정을 거칩니다. EmoScope는 계획이 이미지별 가능성으로 표현되므로, 기존 템플릿 기반 방식과 달리, 사용자가 계획 수준에서 편집 전략을 상호 작용적으로 개선할 수 있는 인터페이스를 제공합니다. 대규모 사용자 평가 결과, 1,824개의 쌍방향 질문에 대한 4,693건의 유효 응답을 바탕으로, 참가자들은 EmoScope가 두 가지 경쟁 모델보다 평균 88.1% 더 높은 선호도를 보였습니다. 추가 분석 결과, EmoScope는 균일한 템플릿을 적용하는 것이 아니라 목표 감정에 적합한 전략을 선택한다는 것을 알 수 있습니다. 또한, 가능성 수준의 계획은 상호 작용적인 시범 연구에서 사용자의 가벼운 수정에도 효과적입니다. 마지막으로, 분류기 기반 지표가 비전형적이거나 맥락에 따른 편집에 대해 감정 조건에 따라 특정 결함을 보이는 현상을 보여주며, 이미지-감정 조합에 따라 EmoScope의 장점이 체계적으로 변하는 콘텐츠-감정 선호도 및 적합성 지도를 제시합니다.
Emotional image editing requires more than applying affective filters or modifying predefined visual factors: an effective edit must identify what a particular image can afford for a target emotion. Existing affective image manipulation methods, including recent agentic variants, largely operate within bounded strategy spaces based on predefined factor taxonomies, knowledge libraries, or conventional editing templates, and therefore often miss image-specific, context-grounded strategies. We introduce EmoScope, a multi-agent framework that reframes the task from "how should I edit?" to "what can I edit?" EmoScope first discovers an image-specific editable space through emotion-conditioned affordance reasoning, then uses a semantic hierarchy of anchors, variables, and context to balance content consistency and emotional expressiveness before executing and verifying the edit. Because its plans are expressed as image-specific affordances rather than retrieved templates, EmoScope also exposes the editing strategy as an interactive surface for user refinement at the plan level. In a large-scale human evaluation covering all eight Mikels emotion categories, with 4,693 valid responses across 1,824 pairwise questions, participants preferred EmoScope over two competitive baselines by 88.1% on average. Attribution analysis further shows that EmoScope selects target-emotion-adaptive strategies rather than applying a uniform template. The same affordance-level plan also supports lightweight user refinement in an interactive pilot. Finally, we show that classifier-based metrics exhibit emotion-conditional blind spots toward non-stereotypical, context-grounded edits, and present a relative content-emotion preference-affinity landscape showing that EmoScope's advantage varies systematically across image-emotion combinations.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.