2608.09316v1 Aug 10, 2026 cs.CV

MemeMind: 참조 기반 추적 생성 시스템을 활용한 오프라인 컨텍스트 최적화

MemeMind: Reference-Guided Trace Construction for Offline Context Optimization

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Qiang Sun
Qiang Sun
Citations: 55
h-index: 5
Yuchen He
Yuchen He
Citations: 20
h-index: 2
Boheng Sheng
Boheng Sheng
Citations: 9
h-index: 1
Longwen Gao
Longwen Gao
Citations: 16
h-index: 2
Jielei Zhang
Jielei Zhang
Citations: 24
h-index: 3

오프라인 컨텍스트 최적화는 모델을 고정한 상태에서 에이전트의 지침과 예시를 수정하여 성능을 향상시키는 방법입니다. 이 접근 방식은 적응 데이터 세트에 대한 시뮬레이션을 통해 학습하지만, 일부 쿼리는 실패한 시뮬레이션 결과만 생성합니다. 이러한 경우, 최적화 도구는 사용 가능한 도구를 사용하여 올바른 답변에 도달하는 성공적인 예시를 얻지 못합니다. 본 논문에서는 오프라인 참조 답변을 활용하여 이러한 누락된 경험을 보완하는 MemeMind라는 시스템을 제안합니다. TraceBuilder는 참조에서 요구되는 증거를 식별하고, 텍스트 검색, 이미지 검색 및 시각적 정렬을 수행하며, 결과적으로 생성된 도구 사용 추적을 검증한 후 적응 버퍼에 추가합니다. ToolGuide는 수집된 추적들을 요약하여 공유 가이드와 각 도구에 대한 개별 지침을 제공합니다. 참조 답변과 생성된 추적은 오프라인 학습 과정에서만 사용되며, 추론 시에는 학습된 가이드를 사용하여 고정된 모델을 활용합니다. 본 연구에서는 애니메이션, 만화 및 게임의 밈 해석이라는 복잡한 문제를 통해 MemeMind를 평가했습니다. 이러한 밈은 편집되고 모호한 시각적 콘텐츠, 텍스트 오버레이, 장기 프랜차이즈 지식 및 특정 문화에 대한 참조 등을 결합합니다. 이러한 밈의 해석에는 조정된 시각적 정렬, 이미지 검색 및 텍스트 검색이 필요하며, 이는 기존 시뮬레이션 그룹이 함께 실패할 수 있는 어려운 환경을 제공합니다. 우리는 전문가가 주석을 단 1,000개의 밈으로 구성된 벤치마크인 MemeX를 사용하여 MemeMind를 평가했습니다. 두 개의 Qwen3-VL 모델, 두 가지 언어 파티션 및 두 명의 독립적인 심사관을 통해, MemeMind는 Qwen3-VL-30B-A3B에서 각각 22.0%와 21.1%, 그리고 Qwen3-VL-235B-A22B (GPT-5 평가 기준)에서 각각 8.1%와 8.0%의 성능 향상을 보였습니다. 분석 결과, 실패한 그룹에 대한 성공적인 도구 사용 추적을 생성하는 것이 가장 큰 성능 개선을 가져오며, 추론 시 더 효과적인 증거 확보를 가능하게 한다는 것을 확인했습니다.

Original Abstract

Offline context optimization improves an agent by revising its instructions and examples while keeping the model frozen. This approach learns from rollouts on an adaptation set, but some queries produce only failed rollouts. In these cases, the optimizer sees no successful example of how the available tools can reach the correct answer. We introduce MemeMind, which uses an offline reference answer to recover this missing experience. TraceBuilder identifies the evidence required by the reference, executes text search, image retrieval, and visual grounding, and verifies the resulting tool trace before adding it to the adaptation buffer. ToolGuide then summarizes the collected traces into a shared guide and separate instructions for each tool. The reference answers and constructed traces are used only during adaptation, while inference uses the learned guides with a frozen model. We study this problem through Anime, Comic, and Game meme interpretation. These memes combine edited and ambiguous visual content, overlaid text, long tail franchise knowledge, and culture specific references. Their interpretation can require coordinated visual grounding, image retrieval, and text search, making them a demanding setting in which native rollout groups may fail together. We evaluate MemeMind on MemeX, a benchmark of 1,000 such memes annotated by experts. Across two Qwen3-VL models, two language partitions, and two independent judges, MemeMind improves over the strongest context optimization baseline by 22.0% and 21.1% on Qwen3-VL-30B-A3B, and by 8.1% and 8.0% on Qwen3-VL-235B-A22B under GPT-5 judging. Ablations and held out traces show that constructing successful tool use for failed groups provides the largest component gain and produces more effective evidence acquisition at inference time.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!