가르칠 수 있는 것 너머의 탐색: 에이전트 기반 시각 생성에서 지식 경계의 진화
Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation
시각 생성 모델은 뛰어난 렌더링 능력을 갖추고 있지만, 자신이 모르는 내용을 과감하게 날조하는 경향이 있습니다. 사용자 요청은 무한하고, 끊임없이 변화하며, 극단적으로 다양합니다: 새로운 캐릭터, 유행하는 개체, 데이터 종료 시점 이후의 사건 등 다양한 내용이 포함됩니다. 이러한 세계 지식의 제한성은 구조적인 문제입니다. 생성 모델은 고정된 데이터 세트를 기반으로 훈련되지만, 실제 시각 세계는 무한히 확장될 수 있기 때문입니다. 우리는 SearchGen-20K와 SearchGen-Bench를 구축했습니다. 이 데이터 세트는 12가지 오류 유형과 22개 도메인을 포괄하는 20,839개의 프롬프트로 구성되어 있으며, 오프라인에서 재현 가능한 연구를 지원하기 위해 미리 실행된 멀티모달 SearchGen-Corpus-1M과 함께 제공됩니다. SearchGen-Bench에서 최첨단 공개 생성 모델은 100점 만점에 21점에서 28점 사이의 낮은 점수를 기록했습니다. 이는 기존 벤치마크에서는 드러나지 않는 40점이나 감소한 수치입니다. 자연스러운 해결책은 검색 도구를 활용하여 에이전트 기반 시각 생성을 가능하게 하는 것입니다. 그러나 단순한 검색 방식은 효과적이지 않습니다. 왜냐하면, 검색은 무차별적으로 정보를 가져와 생성기가 이미 처리할 수 있는 프롬프트에 노이즈를 추가하기 때문입니다. 우리는 이러한 문제의 근본 원인이 생성기 모델마다 고유하고 끊임없이 변화하는 지식 경계 때문이라는 것을 발견했습니다. 즉, 생성기가 훈련을 통해 내부화할 수 있는 것과 외부 컨텍스트로 유지해야 하는 것 사이의 경계가 존재합니다. 이 경계를 사전에 정확하게 정의하기는 어렵지만, 우리는 '가르치고 검색하는' 공동 훈련 프레임워크를 통해 이를 파악할 수 있음을 보여줍니다. 이러한 공동 훈련 방식의 최소 버전이라도 꾸준한 성능 향상을 가져오며, 세계 지식에 기반한 사용자 요청을 충족시킬 수 있는 시각 생성 모델의 자기 개선 능력을 향상시키는 데 기여합니다. 우리는 전체 데이터 세트, 공동 훈련 코퍼스 및 검색 코퍼스를 공개하여 도구 지원 및 세계 지식 기반 시각 생성을 위한 재사용 가능한 환경을 제공합니다.
Visual generators excel at rendering, but they confidently fabricate what they do not know. User requests are unbounded, evolving, and deeply long-tailed: new characters, trending entities, post-cutoff events, and more. This world-knowledge bottleneck is structural: generators are trained on fixed corpora, but the visual world is open-ended. We construct SearchGen-20K and SearchGen-Bench, with 20,839 prompts spanning twelve failure categories and twenty-two domains, paired with a pre-executed multimodal SearchGen-Corpus-1M to support offline, reproducible research. On SearchGen-Bench, frontier open generators score only 21 to 28 out of 100, a 40-point collapse invisible to existing benchmarks. The natural remedy is to employ search tools, enabling agentic visual generation. However, we find that naive search fails: it retrieves indiscriminately, injecting noise into prompts the generator already handles. We trace the root cause to a generator-specific, evolving knowledge boundary: the divide between what a generator can internalize through training and what must remain in external context. Although this boundary is hard to specify in advance, we show that it is discoverable through a teach-then-search co-training framework. Even a minimal version of this co-training recipe produces monotonic improvement, laying the foundation for recursive self-improvement in visual generation that can meet world-knowledge-grounded requests. We release the full dataset, co-training corpus, and search corpus as a replayable harness for tool-augmented, world-knowledge-grounded visual generation.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.