지식은 어디에서 주입되어야 하는가? 다중 모드 반복 생성 모델에서의 계층적 지식 주입 프레임워크
Where Should Knowledge Enter? A Layered Framework for Knowledge Infusion in Multimodal Iterative Generative Mo
다중 모드 생성 모델은 유창한 결과를 생성하지만, 구조화된, 도메인 특유의 또는 안전에 중요한 지식을 준수해야 할 때 신뢰성이 떨어지는 경향이 있습니다. 기존 방법들은 프롬프트 증강, 가이드, 잠재 공간 편집 또는 미세 조정과 같은 메커니즘을 통해 지식을 통합하지만, 이러한 방법은 일반적으로 사용하는 기술 자체로 분류되는 반면, 생성 프로세스의 어떤 구성 요소를 수정하는지에 따라 분류되지는 않습니다. 우리는 반복적 생성 모델에서 지식 주입이 근본적으로는 '개입 계층' 문제라고 주장합니다. 생성 프로세스는 내부 상태의 경로로 전개되기 때문에, 지식은 이 프로세스의 네 가지 구조적으로 구별되는 구성 요소에 영향을 미칠 수 있습니다: 입력/출력 경계, 변환 함수, 중간 상태 및 모델 파라미터. 이는 표면(surface), 궤적(trajectory), 잠재 공간(latent) 및 파라미터(parametric) 주입의 네 가지 개입 계층으로 연결됩니다. 우리는 이 프레임워크를 확산 모델에 적용하고, 대표적인 방법들을 네 가지 계층에 매핑하며, 다중 계층 조합을 위한 설계 원칙을 도출합니다. 두 개의 확산 기반 모델을 사용하는 다중 모드 지식 그래프를 활용한 안전 정렬 실험에서, 우리는 표면(입력 측 및 출력 측)과 궤적-잠재 공간(생성 과정 중간)을 포함하여 네 가지 계층 중 세 가지를 누적적으로 구현했습니다. 실험 결과, 각 추가 계층은 이전 계층으로는 해결할 수 없는 오류 유형을 해결하며, 일반적인 생성 방식에 비해 지식 위반 결과를 70.97% 줄이는 것을 경험적으로 확인했습니다. 이는 프레임워크의 상호 보완성 예측을 실증적으로 뒷받침합니다.
Multimodal generative models produce fluent outputs but remain unreliable when generation must respect structured, domain-specific, or safety-critical knowledge. Existing methods incorporate knowledge through mechanisms such as prompt augmentation, guidance, latent editing, or fine-tuning, yet they are typically categorized by technique rather than by the component of the generative process they modify. We argue that knowledge infusion in iterative generative models is fundamentally anintervention-layer problem. Since thegenerative process unfolds as a trajectory of internal states, knowledge can act on four structurally distinct components of this process: the input/output boundary, the transition function, the intermediate state, and the model parameters. This maps to four intervention layers: surface, trajectory, latent, and parametric infusion. We instantiate the framework in diffusion models, map representative methods to all four layers, and derive design principles for multi-layer composition. In a controlled safety-alignment experiment using a multimodal knowledge graph with two diffusion backbones, we implement three of the four layers cumulatively, surface (input-side and output-side) and trajectory--latent (mid-generation). We show empirically that each additional layer addresses failure classes that prior layers cannot reach, reducing knowledge-violating outputs by 70.97% compared to vanilla generation and empirically confirming the framework's complementarity prediction.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.