2606.12018v1 Jun 10, 2026 cs.AI

MODF-SIR: 다중 에이전트 기반의 범용 모달 지식 증류 프레임워크를 이용한 사회적 지능 추론

MODF-SIR: A Multi-agent Omni-modal Distilled Framework for Social Intelligence Reasoning

Jisheng Dang
Jisheng Dang
Citations: 119
h-index: 4
Bimei Wang
Bimei Wang
Citations: 27
h-index: 3
Wencan Zhang
Wencan Zhang
Citations: 305
h-index: 7
Hong Peng
Hong Peng
Citations: 84
h-index: 5
Shangyuan Ma
Shangyuan Ma
Citations: 0
h-index: 0
Binbin Hu
Binbin Hu
Citations: 346
h-index: 9
Qi Tian
Qi Tian
Citations: 47
h-index: 3
Tat-Seng Chua
Tat-Seng Chua
Citations: 23
h-index: 2
Yifan Zhang
Yifan Zhang
Citations: 116
h-index: 5

본 논문에서는 경량화된 다중 모드 대규모 언어 모델(MLLM)을 기반으로 구축된 다중 에이전트 협업 프레임워크를 제안하며, 이는 사회적 지능 추론에 특화되어 있습니다. 저희 접근 방식의 핵심은 학습 및 추론 단계 모두에서 지식 증류 기술을 활용한다는 것입니다. 이 아키텍처 내에서 사회적 지능과 관련된 다중 모드 데이터가 정확하게 식별됩니다. 또한, 관련성이 높은 희귀 이벤트(long-tail events)를 식별하고 추출하여 구조화된 명시적인 텍스트 형태로 표현합니다. 이러한 구조화 전략은 토큰화 과정에서 중요한 희귀 정보가 일반적인 이벤트 및 환경적 노이즈에 가려지는 현상을 방지합니다. 특히, 저희는 전체 추론 파이프라인 전반에 걸쳐 테스트 시간 적응(Test-Time Adaptation, TTA)을 통합하여, 희귀 이벤트의 추출 및 표현, 사고 과정(Chain-of-Thought, CoT) 프롬프트, 그리고 자기 성찰 기능을 포함합니다. 이 TTA 메커니즘은 또한 지식 증류를 통해 강화되며, 저랭크 적응(Low-Rank Adaptation, LoRA)을 사용하여 기초 모델을 인스턴스 수준의 추론에만 맞게 미세 조정합니다. 다양한 공개 및 독점 AI 모델에 대한 광범위한 실험 결과는 제안된 프레임워크의 효과를 입증합니다. IntentTrain 데이터셋에서 약 30%의 학습 데이터를 활용하여 최첨단 성능을 달성했습니다. 코드, 데모, LoRA 모델 및 라우터 학습용 데이터셋은 각각 다음 링크에서 확인할 수 있습니다: https://github.com/eeee-sys/MODF-SIR, https://huggingface.co/spaces/Harry-1234/MODF-SIR, https://huggingface.co/Harry-1234/MODF-SIR 및 https://huggingface.co/datasets/Harry-1234/IntentRouterTrain.

Original Abstract

We propose a multi-agent collaborative framework built upon a lightweight Multimodal Large Language Model (MLLM), specifically designed for social intelligence reasoning. A key feature of our approach is that both the training and inference phases are augmented via knowledge distillation. Within this architecture, multi-modal data pertinent to social intelligence is precisely localized. Furthermore, relevant long-tail events are identified, extracted, and rendered as formatted, explicit text. This formatting strategy prevents critical long-tail information from being overshadowed by head events and environmental noise during the tokenization process. Specifically, we integrate Test-Time Adaptation (TTA) across the entire reasoning pipeline, encompassing the extraction and representation of long-tail events, Chain-of-Thought (CoT) prompting, and self-reflection. This TTA mechanism is also distillation-enhanced, utilizing Low-Rank Adaptation (LoRA) to fine-tune the foundation model exclusively for instance-level reasoning. Extensive evaluations against various open-source and proprietary AI models across multiple benchmarks demonstrate the effectiveness of the proposed framework. With around 30% of training data from IntentTrain, we achieve state-of-the-art results. Codes are available at https://github.com/eeee-sys/MODF-SIR, demo is available at https://huggingface.co/spaces/Harry-1234/MODF-SIR, LoRA is available at https://huggingface.co/Harry-1234/MODF-SIR and the dataset for training router is available at https://huggingface.co/datasets/Harry-1234/IntentRouterTrain.

0 Citations
0 Influential
24.5 Altmetric
122.5 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!