LLM에게 그래프를 읽게 하지 마세요: 그래프가 생각하게 하세요
Don't Make the LLM Read the Graph: Make the Graph Think
본 연구에서는 명시적인 신념 그래프가 협력적 다중 에이전트 추론에서 LLM의 성능을 향상시키는지 조사합니다. 협력 카드 게임인 Hanabi를 활용하여 4가지 LLM 패밀리에 걸쳐 3,000건 이상의 통제된 실험을 진행한 결과, 다음과 같은 4가지 주요 결과를 얻었습니다. 첫째, 통합 아키텍처는 신념 그래프가 제공하는 가치를 결정합니다. 프롬프트 컨텍스트로 사용될 경우, 그래프는 강력한 모델에게는 부가적인 요소일 뿐이며, 약한 모델에게만 유익한 경우가 많습니다(2차 수준의 마음 이론, 80% vs 10%, p<0.0001, OR=36.0). 반면, 그래프가 순위가 매겨진 목록을 통해 행동 선택을 제어할 때, 그래프는 강력한 모델에게도 구조적으로 필수적인 요소가 됩니다(2차 수준의 마음 이론, 100% vs 20%, p<0.001). 둘째, 모델 패밀리별 특정 실패 현상인 "계획 거부(Planner Defiance)"를 확인했습니다. LLM이 부분적인 역량을 갖추고 있음에도 불구하고, 올바른 계획 추천을 무시하는 현상인데, 이는 약 90%의 비율로 나타났습니다(N=20). Gemini 모델은 거의 거부 현상을 보이지 않는 반면, Llama 70B는 90%의 거부율을 보이며, 모델들은 사실적인 정보(참조)와 조언(무시)을 구분합니다. 셋째, 전체 게임 데이터를 분석한 결과, 에이전트 간의 규칙(+128% vs. 기준선, p=0.003)이 모든 단일 에이전트 개입보다 우수하며, 신념 그래프의 개별 구성 요소들을 결합해야만 성능 향상을 얻을 수 있습니다. 넷째, 초기 확장 분석(N=10/셀, 탐색적) 결과, 그래프 깊이가 성능 향상에 기여하는 정도가 감소하는 경향을 보입니다. 얕은 그래프가 가장 좋은 비용-효율 비율을 제공하며, 깊은 마음 이론 그래프는 플레이어 수가 증가할수록 오히려 해로울 수 있습니다(5인 게임에서 -1.5점, p=0.029).
We investigate whether explicit belief graphs improve LLM performance in cooperative multi-agent reasoning. Through 3,000+ controlled trials across four LLM families in the cooperative card game Hanabi, we establish four findings. First, integration architecture determines whether belief graphs provide value: as prompt context, graphs are decorative for strong models and beneficial only for weak models on 2nd-order Theory of Mind (80% vs 10%, p<0.0001, OR=36.0); when graphs gate action selection through ranked shortlists, they become structurally essential even for strong models (100% vs 20% on 2nd-order ToM, p<0.001). Second, we identify "Planner Defiance," a model-family-specific failure where LLMs override correct planner recommendations at partial competence (90% override, replicated N=20); Gemini models show near-zero defiance while Llama 70B shows 90%, and models distinguish factual context (deferred to) from advisory recommendations (overridden). Third, full-game evidence confirms inter-agent conventions (+128% over baseline, p=0.003) outperform all single-agent interventions, and individual belief-graph components must be combined to produce gains. Fourth, preliminary scaling analysis (N=10/cell, exploratory) suggests graph depth has diminishing returns: shallow graphs provide the best cost-benefit ratio, while deeper ToM graphs appear harmful at larger player counts (-1.5 pts at 5-player, p=0.029).
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.