2604.02923v1 Apr 03, 2026 cs.CL

의회 모드: 멀티 에이전트 합의를 통한 LLM의 환각 및 편향 완화

Council Mode: Mitigating Hallucination and Bias in LLMs via Multi-Agent Consensus

S. Wu
S. Wu
Citations: 16
h-index: 3
Xue Li
Xue Li
Citations: 15
h-index: 2
Yan Feng
Yan Feng
Citations: 175
h-index: 6
Zhijun Wang
Zhijun Wang
Citations: 51
h-index: 3
Yufang Li
Yufang Li
Citations: 7
h-index: 2
Ran Wang
Ran Wang
Citations: 92
h-index: 3

대규모 언어 모델(LLM), 특히 Mixture-of-Experts(MoE) 아키텍처를 사용하는 모델들은 다양한 자연어 처리 작업에서 놀라운 성능을 보여줍니다. 그러나 이러한 모델들은 종종 환각 현상, 즉 그럴듯하지만 사실과 다른 내용을 생성하는 현상을 겪으며, 추론 과정에서 전문가의 활성화가 불균형하게 이루어지면서 체계적인 편향이 증폭되는 경향이 있습니다. 본 논문에서는 이러한 한계점을 해결하기 위해, 여러 개의 이기종 최첨단 LLM에 쿼리를 동시에 전달하고, 전용 합의 모델을 통해 그 결과를 종합하는 새로운 멀티 에이전트 합의 프레임워크인 '의회 모드'를 제안합니다. 의회 파이프라인은 세 단계로 구성됩니다: (1) 복잡성에 따라 쿼리를 분류하는 지능형 트라이어지 분류기, (2) 다양한 아키텍처를 가진 모델에 대한 전문가 생성, (3) 최종 응답을 생성하기 전에 합의, 불일치, 그리고 고유한 발견 사항을 명시적으로 식별하는 구조화된 합의 종합. 우리는 이 아키텍처를 오픈 소스 AI 환경에서 구현하고 평가했습니다. 여러 벤치마크에 대한 종합적인 평가 결과, '의회 모드'는 HaluEval 벤치마크에서 환각 발생률을 35.9% 상대적으로 감소시키고, TruthfulQA에서는 성능이 가장 뛰어난 개별 모델보다 7.8점 향상되는 것을 확인했습니다. 또한, 다양한 분야에서 편향의 변동성을 현저히 낮추었습니다. 우리는 합의 메커니즘의 수학적 공식, 시스템 아키텍처에 대한 상세 설명, 그리고 다양한 실험 결과와 분석 결과를 제시합니다.

Original Abstract

Large Language Models (LLMs), particularly those employing Mixture-of-Experts (MoE) architectures, have achieved remarkable capabilities across diverse natural language processing tasks. However, these models frequently suffer from hallucinations -- generating plausible but factually incorrect content -- and exhibit systematic biases that are amplified by uneven expert activation during inference. In this paper, we propose the Council Mode, a novel multi-agent consensus framework that addresses these limitations by dispatching queries to multiple heterogeneous frontier LLMs in parallel and synthesizing their outputs through a dedicated consensus model. The Council pipeline operates in three phases: (1) an intelligent triage classifier that routes queries based on complexity, (2) parallel expert generation across architecturally diverse models, and (3) a structured consensus synthesis that explicitly identifies agreement, disagreement, and unique findings before producing the final response. We implement and evaluate this architecture within an open-source AI workspace. Our comprehensive evaluation across multiple benchmarks demonstrates that the Council Mode achieves a 35.9% relative reduction in hallucination rates on the HaluEval benchmark and a 7.8-point improvement on TruthfulQA compared to the best-performing individual model, while maintaining significantly lower bias variance across domains. We provide the mathematical formulation of the consensus mechanism, detail the system architecture, and present extensive empirical results with ablation studies.

1 Citations
0 Influential
3 Altmetric
16.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!