2602.19843v1 Feb 23, 2026 cs.SE

MAS-FIRE: LLM 기반 다중 에이전트 시스템을 위한 결함 주입 및 신뢰성 평가

MAS-FIRE: Fault Injection and Reliability Evaluation for LLM-Based Multi-Agent Systems

Yingqi Wang
Yingqi Wang
Citations: 5
h-index: 1
Zibin Zheng
Zibin Zheng
Citations: 17
h-index: 3
Jin Jia
Jin Jia
Citations: 12
h-index: 2
Zhiling Deng
Zhiling Deng
Citations: 5
h-index: 1
Zhuangbin Chen
Zhuangbin Chen
Citations: 26
h-index: 4

LLM 기반 다중 에이전트 시스템(MAS)이 복잡한 작업에 점점 더 많이 배포됨에 따라, 이들의 신뢰성을 보장하는 것이 시급한 과제가 되었다. MAS는 엄격한 프로토콜이 아닌 비정형 자연어를 통해 조정되기 때문에, 런타임 예외를 발생시키지 않고 조용히 전파되는 의미론적 오류(예: 환각, 지시 오해, 추론 이탈 등)에 취약하다. 종단 간(end-to-end) 작업 성공 여부만을 측정하는 기존의 평가 방식은 이러한 오류가 어떻게 발생하는지, 혹은 에이전트가 이를 얼마나 효과적으로 복구하는지에 대해 제한적인 통찰만을 제공한다. 이러한 격차를 해소하기 위해, 우리는 MAS의 결함 주입 및 신뢰성 평가를 위한 체계적인 프레임워크인 MAS-FIRE를 제안한다. 우리는 에이전트 내부의 인지 오류와 에이전트 간의 조정 실패를 아우르는 15가지 결함 유형의 분류 체계를 정의하고, 프롬프트 수정, 응답 재작성, 메시지 라우팅 조작이라는 세 가지 비침습적 메커니즘을 통해 이를 주입한다. 세 가지 대표적인 MAS 아키텍처에 MAS-FIRE를 적용하여, 우리는 메커니즘, 규칙, 프롬프트, 추론의 네 가지 계층으로 분류되는 다양한 내결함성 동작들을 발견했다. 이러한 계층적 관점은 시스템이 어디서, 왜 성공하거나 실패하는지에 대한 세밀한 진단을 가능하게 한다. 우리의 연구 결과는 더 강력한 파운데이션 모델이 일률적으로 견고성을 향상시키는 것은 아님을 밝혀냈다. 또한 아키텍처 토폴로지 역시 그에 못지않게 결정적인 역할을 한다는 것을 보여주며, 반복적이고 닫힌 루프(closed-loop) 설계는 선형 워크플로우에서 파국적 붕괴를 일으키는 결함의 40% 이상을 무력화하는 것으로 나타났다. MAS-FIRE는 다중 에이전트 시스템을 체계적으로 개선하는 데 필요한 프로세스 수준의 관찰 가능성과 실행 가능한 지침을 제공한다.

Original Abstract

As LLM-based Multi-Agent Systems (MAS) are increasingly deployed for complex tasks, ensuring their reliability has become a pressing challenge. Since MAS coordinate through unstructured natural language rather than rigid protocols, they are prone to semantic failures (e.g., hallucinations, misinterpreted instructions, and reasoning drift) that propagate silently without raising runtime exceptions. Prevailing evaluation approaches, which measure only end-to-end task success, offer limited insight into how these failures arise or how effectively agents recover from them. To bridge this gap, we propose MAS-FIRE, a systematic framework for fault injection and reliability evaluation of MAS. We define a taxonomy of 15 fault types covering intra-agent cognitive errors and inter-agent coordination failures, and inject them via three non-invasive mechanisms: prompt modification, response rewriting, and message routing manipulation. Applying MAS-FIRE to three representative MAS architectures, we uncover a rich set of fault-tolerant behaviors that we organize into four tiers: mechanism, rule, prompt, and reasoning. This tiered view enables fine-grained diagnosis of where and why systems succeed or fail. Our findings reveal that stronger foundation models do not uniformly improve robustness. We further show that architectural topology plays an equally decisive role, with iterative, closed-loop designs neutralizing over 40% of faults that cause catastrophic collapse in linear workflows. MAS-FIRE provides the process-level observability and actionable guidance needed to systematically improve multi-agent systems.

8 Citations
0 Influential
2 Altmetric
18.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!