추론 깊이에서 추론 폭으로: 대규모 언어 모델의 다중 지점 연관 추론 평가
From Reasoning Depth to Reasoning Breadth: Evaluating Multi-Point Associative Reasoning in Large Language Models
대규모 언어 모델(LLM)은 점점 더 길고 복잡한 추론 과정을 요구하는 작업에서 상당한 발전을 이루었습니다. 이러한 발전은 주로 추론의 깊이와 관련이 있습니다. 보완적이고 상대적으로 덜 연구된 능력은 추론의 폭입니다. 이는 여러 의미 방향을 병렬로 탐색하고, 그 결과 얻은 단서들을 하나의 일관된 답변으로 통합하는 것을 의미합니다. 본 논문에서는 다중 지점 연관 추론을 통해 추론의 폭을 평가하기 위한 영어-중국어 이중 언어 벤치마크인 MPAR-Bench를 소개합니다. 'Just One'이라는 협력 게임에서 영감을 받아, 각 항목은 모델에게 여러 개의 독립적으로 생성되고 의미적으로 다양한 단서로부터 숨겨진 목표를 추론하도록 요구합니다. 우리는 다중 에이전트 기반의 단서 생성 파이프라인, 임베딩 기반 다양성 필터링 및 인간 검증을 통해 1,000개의 항목을 구성했습니다. 답변 공간은 공개된 어휘 목록에서 가져온 반면, 모든 단서 집합은 처음부터 생성되었습니다. 정확도와 함께, 우리는 모델의 성능을 정확도, ANLS(Answer Non-Logical Sequence), 임베딩 유사성, 추론 과정 검증 및 네 가지 유형의 교란(단서 마스킹, 순서 변경, 주의 산만 삽입, 다단계 단서)을 사용하여 평가했습니다. 평가된 모델에서, 교란은 영어에서 9~18%p, 중국어에서 5~12%p만큼 정확도를 감소시켰습니다. 사고 모드는 표준 설정 정확도를 향상시키지만, 특히 영어에서 교란에 대한 민감도를 일관되게 줄이지는 않습니다. 사례별 분석 결과, 확장된 추론은 처음에는 올바른 가설을 뒤집을 수 있음을 보여줍니다. 이러한 결과는 더 깊은 추론 능력이 반드시 강력한 추론의 폭을 보장하지 않으며, 현재 벤치마크로는 추론의 폭이 여전히 충분히 평가되지 않고 있다는 것을 시사합니다.
Large language models (LLMs) have made substantial progress on reasoning tasks that require increasingly long and complex inferential chains. This progress primarily reflects reasoning depth. A complementary and comparatively unexamined capability is reasoning breadth: exploring multiple semantic directions in parallel and integrating the resulting clues into one coherent answer. We introduce MPAR-Bench, a bilingual English-Chinese benchmark that isolates reasoning breadth through multi-point associative reasoning. Inspired by the cooperative game Just One, each item asks a model to recover a hidden target from several independently generated, semantically diverse clues. We construct 1,000 items using a multi-agent clue-generation pipeline, embedding-based diversity filtering, and human verification. Only the answer space is drawn from public word lists, whereas every clue set is generated from scratch. Beyond exact-match accuracy, we evaluate models using accuracy, ANLS, embedding similarity, reasoning-trace verification, and four perturbations: clue masking, order shuffling, distractor injection, and multi-step clues. Across evaluated models, perturbations reduce accuracy by 9-18 percentage points in English and 5-12 percentage points in Chinese. Thinking mode improves standard-setting accuracy, especially in English, but does not consistently reduce sensitivity to perturbations. Case-level analysis also shows that extended reasoning can overturn an initially correct hypothesis. These results indicate that greater reasoning depth does not automatically confer robust reasoning breadth, and that reasoning breadth remains largely uncovered by current benchmarks.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.