2605.26530v1 May 26, 2026 cs.AI

어떤 변화가 중요한가? 관련성 기반 평가 및 솔버 기반 추론을 통한 신뢰할 수 있는 법률 AI 개발

Which Changes Matter? Towards Trustworthy Legal AI via Relevance-Sensitive Evaluation and Solver-Grounded Reasoning

Yufan Cai
Yufan Cai
Citations: 90
h-index: 5
Z. Hóu
Z. Hóu
Citations: 11
h-index: 2
Linze Chen
Linze Chen
Citations: 6
h-index: 2
Jin Song Dong
Jin Song Dong
Citations: 145
h-index: 6

법률적 판단은 중요하지 않은 변화와 중요한 변화를 구별하는 것을 요구합니다. 법률 AI는 법적으로 무관한 변화에는 안정적이어야 하지만, 법적으로 의미 있는 부분을 변경하는 변화에는 대응해야 합니다. 우리는 이러한 요구 사항을 법적 관련성을 고려한 평가 문제로 정의하며, LLM은 오직 법적으로 관련된 변화에만 민감해야 합니다. 본 연구에서는 사법 공정성, 견고성, 그리고 법률 해석 오류 시나리오를 포괄하는 통합적인 평가 도구를 제시합니다. 실험 결과, 기존의 법률 LLM은 법적으로 무관한 변형에 체계적으로 민감하게 반응하며, 종종 관련된 법적 요소 및 규정을 구별하지 못하는 것을 확인했습니다. 이러한 문제점을 해결하기 위해, 우리는 형식 논리를 기반으로 하는 적대적 다중 에이전트 프레임워크인 LexGuard를 제안합니다. LexGuard는 규정을 실행 가능한 제약 조건으로 공식화하고, 적대적인 에이전트를 사용하여 상반되는 사실-규정 주장을 추출하며, SMT 솔버를 사용하여 법적 만족도 및 논리적 일관성을 검증합니다. 실험 결과, LexGuard는 조작적인 프레임에 대한 취약점을 줄이고, 유사한 규정 간의 모호성을 개선하고, 법적으로 무관한 속성의 영향을 제한하고, 온건한 재구성 하에서의 일관성을 높여 법률적 추론의 신뢰성을 향상시킴을 보여줍니다. 본 연구는 법률 AI의 신뢰성은 정확성뿐만 아니라 법적으로 중요한 변화에 대한 적절한 민감도를 갖추는 데 달려 있음을 입증합니다.

Original Abstract

Legal reasoning requires distinguishing changes that matter from those that do not. Legal AI should remain stable under legally irrelevant perturbations, but should change when perturbations alter legally material points. We formulate this requirement as a legal-relevance-sensitive evaluation problem: LLMs should only be sensitive to the legally relevant change. We introduce a unified evaluation suite covering should-change and should-not-change evaluation across judicial fairness, robustness, and statute-confusion scenarios. Our evaluation shows that existing legal LLMs are systematically sensitive to legally irrelevant variations and often fail to distinguish related legal elements and statutory rules. To mitigate these failures, we present LexGuard, an adversarial multi-agent framework grounded in formal reasoning. LexGuard formalizes statutes into executable constraints, uses adversarial agents to extract competing fact-statute arguments, and invokes SMT solvers to verify legal satisfaction and logical consistency. Experiments show that LexGuard improves legal reasoning reliability by reducing vulnerability to manipulative framing, improving disambiguation among similar statutes, limiting the influence of legally irrelevant attributes, and increasing consistency under benign reformulations. We show that legal trustworthiness requires not only accuracy, but calibrated sensitivity to legally material changes.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!