LJP가 다루지 못하는 사례들: 보다 완전한 형사 책임 평가를 위한 검찰 결정 예측
The Cases LJP Never Sees: Prosecution Decision Prediction for More Complete Criminal Liability Assessment
법률 판결 예측(Legal Judgment Prediction, LJP)은 형사 법률 분야에서 인공지능을 평가하는 핵심 지표로 자리 잡았지만, 이는 이미 검찰의 심사를 거쳐 공식 기소된 사건만을 다룹니다. 결과적으로 LJP는 형사 책임 평가에 상당한 사각지대를 남기며, 증거 부족, 무죄 또는 처벌 면제 대상인 사례들을 간과합니다. 이러한 격차를 해소하기 위해, 본 연구에서는 검찰 심사를 중심으로 구축된 최초의 법률 인공지능 과제인 extbf{검찰 결정 예측(Prosecution Decision Prediction, PDP)}을 제안합니다. PDP은 각 사건을 기소 또는 세 가지 비기소 결정 중 하나로 분류하며, 증거 평가, 법률 적용 및 가치 기반 재량 결정 능력을 반영합니다. 또한, 190개 범죄 유형에 걸쳐 총 4,630건의 실제 중국 검찰 결정을 포함하는 벤치마크인 extbf{PDP-Bench}를 구축했습니다. 광범위한 실험 결과, 최첨단 LLM은 PDP에서 LJP보다 현저히 낮은 성능을 보였으며, 일반적인 성능 향상 기법들은 이러한 격차를 해소하지 못했습니다. 또한, 제어된 RLVR(Reinforcement Learning with Value Reinforcement) 개입 실험 결과, 단순한 결과 기반 보상은 일반화 가능한 PDP 예측 능력을 얻는 데 실패하는 것으로 나타났습니다.
Legal Judgment Prediction (LJP) has become a core benchmark for evaluating AI in the criminal legal domain, but it only sees criminal cases that have already passed prosecutorial review and been formally indicted. As a result, LJP leaves a substantial blind spot in assessing criminal liability, overlooking cases involving insufficient evidence, no criminal liability, or guilt exempted from punishment. To fill this gap, we propose \textbf{Prosecution Decision Prediction (PDP)}, the first Legal AI task built around prosecutorial review, which classifies each case into prosecution or one of three non-prosecution decisions and reflects legal AI's capabilities in evidence evaluation, legal subsumption, and value-based discretion. We further construct \textbf{PDP-Bench}, a benchmark of 4{,}630 real Chinese prosecutorial decisions spanning 190 charges. Extensive experiments show that state-of-the-art LLMs perform substantially worse on PDP than on LJP and that mainstream enhancement routes fail to close the gap. Moreover, controlled RLVR interventions show that simple outcome rewards fail to produce generalizable PDP discrimination.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.