AI 기반 모의 법정: 구두 변론 시 법관의 특정 질문 패턴 시뮬레이션
AI-Assisted Moot Courts: Simulating Justice-Specific Questioning in Oral Arguments
구두 변론 과정에서 재판관들은 사실 관계, 법률 주장, 그리고 변호사의 주장의 타당성에 대해 질문을 던집니다. 이러한 질문에 대비하기 위해 법학대학 및 실무 변호사들은 모의 법정, 즉 항소 심리 연습을 위한 시뮬레이션을 활용합니다. 본 연구에서는 미국 대법원 구두 변론 기록 데이터를 활용하여, AI 모델이 모의 법정 훈련을 위해 법관의 특정 질문 패턴을 효과적으로 시뮬레이션할 수 있는지 조사합니다. 구두 변론 시뮬레이션의 평가는, 특정 상황에 대해 단 하나의 정답 질문이 없기 때문에 어렵습니다. 대신, 효과적인 질문은 실질적인 법률 문제 예측, 논리적 약점 파악, 그리고 적절한 대립적인 어조 유지와 같은 바람직한 특성을 반영해야 합니다. 본 연구에서는 시뮬레이션된 질문의 현실감과 교육적 유용성을 평가하기 위한 두 계층의 평가 프레임워크를 제시하며, 상호 보완적인 지표를 활용합니다. 프롬프트 기반 및 에이전트 기반 구두 변론 시뮬레이터를 구축하고 평가합니다. 연구 결과, 시뮬레이션된 질문은 인간 평가자가 종종 현실적으로 인식하며, 실제 법률 문제에 대한 높은 수준의 이해도를 보여줍니다. 그러나 모델은 여전히 질문 유형의 다양성 부족, 지나치게 아첨하는 경향 등 상당한 한계를 가지고 있습니다. 특히, 이러한 한계는 단순한 평가 방법을 사용할 경우 감지되지 않을 수 있습니다.
In oral arguments, judges probe attorneys with questions about the factual record, legal claims, and the strength of their arguments. To prepare for this questioning, both law schools and practicing attorneys rely on moot courts: practice simulations of appellate hearings. Leveraging a dataset of U.S. Supreme Court oral argument transcripts, we examine whether AI models can effectively simulate justice-specific questioning for moot court-style training. Evaluating oral argument simulation is challenging because there is no single correct question for any given turn. Instead, effective questioning should reflect a combination of desirable qualities, such as anticipating substantive legal issues, detecting logical weaknesses, and maintaining an appropriately adversarial tone. We introduce a two-layer evaluation framework that assesses both the realism and pedagogical usefulness of simulated questions using complementary proxy metrics. We construct and evaluate both prompt-based and agentic oral argument simulators. We find that simulated questions are often perceived as realistic by human annotators and achieve high recall of ground truth substantive legal issues. However, models still face substantial shortcomings, including low diversity in question types and sycophancy. Importantly, these shortcomings would remain undetected under naive evaluation approaches.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.