PVminerLLM2: 선호도 최적화를 통한 환자 음성 구조화 추출 성능 향상
PVminerLLM2: Improving Structured Extraction of Patient Voice via Preference Optimization
배경: 환자가 생성한 텍스트는 환자의 실제 경험, 사회적 맥락 및 치료 참여에 대한 중요한 정보를 담고 있지만, 대부분 비정형화되어 있어 환자 중심의 결과 연구에 활용하기 어렵습니다. 이전 연구에서는 PV-Miner 벤치마크와 PVMinerLLM 모델을 사용하여 구조화된 추출 방법을 제시했습니다. 그러나 지도 학습 파인튜닝만으로는 희귀하고 미세하며 불균등하게 분포된 오류, 특히 토큰 수준에서 중요한 구조적 출력에 대한 문제를 해결하기 어렵습니다. 결과: 우리는 구조화된 환자 음성 추출을 위한 개선된 LLM 집합인 PVminerLLM2를 제시합니다. 이는 지도 학습 파인튜닝의 한계를 극복하고 토큰 수준에서 중요한 오류를 해결하기 위해 선호도 최적화를 적용했습니다. 저희 방법은 (i) 절대적인 토큰 가능성의 저하를 방지하는 토큰 수준의 게이티드 안정화 항을 포함하는 선호도 목표, 그리고 (ii) 낮은 구별력을 더 잘 포착하기 위한 혼동 인식 선호 페어 구축을 도입합니다. 또한 토큰 불균형 및 클래스 편향 문제를 해결하기 위해 토큰 중요도 가중치와 역빈도 재가중치를 적용했습니다. 다양한 모델 크기에서 PVminerLLM2는 강력한 기준 모델보다 일관되게 높은 성능을 보이며, 최대 4.43% (Code), 3.50% (Sub-code) 및 1.55% (Span)의 성능 향상을 달성했으며, 기존 선호도 최적화 방법을 사용하여 학습된 기본 LLM보다 우수한 성능을 보였습니다. 가용성 및 구현: PVminerLLM2에 대한 추가 자료, 코드, 평가 스크립트 및 학습 모델은 다음에서 공개적으로 이용할 수 있습니다: https://github.com/Data-Mining-Lab-Yale/PVminerLLM2
Motivation: Patient-generated text contains critical information on patients' lived experiences, social context, and care engagement, but remains largely unstructured, limiting its use in patient-centered outcomes research. Prior work introduced the PV-Miner benchmark and PVMinerLLM models for structured extraction. However, supervised fine-tuning (SFT) alone struggles with rare, fine-grained, and unevenly distributed errors, particularly in token-critical structured outputs. Results: We present PVminerLLM2, an improved set of LLMs for structured patient voice extraction that applies preference optimization to address token-critical errors beyond the reach of supervised fine-tuning. Our method introduces (i) a preference objective with token-level gated stabilization term that prevents degradation of absolute token likelihood under preference optimization, and (ii) confusion-aware preference pair construction to better capture low-separation distinctions. We further incorporate token-importance weighting and inverse-frequency reweighing to address token imbalance and class skew. Across multiple model sizes, PVMinerLLM2 consistently outperforms strong baselines, achieving gains of up to 4.43% (Code), 3.50% (Sub-code), and 1.55% (Span), and outperforms baseline LLM trained with existing preference optimization methods. Availability and Implementation: The supplementary material, code, evaluation scripts, and trained models for PVminerLLM2 are publicly available at: https://github.com/Data-Mining-Lab-Yale/PVminerLLM2
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.