검색, 실패, 복구: 교정 인지 추론을 위한 학습 프레임워크
Search, Fail, Recover: A Training Framework for Correction-Aware Reasoning
많은 추론 작업은 단일한 좌우 방향의 연결로 잘 설명되지 않습니다. 솔버는 실행 가능한 경로를 탐색해야 할 수도 있으며, 지연된 실패를 관찰하고 아직 완료될 수 있는 최신 접두사로 돌아와야 합니다. 우리는 Diligent Learner 프레임워크에서 영감을 받은 Pyligent이라는 학습 및 추론 프레임워크를 소개합니다. Pyligent은 추론을 부분적인 해결 경로에 대한 검증된 검색으로 표현합니다. 작업 검증기는 생성된 연속과 실패를 레이블링하며, 결과적으로 생성된 검색 트리는 세 가지 액션(계속, 완료, 되돌리기)에 대한 지도 학습 목표로 변환됩니다. 선택적으로 폐기된 분기를 요약하는 추적 정보도 제공됩니다. 우리는 Pyligent을 지연된 실패 복구를 분리하도록 설계된 숨겨진 방향 그래프 작업과 정확한 검증기가 있는 구조화된 추론 영역에서 평가했습니다. 여기에는 $4{ imes}4$ 스도쿠, 추론 트레이스가 포함된 스도쿠 및 Blocksworld가 포함됩니다. Pyligent은 금 표준(gold standard) 기반의 지도 학습 미세 조정과 비교하여 숨겨진 그래프에서 해결률을 72.7%p 향상시키고, 혼합 및 전문가 수준의 스도쿠에서 각각 17%p와 18%p 향상시키고, 추론 트레이스가 포함된 혼합 및 전문가 수준의 스도쿠에서 각각 27%p와 14%p 향상시키며, Blocksworld에서 13%p 향상시켰습니다. 이러한 결과는 명시적인 실패 분기 감독이 완성된 해결 경로 모방을 넘어 유용한 복구 행동을 학습시키는 데 도움이 될 수 있음을 시사합니다.
Many reasoning tasks are not well described by a single left-to-right chain: a solver may need to pursue a plausible branch, observe delayed failure, and return to the latest prefix that can still be completed. We introduce Pyligent, a training and inference framework inspired by the Diligent Learner formulation that represents reasoning as validated search over partial solution chains. A task validator labels generated continuations and failures, and the resulting search trees are converted into supervised targets for three actions: continue, finish, and backtrack, with optional traces that summarize abandoned branches. We evaluate Pyligent on a hidden directed graph task designed to isolate delayed-failure recovery, and on structured reasoning domains with exact validators, including $4{\times}4$ Sudoku, Sudoku with reasoning traces, and Blocksworld. Compared with gold-only supervised fine-tuning, Pyligent improves solve rate by $72.7$ percentage points on hidden graphs, by $17$ and $18$ points on mixed and expert Sudoku, by $27$ and $14$ points on mixed and expert Sudoku with reasoning traces, and by $13$ points on Blocksworld. These results suggest that explicit failed-branch supervision can teach useful recovery behavior beyond imitation of polished solution chains.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.