시간적 탐색 및 추론을 활용하여 미래 예측을 위한 LLM 진화: 하니스 기반 효율적인 데이터 합성
Distilling Temporal Search and Reasoning: Evolving LLMs for Future Prediction via Harness-Assisted Efficient Data Synthesis
미래 사건 예측은 사회적으로 큰 영향을 미치지만 여전히 어려운 과제입니다. 최첨단 방법들은 외부 에이전트 프레임워크를 사용하여 LLM을 강화하지만, 이러한 프레임워크가 제거되면 예측 능력이 사라집니다. 최근의 Tool-Integrated Reasoning (TIR) 기술은 다중 홉 검색을 통해 사실을 내부적으로 활용하지만, 정확한 예측을 위해서는 과거 추세와 동적인 변화에 대한 시간적 탐색 및 추론이 더욱 중요합니다. 핵심적인 문제는 데이터입니다. 과거의 질의는 시간적 누수를 유발하여 예측 능력을 검색 능력으로 저하시킵니다. 기존 연구들은 정적인 관찰로 정보 수집을 고정하거나, 거부 샘플링 또는 해결되지 않은 새로운 질의에 의존하여 방대한 양의 데이터를 버려 효율성을 떨어뜨립니다. 본 논문에서는 시간적 절단(time-truncation) 하니스를 제안합니다. 이 하니스는 매 단계마다 시간적 제한을 적용하여 TIR 스타일의 방식으로 과거 사건에서 샘플링하도록 유도하고, 시간적 누수를 줄이며 거부 샘플링 또는 해결되지 않은 질의에 대한 의존성을 낮추어 샘플링 효율성을 향상시킵니다. 또한 대규모 데이터셋과 프로세스 기반 지표를 구축하여 제안하는 하니스가 검색 범위를 넓히고 고품질 데이터의 비율을 높여 효율성을 더욱 향상시키며 복잡한 평가 기준에 대한 의존도를 줄인다는 것을 보여줍니다. 실험 결과, 하니스를 통해 생성된 데이터를 사용하여 학습된 모델이 가장 뛰어난 성능을 보였으며, 이는 하니스 기반 모델 진화가 고품질의 시간적 탐색 및 추론 데이터를 활용하여 학생 모델의 성능을 향상시키는 효과적인 방법임을 입증합니다.
Future event prediction carries broad social impact yet remains challenging. SOTA approaches augment LLMs with external agent frameworks whose predictive capability vanishes once the harness is removed. While recent Tool-Integrated Reasoning (TIR) internalizes deep search for multi-hop retrieval of facts, forecasting further demands temporal search and reasoning over historical trends and dynamic shifts. The key obstacle is data: historical queries induce temporal leakage that degrades forecasting into retrieval. Prior works either freeze information gathering with static observations, or rely on rejection sampling or unresolved fresh queries that discard vast amounts of data, degrading synthesis efficiency. We propose a time-truncation harness that enforces a temporal cut-off at every turn, enabling TIR-style sampling from historical events, reducing temporal leakage and reliance of rejection sampling or unsolved queries, increasing the sampling efficiency. We further build a large-scale corpus and a process-based metric and show that our harness naturally induces a broader temporal breadth of search and raises the proportion of high-quality data, further increasing the efficiency and reducing the reliance on complex rubrics. Distillation experiments show that students trained on harness-intervened data achieve the best performance, demonstrating harness-assisted model evolving that turns higher quality temporal search and reasoning data into a parametric advancement of the students.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.