SibylSense: 메모리 튜닝 및 적대적 탐색을 통한 적응적 루브릭 학습
SibylSense: Adaptive Rubric Learning via Memory Tuning and Adversarial Probing
개방형 생성에 대한 정렬되고 강력한 보상을 설계하는 것은 강화 학습의 사후 훈련에서 여전히 중요한 과제입니다. 루브릭은 체계적이고 해석 가능한 감독 신호를 제공하지만, 루브릭 구축의 확장은 어렵습니다. 전문가 루브릭은 비용이 많이 들고, 프롬프트 기반 루브릭은 종종 피상적이거나 일관성이 없으며, 고정된 풀의 판별적 루브릭은 포화 및 드리프트를 일으켜 보상 공격을 가능하게 합니다. 본 논문에서는 검증된 루브릭 항목의 조정 가능한 메모리 뱅크를 통해 정적인 루브릭 생성기를 적응시키는 추론 시간 학습 접근 방식인 SibylSense를 제안합니다. 메모리는 검증기 기반 항목 보상을 통해 업데이트되며, 이는 소수의 예시로부터 얻은 참조-후보 답변의 판별적 격차를 측정합니다. SibylSense는 메모리 튜닝과 루브릭 만족 후보 답변을 생성하는 루브릭-적대적 정책 업데이트를 반복하여 판별적 격차를 줄이고 루브릭 생성기가 새로운 품질 차원을 포착하도록 유도합니다. 두 가지 개방형 작업에 대한 실험 결과, SibylSense는 더욱 판별적인 루브릭을 생성하고 정적 및 비적응적 기준보다 하위 작업의 강화 학습 성능을 향상시킵니다.
Designing aligned and robust rewards for open-ended generation remains a key barrier to RL post-training. Rubrics provide structured, interpretable supervision, but scaling rubric construction is difficult: expert rubrics are costly, prompted rubrics are often superficial or inconsistent, and fixed-pool discriminative rubrics can saturate and drift, enabling reward hacking. We present SibylSense, an inference-time learning approach that adapts a frozen rubric generator through a tunable memory bank of validated rubric items. Memory is updated via verifier-based item rewards measured by reference-candidate answer discriminative gaps from a handful of examples. SibylSense alternates memory tuning with a rubric-adversarial policy update that produces rubric-satisfying candidate answers, shrinking discriminative gaps and driving the rubric generator to capture new quality dimensions. Experiments on two open-ended tasks show that SibylSense yields more discriminative rubrics and improves downstream RL performance over static and non-adaptive baselines.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.