2602.20751v1 Feb 24, 2026 cs.CL

SibylSense: 메모리 튜닝 및 적대적 탐색을 통한 적응적 루브릭 학습

SibylSense: Adaptive Rubric Learning via Memory Tuning and Adversarial Probing

Guilherme Potje
Guilherme Potje
Citations: 4
h-index: 1
S. Shandilya
S. Shandilya
Citations: 325
h-index: 6
Tiancheng Yuan
Tiancheng Yuan
Citations: 14
h-index: 2
Leonardo Nunes
Leonardo Nunes
Citations: 291
h-index: 5
Saeid Asgari
Saeid Asgari
Citations: 10
h-index: 2
Adam Atkinson
Adam Atkinson
Citations: 482
h-index: 4
Emre Kıcıman
Emre Kıcıman
Citations: 598
h-index: 8
Songwu Lu
Songwu Lu
Citations: 66
h-index: 4
Ranveer Chandra
Ranveer Chandra
Citations: 309
h-index: 6
Tusher Chakraborty
Tusher Chakraborty
Citations: 271
h-index: 9
Yifei Xu
Yifei Xu
Citations: 70
h-index: 4
Raksha Agarwal
Raksha Agarwal
Citations: 66
h-index: 6

개방형 생성에 대한 정렬되고 강력한 보상을 설계하는 것은 강화 학습의 사후 훈련에서 여전히 중요한 과제입니다. 루브릭은 체계적이고 해석 가능한 감독 신호를 제공하지만, 루브릭 구축의 확장은 어렵습니다. 전문가 루브릭은 비용이 많이 들고, 프롬프트 기반 루브릭은 종종 피상적이거나 일관성이 없으며, 고정된 풀의 판별적 루브릭은 포화 및 드리프트를 일으켜 보상 공격을 가능하게 합니다. 본 논문에서는 검증된 루브릭 항목의 조정 가능한 메모리 뱅크를 통해 정적인 루브릭 생성기를 적응시키는 추론 시간 학습 접근 방식인 SibylSense를 제안합니다. 메모리는 검증기 기반 항목 보상을 통해 업데이트되며, 이는 소수의 예시로부터 얻은 참조-후보 답변의 판별적 격차를 측정합니다. SibylSense는 메모리 튜닝과 루브릭 만족 후보 답변을 생성하는 루브릭-적대적 정책 업데이트를 반복하여 판별적 격차를 줄이고 루브릭 생성기가 새로운 품질 차원을 포착하도록 유도합니다. 두 가지 개방형 작업에 대한 실험 결과, SibylSense는 더욱 판별적인 루브릭을 생성하고 정적 및 비적응적 기준보다 하위 작업의 강화 학습 성능을 향상시킵니다.

Original Abstract

Designing aligned and robust rewards for open-ended generation remains a key barrier to RL post-training. Rubrics provide structured, interpretable supervision, but scaling rubric construction is difficult: expert rubrics are costly, prompted rubrics are often superficial or inconsistent, and fixed-pool discriminative rubrics can saturate and drift, enabling reward hacking. We present SibylSense, an inference-time learning approach that adapts a frozen rubric generator through a tunable memory bank of validated rubric items. Memory is updated via verifier-based item rewards measured by reference-candidate answer discriminative gaps from a handful of examples. SibylSense alternates memory tuning with a rubric-adversarial policy update that produces rubric-satisfying candidate answers, shrinking discriminative gaps and driving the rubric generator to capture new quality dimensions. Experiments on two open-ended tasks show that SibylSense yields more discriminative rubrics and improves downstream RL performance over static and non-adaptive baselines.

4 Citations
1 Influential
4.5 Altmetric
28.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!