2608.04567v1 Aug 05, 2026 cs.CL

STRIVE: 등급별 신뢰성 생성 및 평가를 통한 추론 능력의 한계 탐구

STRIVE: Probing Reasoning Limits in Graded Plausibility Generation and Evaluation

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Xiang Li
Xiang Li
Citations: 21
h-index: 2
A. Chrabaszcz
A. Chrabaszcz
Citations: 438
h-index: 11
Haley C. Dresang
Haley C. Dresang
Citations: 147
h-index: 6
Bhiman Kumar Baghel
Bhiman Kumar Baghel
Citations: 28
h-index: 3

사건 지식은 누가 누구에게 무엇을 하는지에 대한 정보를 포함합니다. 심리 언어학자들은 사건의 신뢰성 판단을 사용하여 이러한 지식이 인간의 언어 처리 과정을 어떻게 지원하는지 연구합니다. 이러한 연구에서는 신뢰성에 미치는 영향을 명확히 하기 위해, 하나의 사건 요소만 신뢰성 수준에 따라 변하고 다른 모든 요소는 고정된 통제된 사건 집합이 필요합니다. 이러한 집합을 수동으로 구성하는 것은 많은 노력이 소요됩니다. 따라서 우리는 STRIVE라는 LLM 기반 프레임워크를 소개합니다. 이 프레임워크는 의도적인 분류 난이도(쉬움 vs. 어려움)와 함께 신뢰성 범주(신뢰 가능 vs. 비신뢰 가능)에 따라 통제된 사건 집합을 동시에 생성하고 평가하는 데 사용됩니다. STRIVE는 주어진 동사에 대해 공유된 사건 프레임을 구성한 다음, 다른 모든 요소를 고정하면서 하나의 요소만 변경하여 각 조건에 해당하는 하나의 사건을 생성합니다. 6개의 모델과 60개의 동사를 사용하여 수행한 실험에서, 기본 생성 프롬프트를 사용했을 때 GPT-5.1은 16.7%의 경우에만 고품질의 집합을 생성했습니다. 전역적인 추론 공간(scratchpad)을 추가하고 평가자 중심의 개선 과정을 적용하면 이 비율이 75.0%로 증가했습니다. 더 많은 추론 노력을 기울이는 것이 평가자와 인간 간의 일치도를 향상시키는 데 도움이 되었습니다. 그러나 신뢰성 경계 근처의 사건은 여전히 가장 어렵습니다. 이러한 사건은 가장 큰 인간의 의견 불일치를 야기하며, 최고의 평가자는 비신뢰 가능-어려운 조건에서 57%의 정확도에 도달하는 데 그치므로, 여전히 인간의 입력이 필요합니다. 전체적으로 STRIVE는 심리 언어학 연구를 위한 초기 사건 집합 생성 및 평가 과정을 자동화하여 수동 노력을 줄이는 확장 가능한 접근 방식을 제공합니다.

Original Abstract

Event knowledge concerns who does what to whom. Psycholinguists use event-plausibility judgments to examine how this knowledge supports human language processing. To isolate plausibility effects, these studies require controlled event sets in which one event slot varies across plausibility levels while all other event features remain fixed. Constructing such sets manually is labor-intensive. We therefore introduce STRIVE, an LLM-based framework for jointly generating and evaluating controlled event sets crossing plausibility class (plausible vs. implausible) with intended classification difficulty (easy vs. hard). Given a verb, STRIVE constructs a shared event frame, then produces one event per condition by varying one slot while holding all others fixed. In experiments with six models across 60 verbs, GPT-5.1 produced high-quality sets only 16.7% of the time using the baseline generation prompt. Adding a global reasoning scratchpad and evaluator-guided refinement raised this rate to 75.0%. Greater reasoning effort also improved evaluator--human agreement. Nevertheless, events near the plausibility boundary remain most difficult. They elicit the greatest human disagreement, and the best evaluator reaches only 57% accuracy on the implausible-hard condition, indicating a need for human input. Overall, STRIVE offers a scalable approach to reducing manual effort by automating initial event-set generation and evaluation for psycholinguistic studies.

0 Citations
0 Influential
5.5 Altmetric
27.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!