LLM 활성화 데이터를 활용한 개념 내용 측정: 개념 벡터 및 선형 탐색을 통한 ESG 사례 연구
Measuring Concept Content in Text from LLM Activations: ESG Evidence from Concept Vectors and Linear Probes
기존의 텍스트 내 개념 관련성 측정 방법은 단어 공유 비율, 주제 분포, 임베딩 유사도 등 표면적인 특성을 기반으로 합니다. 이러한 방법들은 텍스트에 사용된 단어를 평가하지만, 독자가 해당 텍스트에 대해 형성하는 판단과는 거리가 있습니다. 최근 연구에서는 대규모 언어 모델(LLM)이 내부적으로 알고 있는 지식과 실제 응답에서 드러내는 지식 사이에 간극이 존재한다는 사실이 밝혀졌습니다. 본 논문은 동결된, 즉 별도의 추가 학습을 거치지 않은 LLM의 활성화 데이터를 분석하여 개념 내용을 측정할 때, 특정 작업에 맞게 미세 조정된 모델의 성능을 대체할 수 있는지, 그리고 어떤 추출 방법이 가장 효과적인지를 탐구합니다. 우리는 Recursive Feature Machine (RFM) 알고리즘과 선형 탐색 방법을 사용하여 이러한 지표를 추출하고, 이를 임베딩 기반 측정값, 표면 특징 기반 측정값 및 동일 모델 자체의 답변 결과와 비교합니다. 본 연구는 광범위하게 연구되고 관련 데이터셋이 구축된 금융 텍스트 분야에서 인간이 직접 라벨링한 ESG 데이터를 활용하여 제안하는 방법을 검증합니다. 가장 우수한 성능을 보이는 선형 탐색 방법은 별도의 작업별 미세 조정 없이도, 특정 도메인 분류 모델의 정확도와 0.6%p 이내의 오차를 보이며, 12개의 비교 항목 중 11개에서 동일 모델 자체의 답변보다 더 높은 점수를 얻었습니다. 이는 활성화 데이터가 응답에 나타나지 않는 개념 내용을 담고 있음을 시사합니다. 간단한 선형 탐색 방법은 RFM 기반 개념 벡터보다 일관되게 우수한 성능을 보이며, RFM 방식은 분류만으로는 제공할 수 없는 연속적인 점수를 제공하여 텍스트 내 특정 개념의 존재 정도를 나타내지만, 이 점수의 유효성은 추가적인 등급 라벨링을 통해 검증될 예정입니다.
Existing measures of how much a text is about a concept read the surface of the text: dictionary word shares, topic proportions, embedding similarities. They score the words a text uses, not the judgment a reader forms about it. Recent work has shown that a gap exists in what Large Language Models (LLMs) know internally versus what they express in their response. This paper asks whether that internal knowledge, read by monitoring the activations of frozen, out-of-the-box LLMs, can stand in for task-specific fine-tuning when measuring concept content, and which extraction method reads it best. We extract such measures via the Recursive Feature Machine (RFM) algorithm and via linear probing, and compare these against an embedding baseline, surface baselines, and the same model's own answer to the question. We demonstrate the approach on financial text, a domain studied extensively and served by established annotated resources, using a human-annotated Environmental, Social and Governance (ESG) dataset. The best linear probe comes within 0.6 percentage points of a fine-tuned domain classifier's accuracy without any task-specific fine-tuning, and outscores the same model's own answer to the question in eleven of twelve comparisons, so the activations carry concept content the response does not report. The simple probe consistently beats the RFM concept vectors, which in turn provide what classification alone does not: a continuous score intended to reflect how strongly a concept is present in a text, whose validation awaits graded labels.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.