2607.26397v1 Jul 29, 2026 cs.CL

추론 이전의 지식: EC-Reason-Bench - LLM 효소 분류를 위한 학습 불필요 진단 벤치마크

Knowledge before Reasoning: EC-Reason-Bench, a Training-Free Diagnostic Benchmark for LLM Enzyme Classification

Huanyao Zhang
Huanyao Zhang
Citations: 33
h-index: 3
Yuanpeng He
Yuanpeng He
Citations: 73
h-index: 5
N. Tashi
N. Tashi
Citations: 937
h-index: 17
Linyu Li
Linyu Li
Citations: 25
h-index: 2
Zhi Jin
Zhi Jin
Citations: 33
h-index: 3
Yichi Zhang
Yichi Zhang
Citations: 7
h-index: 2
Dongming Jin
Dongming Jin
Citations: 120
h-index: 5
Xuan Zhang
Xuan Zhang
Citations: 13
h-index: 2
Gadeng Luosang
Gadeng Luosang
Citations: 65
h-index: 5

효소 기능 예측은 계층적이고 지식 집약적인 단백질 기능 분류의 한 형태입니다. 기존 벤치마크는 다음과 같은 현상을 보여줍니다. 일반적인 LLM은 일반적으로 첫 번째 수준에서는 비교적 정확한 결과를 보이지만, EC 번호를 완벽하게 예측하도록 요청하면 두 번째부터 네 번째 수준까지의 정확도가 거의 0으로 떨어지는 반면, 전문 모델과 도구는 여전히 유용합니다. 우리는 학습이 필요 없는 진단 평가 프로토콜인 EC-Reason-Bench를 제안합니다. 이는 다음 두 가지 질문에 답하기 위해 설계되었습니다. 일반적인 LLM이 왜 EC 번호 예측에서 거의 0에 가까운 성능을 보이는가? 그리고 단일 가중치를 업데이트하지 않고도 얼마나 많은 성능 향상을 얻을 수 있는가? 우리는 효소 분류 능력을 네 가지 독립적인 요소로 나누어 각 요소를 개별적으로 측정합니다. 이러한 요소는 출력 구조, 외부 지식, 추론 구조 및 추론 견고성입니다. 우리는 각 요소를 추론 시간을 이용한 방법으로 테스트하고, 공유된 제로샷 기반을 사용하여 이전에 보고된 거의 0에 가까운 성능을 재현합니다. 여러 강력한 추론 LLM에 대한 실험 결과, 네 가지 주요 결과를 얻었습니다. 첫째, 외부 지식은 결정적인 역할을 하며, 추론보다 먼저 제공되어야 합니다. 제한된 정보 환경에서의 낮은 성능은 개방형 학습 환경으로 전환하면 크게 향상되며, 모델 간의 격차를 좁힙니다. 둘째, 제한된 정보 환경에서 연결 추론(cascading)과 연쇄적 사고(chain-of-thought)가 도움이 되는지 해로운지는 모델이 회피하는 경향에 따라 달라집니다. 셋째, 증거가 제공되면 최상의 LLM 설정의 종합 점수는 가장 가까운 이웃의 EC 번호에 단순히 투표하는 것과 구별할 수 없습니다. 이러한 동률은 평균화의 결과이며, 이는 다중 기능 효소에서 큰 손실을 숨기면서 적대적인 증거 세트에 대한 상당한 이득을 가립니다. 따라서 추론은 지식의 원천이 아니라 충돌하는 이웃 간의 중재자 역할을 하며, 단일 숫자로 표현된 순위표에서는 이러한 현상을 파악할 수 없습니다. 넷째, 정확도는 이용 가능한 유사성 정보의 양에 따라 결정됩니다.

Original Abstract

Enzyme function prediction is a hierarchical, knowledge-intensive form of protein function classification. Existing benchmarks expose an anomaly: general LLMs often get the coarse first level right, yet once asked for a complete EC number their accuracy at levels two through four drops to almost zero, while specialized models and tools stay usable. We propose EC-Reason-Bench, a training-free, diagnostic evaluation protocol built to answer two questions: why general LLMs score close to nothing on EC number prediction, and how much of that loss can be recovered without updating a single weight. We break enzyme classification ability into four orthogonal levers that can each be measured on their own: output structure, external knowledge, reasoning structure, and reasoning robustness. We test each lever with an inference-time method against a shared zero-shot baseline reproducing previously reported near-zero performance. Experiments with several strong reasoning LLMs yield four main findings. First, external knowledge is decisive and must precede reasoning: uniformly low closed-book performance rises sharply with open-book access, narrowing model gaps. Second, in closed-book settings, whether cascading and chain-of-thought help or hurt depends on a model's tendency to abstain. Third, once evidence is available the aggregate score of the best LLM setting is indistinguishable from simply voting the EC numbers of the nearest retrieved neighbors; that tie is an artifact of averaging, and it hides a large gain on adversarial evidence set against an equally large loss on multi-functional enzymes. Reasoning over evidence therefore acts as an arbiter of conflicting neighbors rather than as a source of knowledge, and no single-number leaderboard can see it. Fourth, accuracy obeys a law of homology availability.

0 Citations
0 Influential
8.5 Altmetric
42.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!