LSR-Synth에서의 라이브러리 접근성: 반암기화 설계가 상징적 발견 측정 방식에 미치는 영향
Library Reachability in LSR-Synth: How Anti-Memorization Design Changes the Measurement of Symbolic Discovery
과학 방정식 발견을 위한 기존의 벤치마크는 대부분 공개적으로 이용 가능한 잘 알려진 방정식으로 구성되어 있어, 모델이 데이터로부터 법칙을 발견하는지 아니면 단순히 학습된 자료에서 답을 회상하는지에 대한 판단을 어렵게 만듭니다. LSR-Synth는 새로운 합성 용어를 기존의 과학적 메커니즘에 도입하고, 생성된 작업의 참신성, 해결 가능성 및 과학적 타당성을 필터링하여 이러한 문제를 완화합니다. 본 논문에서는 더 좁은 측정 질문을 다룹니다: 이러한 작업들이 언어 모델이 제공하는 과학적 사전 지식과 작업 의미를 접근하지 않는 기존의 연산자 탐색 방법을 더욱 명확하게 구분할 수 있는가? 공개적으로 문서화된 출처를 가진 고정 어휘 집합을 사용하여 의미 기반(semantics-free) 기준선을 구축하고, 의미 블라인딩, 라이브러리 약화 및 일치하는 연산자 패밀리의 제거를 통해 후보 커버리지의 역할을 평가합니다. 현재 작업 스냅샷, 탐색 예산 및 점수 부여 프로토콜 하에서, 고정 어휘는 대부분의 작업을 이미 포괄하며, 언어 모델이 생성한 후보들은 해결 가능한 인스턴스의 범위를 거의 확장하지 않습니다. 어휘 커버리지가 선택적으로 방해될 때에만 이러한 기여도가 상당해집니다. 엄격한 외부 데이터 평가(out-of-distribution evaluation)는 모든 방법의 절대 성공률을 낮추지만, 이러한 관계를 변경하지는 않습니다. 이러한 결과는 LSR-Synth가 완전한 공식 암기를 방지하는 제어 메커니즘을 무효화하지 않으며, 언어 모델 기반 사전 지식이 일반적으로 도움이 되지 않는다는 것을 의미하지도 않습니다. 오히려, 이는 다음과 같은 제한적인 결론을 뒷받침합니다: 대부분의 현재 작업은 여전히 이전에 보지 못한 표현식을 적합하고 재조합하는 능력을 평가하기에 적합하지만, 고정된 탐색 공간 외의 사전 지식의 기여를 식별하기에는 충분하지 않습니다.
Existing benchmarks for scientific equation discovery are largely composed of well-known equations available in the public domain, making it difficult to determine whether a model is discovering laws from data or merely recalling answers from its training corpus. LSR-Synth mitigates this problem by introducing novel synthetic terms into established scientific mechanisms and filtering the resulting tasks for novelty, solvability, and scientific plausibility. This paper examines a narrower measurement question: can these tasks further distinguish scientific priors supplied by language models from conventional operator search that does not access task semantics? We construct a semantics-free baseline using a fixed vocabulary with publicly documented provenance, and assess the role of candidate coverage through semantic blinding, library weakening, and matched operator-family knockouts. Under the current task snapshot, search budget, and scoring protocol, the fixed vocabulary already covers most tasks, while language-model-generated candidates rarely expand the set of solvable instances. Their marginal contribution becomes substantial only when vocabulary coverage is selectively disrupted. Strict out-of-distribution evaluation lowers the absolute success rates of all methods but does not alter this relationship. These findings neither invalidate LSR-Synth's controls against memorization of complete formulas nor imply that language-model priors are generally unhelpful. Rather, they support a more limited conclusion: most current tasks remain suitable for evaluating the fitting and recombination of previously unseen expressions, but are insufficient on their own to identify contributions from priors beyond a fixed search space.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.