구조 기반 삽입: 대규모 언어 모델을 위한 언어학적 기반 편집 기반 코드 혼합 지문
Construction-Driven Injection: Linguistically-Grounded Edit-Based Code-Mixing Fingerprints for Large Language Models
대규모 언어 모델(LLM)은 상당한 가치를 지니지만, 무단 배포 및 상업적 남용에 취약합니다. 모델의 동작 내에 삽입된 지문(트리거-타겟 쌍)은 실질적인 소유권 증명 역할을 하며, 블랙박스 환경에서도 검증 가능하지만, 기존 방법들은 지문 생성과 삽입이라는 두 단계를 분리합니다. 기존 지문 프레임워크는 두 가지 한계를 가지고 있습니다. 자연어 기반 지문은 의도치 않은 활성화될 위험이 있으며, 의미가 불분명한 지문은 퍼플렉시티 기반 탐지로 쉽게 걸러집니다. 또한, 생성과 삽입을 분리하면 삽입 과정에서 트리거의 언어적 구조를 알 수 없어, 목표에 맞는 최적화를 수행할 기회를 놓치게 됩니다. 우리는 지문 생성이 삽입을 이끌어야 한다고 주장하며, 두 단계를 동시에 최적화하는 통합 지문 프레임워크를 제시합니다. 먼저, LCF(Language-Constructed Fingerprint)는 의미 밀도 대체 규칙과 문법 기반 혼합을 사용하여 저자원 언어를 결합하여 코드 혼합 지문을 생성합니다. 이 방법은 의미가 불분명한 기준보다 훨씬 낮은 퍼플렉시티를 가지는 트리거를 생성하며, 자연어 기반 트리거의 의도치 않은 활성화 문제를 해결합니다. 둘째, LCFEdit는 각 지문을 고자원 다국어 표현에서 파생된 null-space 투영을 사용하여 삽입합니다. 이 과정은 지식을 유지하고, 횡단 언어 정렬 단계를 통해 가중치 업데이트를 지문의 언어 표현 공간으로 유도하여 지문 언어의 특징을 반영하도록 합니다. 이러한 구조 인식 기반 삽입은 업데이트가 언어학적으로 정보를 담고 있어 더욱 안정적입니다. 투명성, 탐지 가능성 및 안전성에 대한 광범위한 실험 결과는 유틸리티에 미치는 영향이 거의 없는 상태에서 지속적인 소유권 검증을 보여줍니다.
Large language models (LLMs) are costly intellectual assets that remain exposed to unauthorized redistribution and commercial misuse. Injected fingerprints, i.e., trigger--target pairs embedded in model behavior, offer a practical, black-box-verifiable ownership signal, but existing methods decouple the two stages of the fingerprint life cycle: how a fingerprint is constructed and how it is injected. Existing fingerprinting frameworks suffer from two limitations. Natural-language fingerprints are prone to accidental activation, and garbled fingerprints are easily filtered by perplexity-based detection. Furthermore, decoupling construction from injection leaves the latter unaware of the trigger's linguistic structure, missing the opportunity for targeted optimization. We argue that fingerprint construction should drive injection, and present a unified fingerprinting framework that jointly optimizes both stages. First, LCF constructs code-mixing fingerprints by combining low-resource languages under a semantic-density substitution rule and grammar-biased mixing, yielding triggers whose perplexity sits far below garbled baselines while avoiding the accidental-activation failures of natural-language triggers. Second, LCFEdit injects each fingerprint with a null-space projection derived from high-resource multilingual representations that preserves knowledge, augmented by a cross-lingual alignment step that steers the weight update toward the fingerprint language's representation subspace. This construction-aware injection ensures that the update is linguistically informed and therefore more stable. Extensive evaluations on imperceptibility, detectability, and harmlessness demonstrate persistent ownership verification with negligible impact on utility.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.