2607.08400v1 Jul 09, 2026 cs.CR

TRACE: 보완적인 임베딩을 이용한 이중 채널의 강력한 출처 표시 워터마크 - LLM 에이전트 추적 로그를 위한 방법

TRACE: A Two-Channel Robust Attribution Watermark via Complementary Embeddings for LLM-Agent Trajectories

Zhenchang Xing
Zhenchang Xing
Citations: 434
h-index: 12
Liming Zhu
Liming Zhu
Citations: 1,237
h-index: 19
Yulei Sui
Yulei Sui
UNSW
Citations: 4,612
h-index: 37
Xiaoyu Li
Xiaoyu Li
Citations: 22
h-index: 2
Xiaoyang Feng
Xiaoyang Feng
Citations: 71
h-index: 3
Jiaojiao Jiang
Jiaojiao Jiang
Citations: 39
h-index: 3
Yang Song
Yang Song
Citations: 6
h-index: 2
Zhengguang Gao
Zhengguang Gao
Citations: 2
h-index: 1

LLM 에이전트는 재판매업체를 통해 사용자에게 제공되며, 재판매업체는 개발자의 에이전트를 재브랜딩하거나 더 저렴한 모델로 대체할 수 있습니다. 출처가 분쟁될 경우, 출처 확인은 에이전트의 추적 로그(도구 호출 기록, 관찰 내용 및 실행된 작업 기록이며, 모델의 추론 과정은 아님)에 의존하며, 재판매업체는 사용량 측정을 위해 이 로그를 저장하고 처리합니다. 따라서 워터마크는 감지 대상인 증거 자체에 대한 완전한 읽기/쓰기 권한을 가진 공격자로부터도 생존해야 합니다. 기존 에이전트 워터마크는 그렇지 못하며, 그 이유는 출처 정보가 해당 로그에서 직접 읽혀지기 때문입니다. 본 논문에서는 현재까지 알려진 가장 강력한 에이전트 워터마크인 TRACE를 소개합니다. TRACE는 작업 선택 과정에서 왜곡이 없고, 삭제 시에도 자체적으로 동기화되며, 재작성 시에도 조건부로 불변성을 유지합니다. 삭제는 위치 기반 키를 비동기화시키고, 재작성은 내용을 변경하므로, 삭제에 강한 키는 내용에서 파생되어야 하고, 재작성에 강한 키는 위치에서 파생되어야 하며, 단일 키로는 두 가지 모두를 보장할 수 없습니다. 그러나 추적 로그에는 두 개의 워터마크를 포함할 공간이 있습니다. TRACE는 선택 채널과 집계 채널을 사용합니다. 선택 채널은 로컬 콘텐츠를 기반으로 작동하며, 왜곡 없는 샘플러를 사용하여 어떤 작업이 선택되는지 결정합니다. 이를 통해 에이전트의 분포가 변경되지 않으며, 삭제 후에도 감지가 재동기화됩니다. 또한, 집계 채널은 로그의 기본 구조만을 기반으로 각 의사 결정 그룹에 포함된 레코드 수를 설정하며, 이는 어떠한 강도의 재작성도 수정할 수 없습니다. 우리는 이 행동 기반 워터마크의 신호가 의사 결정 엔트로피를 사용하여 생성되며, 각 의사 결정은 최소 절반 이상의 엔트로피를 사용하고, 결정적인 의사 결정은 아무런 비용도 들지 않으며, 두 채널 모두 삭제하면 재판매업체는 판매하는 추적 로그를 손상시켜야 한다는 것을 증명합니다. ToolBench 및 ALFWorld에서 TRACE는 워터마크가 없는 에이전트의 성공률과 일치하며, 선택 채널은 긴 시간 범위의 추적 로그에서 z = 100에 가까운 높은 감지 점수를 달성하고, 70%의 단계 삭제에도 여전히 감지가 가능하며, LLM을 사용하여 수행된 어떤 강도의 재작성에도 집계 채널은 정확히 동일하게 유지됩니다.

Original Abstract

LLM agents reach users through resellers, who may rebrand a developer's agent or substitute a cheaper model. When provenance is disputed, attribution rests on the trajectory log (the record of tool calls, observations, and executed actions, not the model's reasoning), which the reseller stores and processes to meter usage. A watermark must therefore survive an adversary with full read/write access to the very evidence it is detected from; existing agent watermarks do not, as their attribution is read straight off that log. We present TRACE, to our knowledge the first agent watermark that is distortion-free in its action choices, self-synchronizing under deletion, and unconditionally invariant under rewriting. Deletion desynchronizes a position-derived key and rewriting alters content, so a deletion-robust key must come from content and a rewrite-robust key from position, and no single key serves both. A trajectory, however, has room for two watermarks. TRACE superposes a selection channel that sets which action is chosen, keyed on local content with a distortion-free sampler, so the agent's distribution is provably unchanged and detection resynchronizes after deletions, and a tally channel that sets how many records each decision group holds, keyed on the log's skeleton alone, which no rewriting can touch. We prove this behavioral watermark's signal is bought with decision entropy, each decision paying at least half its entropy and deterministic decisions nothing, and that erasing both channels forces the reseller to corrupt the trajectories it resells. On ToolBench and ALFWorld, TRACE matches the unwatermarked agent's success rate while its selection channel reaches detection scores near z = 100 on long-horizon trajectories, stays detectable under 70% step deletion, and keeps a tally channel exactly unchanged under LLM rewriting of any strength.

1 Citations
0 Influential
18.5 Altmetric
93.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!