2607.27766v1 Jul 30, 2026 cs.CL

그래디언트 기반이 아닌 태스크 조건부 검색을 통한 온디바이스 인컨텍스트 학습

Gradient-free Task-Conditioned Retrieval for On-Device In-Context Learning

Yihua Shao
Yihua Shao
Citations: 230
h-index: 9
Junyi Yang
Junyi Yang
Citations: 86
h-index: 4
Xinyu Luo
Xinyu Luo
Citations: 10
h-index: 2
Hui Liu
Hui Liu
City University of Hong Kong
Citations: 288
h-index: 6
Arindam Basu
Arindam Basu
Citations: 6
h-index: 2
Haoliang Li
Haoliang Li
Citations: 193
h-index: 4

온디바이스 인컨텍스트 학습(ICL)은 다운스트림 모델 추론 전에 유용한 컨텍스트를 선택하기 위해 사전 추론 검색을 활용합니다. 이 검색은 제한된 계산 능력, 메모리 및 데이터 노출 예산 하에서 태스크별 정보를 활용해야 합니다. 본 논문에서는 쌍으로 구성된 후보 입력과 출력을 사용하여 고정된 인코더를 태스크 조건부 검색기로 변환하는 그래디언트 기반이 아닌 프레임워크인 Conditional Retrieval Alignment (CoRA)를 제안합니다. CoRA는 상호 보완적인 인코더 레이어를 선택하고, 후보 메모리에서 출력으로 파생된 컨디셔닝 공간을 구성하며, 닫힌 형태의 ridge regression을 통해 후보 입력 표현을 이 공간에 정렬합니다. 저차원 분해(low-rank factorization)를 통해 압축된 검색 기반을 생성하며, 여기서 후보 출력은 오프라인 인덱스 구축 시에만 사용됩니다. 쿼리 시간 검색에는 쿼리 입력과 사전 계산된 인덱스만 필요합니다. CoRA의 순위 제한 기반이 출력 조건부 맞춤 표현(fitted representation)의 최적의 저차원 압축이라는 것을 보여주고, 전체 맞춤 행렬을 구체화하지 않고 정확한 두 단계 스트리밍 구축 방법을 도출했습니다. 또한, 시각적 표현을 컨디셔닝 및 검색 공간에 통합하여 프레임워크를 다중 모달 예제 검색으로 확장합니다. Llama-3.2-1B, MobileLLM-Pro, OpenFlamingo-3B 및 Qwen3.5-2B 모델과 10개의 텍스트 데이터셋 및 4개의 다중 모달 벤치마크를 사용한 실험 결과는 CoRA가 검색기 미세 조정, 역전파 또는 대상 모델 호출 없이 효과적인 태스크 조건부 검색을 지원한다는 것을 보여줍니다. 또한, Raspberry Pi~5 환경에서의 엔드투엔드 배포를 통해 성능을 검증했습니다.

Original Abstract

On-device in-context learning (ICL) relies on pre-inference retrieval to select demonstrations for useful context before downstream model inference. This retrieval must exploit task-specific information while operating over local memories under limited computation, memory, and data-exposure budgets. We propose Conditional Retrieval Alignment (CoRA), a gradient-free framework that converts a frozen encoder into a task-conditioned retriever using paired candidate inputs and outputs. CoRA selects complementary encoder layers, constructs an output-derived conditioning space from candidate memory, and aligns candidate input representations to this space through closed-form ridge regression. Low-rank factorization then produces a compact retrieval basis where candidate outputs are used only during offline index construction, whereas query-time retrieval requires only the query input and precomputed index. We show that CoRA's rank-constrained basis is the optimal low-rank compression of the output-conditioned fitted representation, and derive an exact two-pass streaming construction that avoids materializing the full fitted matrix. We further extend the framework to multimodal exemplar retrieval by incorporating visual representations into the conditioning and retrieval spaces. Experiments across ten textual datasets and four multimodal benchmarks with Llama-3.2-1B, MobileLLM-Pro, OpenFlamingo-3B, and Qwen3.5-2B, as well as end-to-end Raspberry Pi~5 deployment demonstrate that CoRA supports effective task-conditioned retrieval without retriever fine-tuning, backpropagation, or target-model calls.

0 Citations
0 Influential
4.5 Altmetric
22.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!