2606.16253v1 Jun 15, 2026 cs.CV

시각-언어-행동 모델을 위한 학습 기반 이미지 압축

Learned Image Compression for Vision-Language-Action Models

Hyeonjun Kim
Hyeonjun Kim
Citations: 51
h-index: 3
Jegwang Ryu
Jegwang Ryu
Citations: 0
h-index: 0
S. Ha
S. Ha
Citations: 59
h-index: 5
Junhyeok Lee
Junhyeok Lee
Citations: 25
h-index: 2
Jun-Hyuk Kim
Jun-Hyuk Kim
Citations: 30
h-index: 2
Hyemin Ahn
Hyemin Ahn
Citations: 584
h-index: 12
Jaeho Lee
Jaeho Lee
POSTECH
Citations: 2,486
h-index: 19

시각-언어-행동(VLA) 모델은 점점 더 고주파 다중 카메라 관찰에 의존하고 있으며, 이는 대역폭이 제한되거나 분산된 환경에서 실시간 로봇 제어를 위한 주요 병목 현상입니다. 기존의 이미지 및 비디오 코덱은 일반적인 시각적 충실도를 유지하도록 설계되었지만, 다운스트림 VLA 정책의 제어 성능을 최적화하도록 설계되지 않았습니다. 본 연구에서는 VLA 기반 로봇에 특화된 학습 기반 이미지 압축 프레임워크인 SPARC(SPatially Adaptive Rate Control)를 소개합니다. 우리의 주요 관찰은 시각 정보의 중요성이 카메라 뷰와 이미지 내 공간 영역 모두에서 크게 다르다는 것입니다. 이러한 관찰을 바탕으로, SPARC는 가벼운 시간 마스크 선택기를 사용하여 작업 관련성에 따라 잠재 표현에 비트율을 적응적으로 할당하고 동시에 시간적 맥락을 활용합니다. 또한, 우리는 엔트로피 기반 목표가 드물지만 작업에 중요한 시각 패턴을 과도하게 억제하는 경향을 줄여 학습의 안정성을 높이는 기울어진 속도 손실 함수를 도입했습니다. RoboCasa365, VLABench 및 LIBERO 등 다양한 로봇 벤치마크에서 수행한 실험 결과, SPARC는 동일한 비트율 예산 하에서 기존 이미지/비디오 코덱 및 최근 학습 기반 압축 방법보다 일관되게 더 강력한 제어 성능을 달성합니다. 또한, 원격 제어 환경에서 실제 배포 이점을 보여주며, 우리의 방법은 비트율과 성공률 간의 균형을 크게 향상시킵니다.

Original Abstract

Vision-language-action (VLA) models increasingly rely on high-frequency multi-camera observations, making visual communication a major bottleneck for real-time robotic control in bandwidth-constrained or distributed deployment settings. Existing image and video codecs, however, are designed to preserve generic visual fidelity rather than the control performance of downstream VLA policies. In this work, we introduce SPARC (SPatially Adaptive Rate Control), a learned image compression framework tailored for VLA-driven robots. Our key observation is that the importance of visual information varies substantially across both camera views and spatial regions within an image. Based on this observation, SPARC employs a lightweight temporal mask selector that adaptively allocates bitrate over latent representations according to task relevance while leveraging temporal context. We further introduce a tilted rate loss that stabilizes training by reducing the tendency of entropy-based objectives to over-suppress rare yet task-critical visual patterns. Experiments on diverse robotic benchmarks, including RoboCasa365, VLABench, and LIBERO, show that SPARC consistently achieves stronger control performance than conventional image/video codecs and recent learned compression methods under the same bitrate budget. We additionally demonstrate real-world deployment benefits in remote-control settings, where our method substantially improves the bitrate-success tradeoff.

0 Citations
0 Influential
9.5 Altmetric
47.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!