2607.06402v1 Jul 07, 2026 cs.CV

이미지가 말할 수 없는 것: 언어 기반의 후각 표현 학습

What Images Cannot Say: Language-Guided Olfactory Representation Learning

Xi Wang
Xi Wang
Citations: 149
h-index: 6
Vicky Kalogeiton
Vicky Kalogeiton
Citations: 164
h-index: 5
Eleftherios Tsonis
Eleftherios Tsonis
Citations: 6
h-index: 2

이미지는 특정 장면이 어떻게 보이는지를 알려주지만, 그 장면에 실제로 있는 듯한 느낌을 전달하는 경우는 드뭅니다. 최근에는 시각적 장면과 전자 코 측정 데이터를 결합한 데이터셋들이 등장했지만, 여전히 냄새 신호를 이미지와 연결하는 것은 어렵습니다. 왜냐하면 많은 후각 정보는 직접적으로 픽셀에 드러나지 않는 맥락적인 환경 요인에서 비롯되기 때문입니다. 본 논문에서는 SCENT라는 다중 모드 프레임워크를 제안합니다. 이 프레임워크는 언어 지침을 활용하여 시각과 후각 사이의 의미적 연결 고리를 제공합니다. 우리는 Vision-Language Model (VLM)을 사용하여 객체, 환경 맥락, 그리고 시각적인 장면에서 유추할 수 있는 주변 냄새 정보를 포괄하는 장면 설명을 생성합니다. 이러한 설명은 후각 표현 학습에 대한 의미론적 지침을 제공합니다. 우리는 전자 코 신호를 시각 및 텍스트 표현과 일치하는 공유 임베딩 공간으로 매핑하는 후각 인코더를 학습시키고, 객체별 냄새와 맥락적인 환경 요인의 영향을 분리하는 언어 기반 잠재 변수 분해 방식을 도입합니다. New York Smells 데이터셋에 대한 실험 결과, SCENT는 시각 정보만 사용하는 기본 모델보다 크로스모달 검색 성능을 크게 향상시켜, 냄새-이미지 및 냄새-텍스트 검색 작업에서 최고 수준의 성능을 달성했습니다. 또한, 본 프레임워크는 해석 가능한 후각 표현을 생성하여 복잡한 냄새 혼합물을 분리하는 데 도움을 줍니다. 우리의 연구 결과는 다중 모드 학습에서 후각 인식을 이해하는 데 있어 맥락적인 의미 정보의 중요성을 보여주며, 이 분야의 향후 연구를 위한 기반을 마련합니다.

Original Abstract

Images tell us what a scene looks like, but rarely what it would feel like to be there. While recent datasets pair visual scenes with electronic-nose measurements, aligning smell signals with images remains challenging because many olfactory cues arise from contextual environmental factors that are not directly visible in pixels. We introduce SCENT, a multimodal framework that uses language guidance as a semantic bridge between vision and olfaction. Our approach leverages Vision-Language Models (VLMs) to generate scene descriptors capturing objects, environmental context, and plausible ambient smell cues suggested by the visual scene. These descriptors provide semantic guidance for learning olfactory representations. We train a smell encoder that maps electronic-nose signals into a shared embedding space aligned with both visual and textual representations, and introduce a languageguided latent decomposition that separates object-specific odors from contextual environmental contributions. Experiments on the New York Smells dataset demonstrate that SCENT significantly improves crossmodal retrieval compared to vision-only baselines, achieving state-of-theart performance on smell-to-image and smell-to-text retrieval tasks. In addition, our framework produces interpretable olfactory representations that enable the disentanglement of complex smell mixtures. Our results reveal the importance of contextual semantic information for grounding olfactory perception in multimodal learning and pave the way for future research in this area.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!