범용 시각 언어 모델(VLM)이 천문 기초 모델의 은하 형태 인식 성능을 향상시키는 방법
A General-Purpose VLM Can Teach an Astronomy Foundation Model to Better Recognize Galaxy Morphology
기존의 천문 기초 모델은 우수한 은하 표현 방식을 제공하지만, 새로운 관측 조건에 적응하거나 특정 관측 데이터에 특화된 형태 인식 작업을 수행하려면 여전히 상당한 수준의 인간 감독이 필요합니다. 본 연구에서는 VLM 기반 질의 응답(VQA) 시스템이 의미 있는 시각-의미 연결 정보를 포함하고 있으며, 이러한 정보가 하위 작업의 형태 분류기에 대한 약한 형태의 지도 학습으로 작용하여 제한된 인간 라벨 예산 환경에서 형태 분류 성능을 향상시킬 수 있음을 보여줍니다. 먼저 두 가지 대표적인 관측 방식을 포괄하는 VQA 벤치마크를 제시하고, 최첨단 VLM 모델들을 사용하여 은하 형태 관련 질문에 대한 성능을 평가했습니다. 결과는 이러한 모델들이 유용한 형태 정보를 그리고 정보적인 불확실성을 파악하지만, 인간 주석자의 역할을 대체하기에는 충분히 신뢰성이 높지 않음을 나타냅니다. 이러한 결과를 바탕으로, 범용 VLM을 Zoobot이라는 천문 기초 모델의 형태 인식 '교사'로 활용했습니다. Zoobot은 대규모 Galaxy Zoo 데이터셋으로 사전 학습된 모델입니다. 두 가지 관측 영역과 다양한 라벨 예산에서, VLM 교사는 Zoobot의 하위 작업 형태 분류 성능을 꾸준히 향상시켰습니다. 이러한 결과는 범용 VLM이 천문 기초 모델에 보완적인 지식을 제공하며, 제한된 인간 감독 환경에서 은하 형태 인식 능력을 향상시키는 데 기여할 수 있음을 보여줍니다. 개발된 파이프라인은 Vera C. Rubin Observatory의 Legacy Survey of Space and Time (LSST) 및 Nancy Grace Roman Space Telescope를 포함한 향후 대규모 관측 프로젝트에 효율적으로 적용될 수 있도록 설계되었습니다. 벤치마크와 코드는 https://github.com/fw-ic/VLM-morphology-teacher 에서 공개적으로 이용할 수 있습니다.
Existing astronomy foundation models provide strong galaxy representations, but adapting them to new survey conditions and survey-specific morphology recognition tasks still requires substantial human supervision. We show that VLM-based VQA systems contain meaningful visual-semantic priors that can serve as weak supervision for downstream morphology classifiers and improve morphology classification under limited human-label budgets. We first introduce a survey-oriented VQA benchmark spanning two representative imaging regimes and evaluate state-of-the-art VLMs on galaxy morphology questions. The results show that these models capture useful morphology signals and informative uncertainty, but are not sufficiently reliable to replace human annotators. Motivated by this finding, we use a general-purpose VLM as a morphology teacher for Zoobot, an astronomy foundation model pretrained on large-scale Galaxy Zoo annotations. Across two survey domains and multiple annotation budgets, the VLM teacher consistently improves Zoobot's downstream morphology classification. These results demonstrate that a general-purpose VLM provides knowledge complementary to an astronomy foundation model and can teach it to better recognize galaxy morphology under limited human supervision. The resulting pipeline is designed for label-efficient adaptation to forthcoming large-scale surveys, including the Vera C. Rubin Observatory's Legacy Survey of Space and Time (LSST) and the Nancy Grace Roman Space Telescope. The benchmark and code are publicly available at https://github.com/fw-ic/VLM-morphology-teacher.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.