2604.01742v2 Apr 02, 2026 cs.CV

강화 학습 기반 포인트 선택을 통한 밀집 포인트-마스크 최적화를 이용한 군중 인스턴스 분할

Dense Point-to-Mask Optimization with Reinforced Point Selection for Crowd Instance Segmentation

Antoni B. Chan
Antoni B. Chan
Citations: 17
h-index: 1
Hongru Chen
Hongru Chen
Citations: 1
h-index: 1
Jiyang Huang
Jiyang Huang
Citations: 0
h-index: 0
Jia Wan
Jia Wan
Citations: 1,199
h-index: 14

군중 인스턴스 분할은 감시 및 교통 시스템 등 다양한 분야에서 중요한 역할을 하는 과제입니다. 현재 대부분의 군중 데이터셋에서는 점(point) 레이블이 흔하게 사용되지만, 영역(region) 레이블 (예: bounding box)은 드물고 정확하지 않은 경우가 많습니다. 분할을 통해 얻어진 마스크는 영역 레이블의 정확도를 향상시키고, 개별 위치 좌표와 군중 밀도 맵 간의 관계를 해결하는 데 도움을 줍니다. 그러나 현재 널리 사용되는 대규모 모델인 SAM과 같은 모델을 직접 적용하면 밀집된 군중 환경에서 최적의 결과를 얻기 어렵습니다. 이에, 본 연구에서는 먼저 SAM과 Nearest Neighbor Exclusive Circle (NNEC) 제약을 통합하여 점 레이블로부터 밀집 인스턴스 분할을 생성하는 Dense Point-to-Mask Optimization (DPMO) 방법을 제안합니다. DPMO와 수동 수정 과정을 통해 기존의 군중 데이터셋에 대한 점 레이블로부터 마스크 레이블을 얻습니다. 또한, 밀집된 군중 환경에서 인스턴스 분할을 예측하기 위해, 초기 점 예측 샘플링으로부터 최적의 예측 점을 선택하는 강화 학습 기반 포인트 선택 (Reinforced Point Selection, RPS) 프레임워크를 제안하며, Group Relative Policy Optimization (GRPO) 방식으로 훈련합니다. 광범위한 실험을 통해 ShanghaiTech, UCF-QNRF, JHU-CROWD++, NWPU-Crowd 데이터셋에서 최첨단 수준의 군중 인스턴스 분할 성능을 달성했습니다. 더불어, 마스크에 의해 감독되는 새로운 손실 함수를 설계하여 다양한 모델의 계수 성능을 향상시켰으며, 이는 마스크 레이블이 계수 정확도를 높이는 데 중요한 역할을 한다는 것을 보여줍니다.

Original Abstract

Crowd instance segmentation is a crucial task with a wide range of applications, including surveillance and transportation. Currently, point labels are common in crowd datasets, while region labels (e.g., boxes) are rare and inaccurate. The masks obtained through segmentation help to improve the accuracy of region labels and resolve the correspondence between individual location coordinates and crowd density maps. However, directly applying currently popular large foundation models such as SAM does not yield ideal results in dense crowds. To this end, we first propose Dense Point-to-Mask Optimization (DPMO), which integrates SAM with the Nearest Neighbor Exclusive Circle (NNEC) constraint to generate dense instance segmentation from point annotations. With DPMO and manual correction, we obtain mask annotations from the existing point annotations for traditional crowd datasets. Then, to predict instance segmentation in dense crowds, we propose a Reinforced Point Selection (RPS) framework trained with Group Relative Policy Optimization (GRPO), which selects the best predicted point from a sampling of the initial point prediction. Through extensive experiments, we achieve state-of-the-art crowd instance segmentation performance on ShanghaiTech, UCF-QNRF, JHU-CROWD++, and NWPU-Crowd datasets. Furthermore, we design new loss functions supervised by masks that boost counting performance across different models, demonstrating the significant role of mask annotations in enhancing counting accuracy.

0 Citations
0 Influential
7 Altmetric
35.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!