2608.07180v1 Aug 07, 2026 cs.LG

Momba: 네트워크 현대화가 다중 목표 강화 학습 성능 향상에 미치는 영향

Momba: Network Modernization Improves Multi-Objective Reinforcement Learning

J. Pajarinen
J. Pajarinen
Citations: 2,576
h-index: 21
Adam Štafa
Adam Štafa
Citations: 0
h-index: 0
Santeri Heiskanen
Santeri Heiskanen
Citations: 1
h-index: 1
Petr Novotný
Petr Novotný
Citations: 616
h-index: 13

최근 딥 러닝 기반 강화 학습 연구에서 신경망 구조 개선이 기본 알고리즘을 변경하지 않고도 샘플 효율성과 최종 성능 면에서 상당한 이점을 가져다줄 수 있음이 밝혀졌습니다. 반면, 상반된 목표 간의 균형을 찾는 정책 집합을 발견하는 것을 목표로 하는 다중 목표 강화 학습(MORL) 연구는 주로 알고리즘 혁신에 집중되어 왔으며, 신경망 구조 분야는 상대적으로 탐구가 부족했습니다. MORL 알고리즘은 일반적으로 무역 관계를 나타내는 데 간단한 피드포워드 네트워크를 사용하는데, 이는 무역 관계에 따라 최적의 정책과 가치 함수가 크게 달라질 수 있다는 점을 고려할 때, 보다 표현력이 뛰어난 함수 근사기를 사용할 경우 알고리즘 성능이 향상될 수 있는지에 대한 의문을 제기합니다. 본 논문에서는 신경망 설계 분야의 최근 발전 사항인 (i) 관측 및 특징 정규화, (ii) 가중치 정규화, 그리고 (iii) 엔트로피 규제 MORL 알고리즘을 사용하여 분포적 보상을 모델링하는 방법을 제시합니다. 표준 연속 제어 벤치마크를 사용한 실험 결과는 이러한 변경 사항이 기본 알고리즘에 큰 변화 없이도 생성된 솔루션 집합의 품질을 크게 향상시킴을 보여줍니다.

Original Abstract

Recent advances in deep reinforcement learning (RL) have shown that improving neural network architectures can yield substantial gains in sample efficiency and asymptotic performance without altering the underlying algorithms. In contrast, work on multi-objective reinforcement learning (MORL), which aims to discover a set of policies that balance trade-offs among conflicting objectives, has predominantly focused on algorithmic innovations, leaving the area of architectures underexplored. While the optimal policies and value functions can differ significantly depending on the trade-offs, MORL algorithms commonly represent them with simple feedforward networks conditioned on the trade-off. This raises the question of whether the performance of the algorithms could be improved with more expressive function approximators. In this paper, we integrate recent advances in neural network design: (i) observation and feature normalization, (ii) weight normalization, and (iii) modeling of distributional returns with an entropy-regularized MORL algorithm. The empirical results across standard continuous control benchmarks demonstrate that these changes substantially improve the quality of the produced solution sets without requiring major changes to the underlying algorithm.

0 Citations
0 Influential
10.5 Altmetric
52.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!