이분법적 도덕 판단을 넘어: 인공지능에서의 윤리적 다원주의 모델링
Beyond Binary Moral Judgment: Modeling Ethical Pluralism in AI
사회적으로 중요한 의사결정 과정에서 다양한 수준의 능력을 가진 인공지능 시스템이 점점 더 많이 활용되고 있습니다. 그러나 자율 시스템이 널리 보급되었음에도 불구하고, 대부분의 자율적인 도덕적 의사결정 방식은 스칼라 또는 이분법적 판단에 의존합니다. 이러한 방법은 적절한 도덕적 논리를 제공하기에는 부족하며, 책임성을 뒷받침하기 위해 필요한 맥락적 정보와 이론적 정보를 충분히 포함하지 못합니다. 이에, 우리는 윤리적 다원주의를 정규화된 윤리 이론의 분포로 모델링하는 프레임워크를 제안합니다. 이 프레임워크는 이러한 이론들을 통합하는 정규 윤리 심플렉스를 도입합니다. 또한, 15개의 세분화된 하위 이론에 걸쳐 450개의 사례 데이터셋을 구축하여 스태킹 앙상블 학습에 활용했습니다. 이 데이터셋은 자연어로 표현된 윤리적 딜레마와 함께 추출된 맥락적 특징 정보를 포함하고 있습니다. 심플렉스 구현은 두 개의 스트림으로 구성된 정규-의미론 아키텍처를 통해 이루어졌습니다. 이후, 정규 정보와 순차적인 스태킹 앙상블을 결합하여 세 가지 주요 이론(결과주의, 덕 윤리 및 의무론)과 15개의 하위 범주에 가장 적합한 모델을 학습합니다. 실험 결과, 맥락적 정보와 정규적 사전 지식을 의미론적 임베딩과 통합하면 분류 성능이 크게 향상되며, 정확도가 88.89%를 달성했습니다. 또한, 구조화된 윤리적 표현은 유사 추론 외에도 기여하며, 선택된 스태킹 아키텍처는 점진적인 학습을 통해 최상의 결과를 제공한다는 것을 ablation 연구를 통해 확인했습니다. 마지막으로, 엔트로피, 신뢰도 및 시각화를 통해 윤리적 다원주의를 분석합니다. 따라서, 윤리적 다원주의를 확률적 정규 분포로 모델링함으로써 인간과 유사한 도덕적 추론, 윤리적 의견 불일치 분석, 그리고 미래의 인공지능 시스템과의 조화로운 발전을 지원할 수 있습니다.
Critical decision-making in socially consequential spaces is increasingly involving AI systems at varying capacities. Yet, despite the ubiquity of autonomous systems, most approaches to handling autonomous moral decision-making resort to scalar or binary judgments. These methods are insufficient for acceptable moral reasoning, as they provide little explanation, leaving out imperative contextual and theoretical information that must be included to support accountability. For this, we propose a framework to model moral reasoning as a distribution over normative ethical theories or ethical pluralism. We introduce a normative ethics simplex that integrates these theories. A benchmark of 450 cases across 15 fine-grained subtheories was also prepared for the purposes of stacked ensemble learning. These cases describe ethical dilemmas in natural language and have associated extracted contextual features. The implementation of the simplex was achieved via a two-stream normative-semantic architecture. This is followed by the fusion of normative information and a sequential, stacking ensemble to learn the best fit of the three broad theories: consequentialism, virtue ethics, and deontology, and the 15 subcategories. Our experiments demonstrate that the integration of contextual and normative priors with the semantic embeddings significantly improves the performance of the classification, displaying an accuracy of 88.89%. We conducted ablation studies to show that structured ethical representations contribute beyond analogical reasoning, and the chosen stacking architecture gives the best results due to the gradual learning of granularity. Ethical pluralism is also analyzed through entropy, confidence, and visualization. Thus, modeling ethical pluralism as a probabilistic normative distribution supports human-like moral reasoning, ethical disagreement analysis, and future alignment in AI systems.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.