당신의 인공지능 여행 에이전트가 투우를 예약해 드릴 수도 있습니다: 최첨단 AI 모델의 잠재적 동물 복지 문제를 평가하는 새로운 방법
Your AI Travel Agent Would Book You a Bullfight: An Agentic Benchmark for Implicit Animal Welfare in Frontier AI Models
인공지능 에이전트는 단순한 조언자 역할을 넘어, 사용자를 대신하여 여행을 예약하고, 메뉴를 계획하며, 구매 업무를 수행하는 등 실제 행동으로 이어지는 단계로 발전하고 있습니다. 기존의 AI 및 동물 복지에 대한 평가지표는 주로 모델의 텍스트 응답을 질문-응답 형태로 평가하는데, 이러한 평가 방식이 실제로 모델이 도구를 사용하여 행동해야 하는 상황에서 나타나는 동물 복지 관련 판단 능력에 제대로 반영되는지 여부는 불분명합니다. 본 연구에서는 TAC(Travel Agent Compassion)라는 새로운 유형의 에이전트 기반 평가지표를 제시합니다. 이는 AI 에이전트가 사용자를 대신하여 특정 행동을 수행할 때, 동물 착취와 관련된 옵션을 회피하는지를 측정하는 지표입니다. TAC는 6가지 동물 착취 범주에 걸쳐 작성된 12개의 여행 예약 시나리오를 제공하며, 가격, 평점, 위치 등의 변수를 통제하기 위해 총 48개의 샘플을 사용합니다. 우리는 4개 기관의 7개의 최첨단 모델을 평가했습니다. 모든 모델의 성능은 무작위 추정치인 64%에 미치지 못했으며, 가장 우수한 모델(Claude Opus 4.7) 역시 53%의 정확도를 기록했습니다. 시스템 프롬프트에 단 하나의 동물 복지에 대한 문장을 추가했을 때, Claude 및 GPT-5.5 모델은 47~63%의 성능 향상을 보였고, GPT-5.2는 26%, DeepSeek와 Gemini는 12% 미만의 성능 향상만 보였습니다. 또한, Gemini 2.5 Flash Lite를 판별기로 사용하여 상위 두 모델의 기본 조건 하에서의 288개 대화 기록을 검토한 결과, 평가에 대한 인지 여부를 판단할 수 있는 기록은 단 하나도 발견되지 않았습니다. 본 연구는 문화 영역 간의 범주별 차이, 텍스트 응답 기반 동물 복지 평가지표의 한계, 그리고 EU 일반 목적 AI 실천 강령의 체계적 위험 관리 프레임워크에 대한 시사점을 논의합니다.
AI agents are moving from advisors to actors, booking travel, planning menus, and running procurement on behalf of users. Existing benchmarks for AI and animal welfare evaluate model text responses to question-answer prompts, leaving open whether the welfare reasoning surfaced in those responses transfers to agentic deployment where the model must take actions with tools. We introduce TAC (Travel Agent Compassion), the first agentic benchmark measuring whether AI agents avoid options involving animal exploitation when acting on behalf of users. TAC presents an AI agent with twelve hand-authored travel booking scenarios across six categories of animal exploitation, augmented to forty-eight samples to control for price, rating, and position confounds. We evaluate seven frontier models from four labs. Every model scores below the chance level of sixty-four percent, with the best performer (Claude Opus 4.7) at fifty-three percent. A single welfare-aware sentence in the system prompt yields gains of forty-seven to sixty-three percentage points in Claude and GPT-5.5, twenty-six points in GPT-5.2, and under twelve points in DeepSeek and Gemini. An auxiliary Inspect Scout audit of 288 base-condition transcripts from the top two performers, using Gemini 2.5 Flash Lite as judge, flags zero transcripts for evaluation awareness, suggesting the below-chance rates do not stem from the models recognising the evaluation. We discuss implications for category-level variation across cultural domains, the limits of text-response welfare benchmarks, and the EU General-Purpose AI Code of Practice systemic risk framework.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.