2607.27686v1 Jul 30, 2026 cs.AI

AI 생성 답변에서 광고의 효과 측정 및 가격 결정

Evaluating and Pricing Advertisements in AI-Generated Responses

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Tonghan Wang
Tonghan Wang
Citations: 33
h-index: 4
John L. Turner-Smith
John L. Turner-Smith
Citations: 0
h-index: 0

검색이 점점 LLM(대규모 언어 모델) 기반 답변 엔진으로 이동함에 따라, 광고는 생성된 응답 자체에 내재화되고 있으며, 따라서 사용자 유용성과 상업적 가치 모두를 고려하여 평가되어야 합니다. 핵심적인 과제는 클릭률 의도 파악입니다. 행동 로그가 없으며, 인간 주석은 교정하기 어렵고, 최첨단 LLM 모델은 의도를 언어적 유창성과 혼동합니다. 이러한 문제점들이 누적되는데, 원칙에 따른 가격 결정은 지속적인 의도 신호를 전제로 하지만, 그러한 신호를 생성하는 것은 현재 이용 가능한 감독 데이터가 부족하기 때문입니다. 우리는 심리적으로 타당한 에이전트 시뮬레이션 프레임워크를 통해 이러한 감독 데이터를 구축하고, 이를 클릭률 의도를 예측하고 광고 품질의 세 가지 측면을 부드럽고 미분 가능한 추정치로 제공하는 효율적인 평가 모델로 변환합니다. 행동 변화를 기반으로 검증된 이 평가 모델은 관련성 민감도(79% vs. 60-67%)에서 최첨단 제로샷 평가 모델보다 우수하며, 콘텐츠 품질 저하 정도를 추적하고, 오류 없이 103개의 가상 제품에 적용 가능하며, 인간의 선호도와 86%의 일치율을 보입니다(5명의 어노테이터 기준). 평가 모델의 추정치를 기반으로 가격 결정 계층을 직접 구축하여, 정직한 입찰이 최적인 고유한 지불 규칙을 도출하고, 베스트-오브-k 할당에 적용하며, 비단조 할당으로 확장합니다. 동일한 미분 가능한 신호는 광고 생성의 학습 목표로 활용될 수 있습니다.

Original Abstract

As search increasingly shifts toward LLM-driven answer engines, advertising is becoming embedded within the generated response itself and should therefore be evaluated for both user utility and commercial value. The key challenge is click-through intent: behavioural logs are unavailable, human annotation resists calibration, and frontier LLM judges conflate intent with linguistic fluency. These gaps compound, as principled pricing presupposes a continuous intent signal, while generating such a signal presupposes supervision that is currently unavailable. We construct the missing supervision through a psychologically grounded agent simulation framework, and distil it into a parameter-efficient evaluator that predicts click-through intent, together with the three companion dimensions of ad quality, as smooth, differentiable estimates. Validated through sign-certain behavioural perturbations, the evaluator surpasses frontier zero-shot judges on relevance sensitivity (79% versus 60-67%), tracks graded content degradation, generalises without error to 103 fictional products, and agrees with human preference in 86% of pairwise judgements across five annotators, with agreement rising in the evaluator's confidence. Upon its estimates we build the pricing layer directly, deriving the unique payment rule under which truthful bidding is optimal, demonstrating it on a best-of-k allocation, and extending the mechanism to non-monotone allocations. The same differentiable signal stands ready as a training objective for ad generation.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!