UrbanAgent: 시스템 간 연동을 위한 도구 기반 에이전트 - 도시 업무 활용
UrbanAgent: A Tool-Augmented Agent for Cross-System Urban Tasks
현대 도시에서는 운영에 필요한 디지털 서비스가 증가하고 있지만, 여전히 주민들의 일상적인 요구를 충족시키는 데 어려움이 있습니다. 서비스는 단편화되어 있고 상호 운용성이 낮아 사용자에게 과도한 부담을 줍니다. 기존의 디지털 플랫폼, 도시 기반 모델 및 지능형 어시스턴트는 각각 도시 업무의 특정 측면만을 다룹니다. 하지만 이러한 시스템들은 복잡한 자연어 요청을 실행 가능한 시스템 간 워크플로우로 안정적으로 변환하는 데 어려움을 겪습니다. 본 연구에서는 시스템 간 연동이 필요한 도시 업무를 위한 도구 기반 에이전트 프레임워크인 Urban-Agent를 제안합니다. Urban-Agent는 대규모 언어 모델의 인지 및 추론 능력과 코드 실행, API 호출, 모델 컨텍스트 프로토콜을 지원하는 도구 세트를 결합합니다. 단일 적응형 폐쇄 루프 시스템을 통해 Urban-Agent는 작업 수행 전에 누락된 정보를 명확히 하고, 도구 사용을 실시간 관찰에 기반하여 수행하며, 최종 응답이 관찰된 증거와 작업 제약 조건에 부합하도록 합니다. 평가 격차를 해소하기 위해, 본 연구에서는 시스템 간 연동 도시 요청을 위한 벤치마크인 Urban-Eval을 소개합니다. 기존의 벤치마크가 일반적인 도구 사용 또는 도시 지식 및 추론 능력을 평가하는 반면, Urban-Eval은 작업 결과와 실행 품질 모두를 평가합니다. 여기에는 필요한 도구 범위, 의존성 유효성 및 증거 추적 가능성이 포함됩니다. 실험 결과에 따르면 Urban-Agent는 71%의 작업 성공률을 달성했으며, 이는 가장 강력한 기준 모델보다 10점 높은 수치입니다. 이러한 우위는 GPT-5-mini, Gemini-2.5-flash, DeepSeek-V4-flash 및 Qwen3-235B-A22B에서 모두 관찰되었습니다.
Modern cities rely on an increasing number of digital services to operate, but residents' daily needs are still difficult to meet. Services are fragmented and have little interoperability, placing a heavy operational burden on users. Existing digital platforms, urban foundation models, and intelligent assistants each address only isolated aspects of an urban task. But they struggle to reliably convert complex natural-language requests into executable cross-system workflows. We propose Urban-Agent, a tool-augmented agent framework for cross-system urban tasks. It couples the cognitive and reasoning capabilities of a large language model with a tool-set supporting code execution, API calls, and Model Context Protocol. Through one adaptive closed loop, it clarifies missing information before acting, grounds tool use in live observations, and aligns the final response with observed evidence and task constraints. To address the evaluation gap, we introduce Urban-Eval, a benchmark specifically designed for cross-system urban request. Unlike prior benchmarks that assess either general tool use or urban knowledge and reasoning, Urban-Eval evaluates both task results and execution quality, including required tool coverage, dependency validity, and evidence traceability. Experimental results indicate that Urban-Agent reaches a 71% task success rate, 10 points above the strongest baseline. This lead holds across GPT-5-mini, Gemini-2.5-flash, DeepSeek-V4-flash, and Qwen3-235B-A22B.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.