2506.07223v2 Jun 08, 2025 cs.AI

먼저 반응하고 나중에 성찰하라: 동적 대응을 위한 지연 시간 인지 자율 에이전트

Reflex First, Reflect Later: Latency-Aware Embodied LLM Agents for Dynamic Response

Yan Zheng
Yan Zheng
Citations: 536
h-index: 13
Weidong Cai
Weidong Cai
Citations: 25
h-index: 3
Shunqi Mao
Shunqi Mao
Citations: 71
h-index: 5
Dingxin Zhang
Dingxin Zhang
University of Sydney
Citations: 81
h-index: 4

대규모 언어 모델(LLM)은 자율 에이전트의 계획 능력을 크게 향상시켜, 동적이고 안전이 중요한 환경에 배치할 수 있도록 합니다. 하지만 이러한 환경에서는 LLM 추론 지연이라는 중요한 문제가 발생합니다. 지연된 LLM 응답은 실시간 대응성을 저하시키고 에이전트의 추론을 빠르게 변화하는 환경 상태와 일치시키지 못하게 할 수 있습니다. 본 논문에서는 동적 환경에서 LLM 기반 자율 에이전트에 대한 추론 지연의 영향을 체계적으로 연구합니다. 우리는 추론 시간을 경과된 시뮬레이션 시간으로 매핑하여 계산 지연이 환경 변화 및 에이전트 결과에 직접적인 영향을 미치도록 하는 FPS 기반 시간 변환 메커니즘(TCM)을 소개합니다. 이 프로토콜을 HAZARD 환경에 적용하고, 응답 지연(RL)과 지연-행동 비율(LAR)을 사용하여 에이전트의 대응성을 평가합니다. 이러한 프레임워크를 바탕으로, 우리는 빠른 반사적 행동과 비동기 LLM 성찰을 통합하여 지연으로 인한 오류를 완화하는 Rapid-Reflex Async-Reflect Agent (RRARA)를 제안합니다. 또한, 모델의 고차원 추론 능력을 유지하면서 반복적인 LLM 호출을 줄이는 LLM 기반 사전 계획기를 도입했습니다. 실험 결과, 추론 지연을 고려하면 자율 에이전트의 성능에 상당한 변화가 있으며, RRARA는 의사 결정 품질과 대응성 간의 균형을 더 잘 제공한다는 것을 보여줍니다.

Original Abstract

Large language models (LLMs) have substantially improved the planning capabilities of embodied agents, enabling their deployment in dynamic and safety-critical environments. However, these settings expose a critical limitation: inference latency. Delayed LLM responses can weaken real-time responsiveness and misalign agent reasoning with rapidly changing environmental states. This paper systematically studies the impact of inference latency on LLM-based embodied agents in dynamic environments. We introduce an FPS-based Time Conversion Mechanism (TCM) that maps inference time to elapsed simulation time, allowing computational delays to directly affect environmental evolution and agent outcomes. We instantiate this protocol in HAZARD and introduce Response Latency (RL) and Latency-to-Action Ratio (LAR) to evaluate agent responsiveness. Building on this framework, we propose the Rapid-Reflex Async-Reflect Agent (RRARA), which integrates rapid reflexive actions with asynchronous LLM reflection to mitigate latency-induced errors. We further introduce an LLM-based PrePlanner that generates cached object-centric subgoals, reducing repeated LLM calls while retaining the model's high-level reasoning capability. Experiments show that accounting for inference latency substantially changes embodied-agent performance and that RRARA achieves a stronger balance between decision quality and responsiveness.

5 Citations
0 Influential
6.5 Altmetric
37.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!