환경 데이터 관리를 위한 견고한 다중 에이전트 워크플로우 연구
Exploring Robust Multi-Agent Workflows for Environmental Data Management
LLM 기반 에이전트를 환경 분야의 FAIR 데이터 관리에 적용하는 것은 매우 유망합니다. 이러한 에이전트는 운영 지식을 외부화하고, 다양한 데이터와 변화하는 규범에 걸쳐 데이터 큐레이션을 확장할 수 있습니다. 그러나, 결정론적인 구성 요소를 확률적인 워크플로우로 대체하면 오류 발생 방식이 바뀝니다. LLM 파이프라인은 표면적인 검사를 통과하지만, 정확하지 않은 결과를 생성할 수 있으며, 이는 DOI 발행 및 공개 배포와 같은 되돌릴 수 없는 작업으로 이어질 수 있습니다. 본 논문에서는 캠퍼스 전체 저장 인프라에 배포된 환경 연구용 생산 데이터 관리 시스템인 EnviSmart를 소개합니다. EnviSmart는 다음과 같은 두 가지 메커니즘을 통해 안정성을 아키텍처의 중요한 속성으로 간주합니다. 첫째, 행동(규제 제약), 도메인 지식(검색 가능한 컨텍스트), 그리고 기술(도구 사용 절차)을 지속적이고 상호 연결된 아티팩트로 외부화하는 세 가지 지식 아키텍처입니다. 둘째, 결정론적인 검증기와 감사된 핸드오프를 사용하여 신뢰 경계에서 안전 정지 기능을 복원하는 역할 분담 다중 에이전트 설계입니다. 본 논문에서는 두 가지 실제 배포 사례를 비교합니다. 대학 GIS 센터의 생태 보관소(849개의 큐레이션된 데이터 세트)는 단일 에이전트의 기준 사례로 사용됩니다. SF2Bench는 2,452개의 모니터링 스테이션과 39년에 걸쳐 8,557개의 게시된 파일을 포함하는 복합 홍수 벤치마크이며, 이는 다중 에이전트 워크플로우를 검증합니다. 다중 에이전트 접근 방식은 효율성을 향상시켰습니다. 단일 운영자가 반복적인 아티팩트 재사용을 통해 2일 만에 완료했으며, 또한 안정성도 향상시켰습니다. 감사된 핸드오프는 모든 2,452개의 스테이션에 영향을 미치는 좌표 변환 오류를 게시 전에 감지하고 차단했습니다. 대표적인 사례(ISS-004)는 경계 기반 격리를 통해 10분 내에 감지되었으며, 사용자에게 미치는 영향은 0이었고, 80분 내에 해결되었습니다. 본 논문은 PEARC 2026에 채택되었습니다.
Embedding LLM-driven agents into environmental FAIR data management is compelling - they can externalize operational knowledge and scale curation across heterogeneous data and evolving conventions. However, replacing deterministic components with probabilistic workflows changes the failure mode: LLM pipelines may generate plausible but incorrect outputs that pass superficial checks and propagate into irreversible actions such as DOI minting and public release. We introduce EnviSmart, a production data management system deployed on campus-wide storage infrastructure for environmental research. EnviSmart treats reliability as an architectural property through two mechanisms: a three-track knowledge architecture that externalizes behaviors (governance constraints), domain knowledge (retrievable context), and skills (tool-using procedures) as persistent, interlocking artifacts; and a role-separated multi-agent design where deterministic validators and audited handoffs restore fail-stop semantics at trust boundaries before irreversible steps. We compare two production deployments. The University's GIS Center Ecological Archive (849 curated datasets) serves as a single-agent baseline. SF2Bench, a compound flooding benchmark comprising 2,452 monitoring stations and 8,557 published files spanning 39 years, validates the multi-agent workflow. The multi-agent approach improved both efficiency - completed by a single operator in two days with repeated artifact reuse across deployments - and reliability: audited handoffs detected and blocked a coordinate transformation error affecting all 2,452 stations before publication. A representative incident (ISS-004) demonstrated boundary-based containment with 10-minute detection latency, zero user exposure, and 80-minute resolution. This paper has been accepted at PEARC 2026.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.