CogEEGAgent: 근거 기반 실행과 선택 인식 검증을 통한 자율적인 인지 EEG 분석 시스템
CogEEGAgent: Toward Autonomous Cognitive EEG Analysis with Grounded Execution and Selection-Aware Verification
인지 연구에서의 뇌전위(EEG) 분석은 전문 지식을 필요로 하며, 대조 조건, 채널, 시간 창, 통계적 검정 등 다양한 측면에서 신중한 선택이 요구됩니다. LLM 에이전트는 다양한 자연어 질문을 분석 결정으로 변환하여 자동화를 위한 유연한 인터페이스를 제공할 수 있습니다. 그러나 단순히 보고서를 생성하는 것만으로는 에이전트가 요청된 분석을 실제로 수행했는지, 또는 적응적 검색에 의존하지 않고 검증 주장을 독립적으로 평가했는지 확인할 수 없습니다. 본 연구에서는 MNE-Python 기반의 인지 EEG 분석 에이전트인 CogEEGAgent를 소개합니다. 이 시스템은 뇌전위 분석에 특화된 과학적 프레임워크를 사용하여 의미론적 부분과 과학적 권한을 분리합니다. LLM은 사용자의 의도를 해석하고 등록된 분석을 제안하며, 결정적인 구성 요소는 타입 검사를 통해 계약을 유효성 검사하고, 확인 접근을 통제하며, 증거 기반의 결과 공개를 허용합니다. 미리 정의된 테스트 데이터셋에서 CogEEGAgent는 동일한 기능의 결정론적 시스템보다 자연어 질문을 등록된 분석으로 더 정확하게 매핑합니다. 또한, 사전 설정된 안전 장치를 통해 두 시스템 모두 필요한 경우 분석 수행을 중단할 수 있습니다. 외부 모델이 작성하고 결과에 영향을 받지 않는 캠페인에서, CogEEGAgent는 지원되는 분석 결과를 참가자별로 독립적으로 확인하여 공개하며, 미리 정의된 위험 요소 및 재사용 요청을 차단합니다. 정책 스트레스 테스트를 통해 확인 절차가 적응적 검색으로 인해 발생할 수 있는 오탐을 줄이는 것을 확인할 수 있었습니다. 이러한 연구 결과들은 인지 EEG 워크플로우에 대한 제한적인 자율성과 감사 가능한 자동화 프레임워크를 구축하는 데 기여하며, 과학적 에이전트가 유연한 자연어 이해 능력과 추론 및 결과 공개에 대한 안전 장치를 결합할 수 있는 방법을 보여줍니다.
Electroencephalography (EEG) analysis in cognitive studies requires specialized expertise and involves many defensible choices over contrasts, channels, time windows, and statistical tests. LLM agents can translate varied natural-language questions into analysis choices, offering a flexible interface for automation. Yet fluent reports alone cannot establish that an agent selected the requested analysis or evaluated a confirmatory claim independently of adaptive search. We present CogEEGAgent, a cognitive-EEG analysis agent grounded in MNE-Python. Its EEG-specific scientific harness separates semantic from scientific authority. The LLM interprets intent and proposes registered analyses, while deterministic components validate typed contracts, control confirmation access, and authorize evidence-bound release. On a prespecified routing benchmark, CogEEGAgent maps language to registered analyses more accurately than a matched deterministic router, while matched preflight makes both systems abstain whenever required. In an externally model-authored, outcome-blind campaign, the complete system releases supported analyses with participant-disjoint confirmation and blocks prespecified capability hazards and lifecycle-reuse requests. Policy stress testing shows that held-out confirmation curbs false positives from uncorrected adaptive search. Together, these studies establish bounded autonomy and an auditable automation framework for cognitive-EEG workflows. More broadly, they show how scientific agents can combine flexible language understanding with fail-closed control over inference and release.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.