Agents-K1: 에이전트 친화적인 지식 조율을 향하여
Agents-K1: Towards Agent-native Knowledge Orchestration
현재 LLM 기반 연구 에이전트는 에이전트 조율을 통해 발전해 왔지만, 과학적 지식 조율은 대부분 간과되고 있습니다. 기존 연구는 종종 논문을 초록, 표면 언급 및 단순한 exttt{cites} 연결로 축소하며, 과학적 추론에 필수적인 핵심 개체, 주장, 증거, 메커니즘 및 방법론 연관성을 누락시키는 경우가 많습니다. 이에, 우리는 raw 문서를 에이전트 친화적인 과학 지식 그래프로 변환하는 엔드투엔드 지식 조율 파이프라인인 extbf{Agents-K1}을 소개합니다. Agents-K1은 세 가지 구성 요소를 통합하며, 이는 통일된 이론적 기반 하에 작동합니다: 다섯 개의 모듈로 구성된 멀티모달 파서는 초록뿐만 아니라 전체 논문에 걸쳐 개체, 멀티모달 증거, 인용 및 유형화된 개체 간 관계를 포착합니다. 40억 개의 매개변수를 가진 정보 추출 백본은 규칙 기반 보상을 사용하여 GRPO 방식으로 학습되었습니다. 또한, graphanything CLI는 웹 검색, 멀티모달 그래프 검색 및 문서 간 탐색을 통합하는 삼중 소스 에이전트 인터페이스입니다. 이 파이프라인을 활용하여 6개 분야의 246만 건의 과학 논문을 처리하고 extbf{Scholar-KG}를 생성했으며, 그 중 100만 건의 하위 집합을 공개합니다. 전체 Scholar-KG는 아래 SCP 링크를 통해 접근할 수 있습니다. 동일한 파이프라인은 일반 도메인 코퍼스 및 스키마 준수 데이터 합성에 적용될 수 있습니다. 광범위한 실험 결과, Agents-K1은 과학 정보 추출, 지식 그래프 구축 및 다중 호프 과학적 추론에서 뛰어난 성능을 달성하는 것으로 나타났습니다.
Current LLM-based research agents have advanced through agent orchestration, yet largely overlook scientific knowledge orchestration. Existing works often reduce papers to abstracts, surface mentions, and flat \texttt{cites} edges, omitting key entities, claims, evidence, mechanisms, and method lineages essential for scientific reasoning. To this end, we introduce \textbf{Agents-K1}, an end-to-end knowledge orchestration pipeline that converts raw documents into agent-native scientific knowledge graphs. Agents-K1 integrates three components under a unifying theoretical foundation: a multimodal parser whose five-module schema captures entities, multimodal evidence, citations, and typed inter-entity relations across the full paper rather than abstracts alone; a 4B information-extraction backbone trained with GRPO under a rule-based reward; and a graphanything CLI, a tri-source agent interface that unifies web search, multimodal graph retrieval, and cross-document traversal. On top of this, we process 2.46 million scientific papers across six subjects to produce \textbf{Scholar-KG}, of which we release a one-million-paper subset, and the full Scholar-KG is accessible via the SCP link below. The same pipeline can be extended to general-domain corpora and to schema-conformant data synthesis. Extensive experiments demonstrate that Agents-K1 achieves superior performance in scientific information extraction, knowledge graph construction, and multi-hop scientific reasoning.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.