2606.18976v1 Jun 17, 2026 cs.SE

CAPRA: 다중 에이전트 기반 LLM 시스템을 활용한 소프트웨어 아키텍처 결과물에 대한 피드백 확장

CAPRA: Scaling Feedback on Software Architecture Deliverables with a Multi-Agent LLM System

N. Caselli
N. Caselli
Citations: 1,012
h-index: 16
Marco Becattini
Marco Becattini
Citations: 33
h-index: 3
Matteo Minin
Matteo Minin
Citations: 0
h-index: 0
Roberto Verdecchia
Roberto Verdecchia
Citations: 73
h-index: 3
Enrico Vicario
Enrico Vicario
Citations: 71
h-index: 5

소프트웨어 공학 교육에서 코드 평가 및 에세이 채점 자동화는 상당한 발전을 이루었지만, 구조적 완전성 및 요구사항 추적성을 분석해야 하는 소프트웨어 아키텍처 결과물의 검토는 아직 완전히 자동화되지 않았습니다. 대규모 언어 모델(LLM)을 이러한 작업에 적용하려면 학생들에게 정확하고 신뢰할 수 있는 기술적 피드백을 제공하기 위한 견고한 시스템 아키텍처가 필요합니다. 본 논문에서는 CAPRA (Configurable Architecture Proficiency Report Assessment)라는 다중 에이전트 기반 LLM 시스템을 소개합니다. CAPRA는 소프트웨어 아키텍처 결과물을 분석하여 개인화되고 템플릿에 부합하는 LaTeX 형식의 피드백을 생성합니다. 핵심 설계 요소로서, CAPRA는 여러 개의 전문 에이전트를 조정하고, PyMuPDF 및 비전 기능을 갖춘 LLM(특히 gpt-4o)을 사용하여 텍스트와 UML 다이어그램을 분석하는 Python 기반 마이크로 서비스를 활용하여 멀티모달 문서 추출을 수행합니다. 교육적 신뢰성을 확보하고 환각 현상을 완화하기 위해, CAPRA는 정규화된 레벤슈타인 거리(Levenshtein distance)를 사용한 퍼지 매칭을 통해 증거를 명확하게 제시하는 단계를 도입하고, 교차 검증, 중복 제거 및 통합 기능을 수행하는 ConsistencyManager 에이전트를 활용합니다. 시스템 성능은 추출 완전성, 기능 유효성, 문제 식별 및 심각도 탐지, 추천의 구체성과 추적 가능성, 템플릿 및 어조 준수 등 8가지 기준으로 구성된 이진 평가 체계를 사용하여 평가되었습니다. 10명의 학생 보고서에 대한 초기 실험 결과, CAPRA는 엄격한 두 명 평가자 결합 규칙 하에서 평가 기준의 88.8%를 충족했으며, 인간 평가자와 중간 수준의 평가자 간 일치도(kappa = 0.582)를 달성하고 각 보고서를 평균 4분 이상 처리했습니다. 이러한 결과는 LLM 기반 아키텍처 피드백의 가능성을 뒷받침하지만, 주관적인 평가 측면에서는 인간의 감독이 여전히 필수적입니다.

Original Abstract

Automated assessment in software engineering education has advanced significantly for code grading and essay scoring. However, reviewing software architecture deliverables, which requires analyzing structural completeness and requirements traceability, has not yet been fully automated. Applying Large Language Models (LLMs) to this task requires robust architectures to ensure technical feedback is accurate and reliable for students. This paper presents CAPRA (Configurable Architecture Proficiency Report Assessment), a multi-agent LLM system that analyzes software architecture deliverables to generate personalized, template-compliant LaTeX feedback. As a core design choice, CAPRA coordinates multiple specialized agents and employs a Python-based microservice for multi-modal document extraction, utilizing PyMuPDF and vision-enabled LLMs (specifically gpt-4o) to parse text and UML diagrams. To ensure educational reliability and mitigate hallucinations, CAPRA introduces a deterministic Evidence Anchoring step using fuzzy matching via normalized Levenshtein distance, along with a ConsistencyManager agent that cross-verifies, deduplicates, and merges findings. System performance is assessed using a structured eight-criterion binary evaluation taxonomy covering: (i) extraction completeness, (ii) feature validation, (iii) issue grounding and severity detection, (iv) recommendation specificity and traceability, and (v) template and tone compliance. A preliminary empirical evaluation on 10 student reports shows that CAPRA satisfied 88.8% of the evaluated criteria under a strict two-rater aggregation rule, achieved moderate inter-rater agreement with human evaluators (kappa = 0.582), and processed each report in slightly over 4 minutes. While these results support the viability of LLM-supported architectural feedback, human oversight remains essential for subjective assessment dimensions.

0 Citations
0 Influential
8 Altmetric
40.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!