2604.03926v1 Apr 05, 2026 cs.AI

CODE-GEN: 인간-루프 기반 RAG (검색 증강 생성) 에이전트 AI 시스템을 활용한 객관식 문제 생성

CODE-GEN: A Human-in-the-Loop RAG-Based Agentic AI System for Multiple-Choice Question Generation

Chaoli Wang
Chaoli Wang
Citations: 1,308
h-index: 20
Xiaojing Duan
Xiaojing Duan
Citations: 125
h-index: 6
Frederick Nwanganga
Frederick Nwanganga
Citations: 50
h-index: 4

본 논문에서는 CODE-GEN을 소개합니다. CODE-GEN은 인간의 개입을 통해 검색 증강 생성(RAG) 기반 에이전트 AI 시스템으로, 학생들의 코딩 추론 및 이해 능력을 향상시키기 위해 문맥에 맞는 객관식 문제를 생성합니다. CODE-GEN은 에이전트 AI 아키텍처를 사용하며, Generator 에이전트는 특정 강의의 학습 목표에 부합하는 객관식 코딩 이해 문제를 생성하고, Validator 에이전트는 7가지 교육적 측면에서 콘텐츠 품질을 독립적으로 평가합니다. 두 에이전트 모두 계산 정확성을 높이고 코드 출력을 검증하는 데 특화된 도구를 사용합니다. CODE-GEN의 효과성을 평가하기 위해, 6명의 전문가가 생성된 288개의 문제를 평가하는 실험을 진행했습니다. 전문가들은 총 2,016개의 인간-AI 평가 쌍을 생성하여, Validator의 평가에 대한 동의 또는 비동의 여부를 표시하고, 131개의 질적 피드백을 제공했습니다. 전문가 평가 결과, CODE-GEN은 높은 성능을 보였으며, 인간이 검증한 성공률은 7가지 교육적 측면에서 79.9%에서 98.6%에 달했습니다. 질적 피드백 분석 결과, CODE-GEN은 문제의 명확성, 코드의 유효성, 개념의 일치성, 정답의 유효성과 같이 계산적 검증 및 명시적인 기준에 적합한 측면에서 높은 신뢰성을 보였습니다. 반면, 교육적으로 의미 있는 오답을 설계하고 이해를 강화하는 고품질 피드백을 제공하는 것과 같이 더 깊은 교육적 판단이 필요한 측면에서는 여전히 인간 전문가의 역할이 중요합니다. 이러한 결과는 AI 지원 교육 콘텐츠 생성에서 인간과 AI의 노력을 전략적으로 배분하는 데 도움이 됩니다.

Original Abstract

We present CODE-GEN, a human-in-the-Loop, retrieval-augmented generation (RAG)-based agentic AI system for generating context-aligned multiple-choice questions to develop student code reasoning and comprehension abilities. CODE-GEN employs an agentic AI architecture in which a Generator agent produces multiple-choice coding comprehension questions aligned with course-specific learning objectives, while a Validator agent independently assesses content quality across seven pedagogical dimensions. Both agents are augmented with specialized tools that enhance computational accuracy and verify code outputs. To evaluate the effectiveness of CODE-GEN, we conducted an evaluation study involving six human subject-matter experts (SMEs) who judged 288 AI-generated questions. The SMEs produced a total of 2,016 human-AI rating pairs, indicating agreement or disagreement with the assessments of Validator, along with 131 instances of qualitative feedback. Analyses of SME judgments show strong system performance, with human-validated success rates ranging from 79.9% to 98.6% across the seven pedagogical dimensions. The analysis of qualitative feedback reveals that CODE-GEN achieves high reliability on dimensions well suited to computational verification and explicit criteria matching, including question clarity, code validity, concept alignment, and correct answer validity. In contrast, human expertise remains essential for dimensions requiring deeper instructional judgment, such as designing pedagogically meaningful distractors and providing high-quality feedback that reinforces understanding. These findings inform the strategic allocation of human and AI effort in AI-assisted educational content generation.

0 Citations
0 Influential
10 Altmetric
50.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!