비기능적 요구사항 평가를 위한 다중 회화 LLM 대화의 정확성과 만족도
Accuracy and Satisfaction in Multi-Turn LLM Dialogues for NFR Assessment
LLM 기반 대화형 어시스턴트는 소프트웨어 개발자들에게 점점 더 보편적인 도구가 되었지만, 현재 평가 벤치마크는 기능적 정확성만을 중점적으로 다룹니다. 이는 비기능적 요구사항(NFR)을 처리할 때 이러한 대화의 품질과 정확성을 평가하는 데 중요한 격차를 만듭니다. NFR은 본질적으로 모호하고 상황에 따라 달라지며 프로그램의 여러 부분을 포함합니다. 이러한 시스템이 NFR에 대한 협업적인 추론을 얼마나 잘 지원하는지를 평가하려면, 시스템 출력의 정확성과 다중 회화 상호작용의 품질을 모두 포괄할 수 있는 방법을 사용해야 합니다. 본 논문에서는 Health Insurance Portability and Accountability Act (HIPAA) 규정 준수 분야에서 개발자와 LLM 기반 에이전트 간의 다중 회화에 대한 정확성과 품질을 조사합니다. 49명의 프로그래머를 대상으로 GitHub Copilot과 상호 작용하며, HIPAA 관련 NFR 148개를 iTrust 코드베이스(HIPAA 규정 준수를 위해 설계된 시스템)에 대해 평가했습니다. 평가 기준은 요구사항 만족도 수준, 추론 및 코드 위치 파악의 세 가지 측면입니다. 연구 결과, 개발자들은 LLM의 평가와 대체로 동의하지만, 전문가의 정확한 답변과 비교했을 때 정확도는 낮은 것으로 나타났습니다. 사용자 만족도를 모델링한 결과, 시스템 응답이 길수록, 정보 제공 횟수가 많을수록 사용자 만족도가 낮아지는 반면, 적극적인 상호작용은 사용자 만족도를 높이는 것으로 확인되었습니다. 본 연구의 결과는 NFR 평가를 지원하는 LLM 기반 대화형 시스템 설계에 대한 통찰력을 제공합니다.
LLM-based dialogue assistants have become mainstream tools for software developers, yet current evaluation benchmarks focus exclusively on functional correctness. This leaves a critical gap in assessing the quality and accuracy of these conversations when handling Non-Functional Requirements (NFRs), which are inherently vague, context-dependent, and involve many parts of a program. Evaluating how well these systems support collaborative reasoning about NFRs requires methods that go beyond single-turn accuracy to capture both the correctness of the system's outputs and the quality of the multi-turn interaction. In this paper, we investigate the accuracy and quality of multi-turn conversations between developers and an LLM-based agent in the domain of Health Insurance Portability and Accountability Act (HIPAA) regulatory compliance. We hired 49 programmers to interact with GitHub Copilot to assess 148 HIPAA-derived NFRs against the iTrust codebase, a system designed to comply with HIPAA regulations, across three dimensions: requirement satisfaction level, reasoning, and code localization. We find that developers tend to agree with LLM assessments, but accuracy against expert ground truth is low. We model user satisfaction and find that longer system responses and more information-providing turns negatively affect user satisfaction, whereas proactive interactions positively affect it. Our findings provide insights for designing LLM-based dialogue systems that support NFR assessment.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.