LLM이 생성한 기술(Skill)은 데이터 과학 분야 AI 전문가를 더 향상시킬 수 있는가? 데이터 과학 워크플로우에 따른 구성 요소 제거 실험
Do LLM-Generated Skills Make Better AI Data Scientists? A Component Ablation Across Data-Science Workflows
제품 데이터 과학자들은 종종 LLM 기반 에이전트에게 데이터 정제, SQL 작성, 통계 검정 선택, 결과 형식화 등 반복적인 작업 수행을 요청합니다. 재사용 가능한 기술 파일은 특정 작업 그룹에 대한 지침을 패키징하여 처음부터 프롬프트를 작성하는 것을 방지하기 위한 것입니다. 전문가가 작성한 기술은 고품질 지침을 포함할 수 있지만, 많은 데이터 과학 작업 그룹에 걸쳐 이를 작성하고 유지 관리하는 것은 수동적인 병목 현상을 야기합니다. 본 연구에서는 LLM이 생성한 기술이 얼마나 유용한 대안이 될 수 있는지 조사합니다: 즉, 단일 프롬프트만 사용하는 것보다 성능 향상을 가져다주는가? 데이터 준비, 데이터 추출, 통계 분석 및 보고라는 네 가지 라이프사이클 단계에 걸쳐 하나의 생성된 기술을 사용하여 이 질문에 답했습니다. 전체적으로 LLM이 생성한 기술이 No-Skill 프롬프팅 방식보다 신뢰성 있는 성능 향상을 가져오지 않는다는 것을 발견했습니다. 이후, 기술의 어느 부분이 유용한지 확인하기 위해 다양한 구성 요소를 제거하는 실험을 진행했습니다. 주요 실험은 56개의 작업, 9가지 모델 구성 및 3개 제공업체를 포함하며 총 7,560회의 실행 결과를 얻었습니다. 결과적으로, 단일 프롬프트만 사용하는 것과 비교했을 때 전체 LLM 생성 기술도, 부분적으로 제거된 기술 변형도 성능을 크게 향상시키지 못했습니다. 모든 p-값은 최소 0.396이었으며, 변형 간의 총 성능 차이는 1.2%에 불과했습니다. 추가적인 토큰 매칭 제어 실험 (1,512회 실행) 결과, 전체 기술은 작업과 관련 없는 기술 형식 콘텐츠와 유사한 성능을 보였습니다. 이러한 결과는 데이터 과학 워크플로우에서 LLM이 생성한 단일 기술을 기본 단일샷 프롬프팅 전략으로 사용하는 것에 대한 주의를 환기시킵니다.
Product data scientists often ask LLM-based agents to help with recurring execution tasks such as cleaning data, writing SQL, choosing statistical tests, and formatting results. Reusable skill files are meant to avoid prompting from scratch by packaging guidance for a task family. Expert-written skills can encode high-quality guidance, but writing and maintaining them across many data-science task families creates a manual bottleneck. We ask whether LLM-generated skills offer a useful low-curation alternative: do they improve performance over the task prompt alone? We test this question across four lifecycle stages: data preparation, data extraction, statistical analysis, and reporting, using one generated skill per stage. We find no reliable improvement from full generated skills over No-Skill prompting. We then ask whether any part of the skill is useful by ablating different skill components. The main ablation covers 56 tasks, nine model configurations, and three providers, yielding 7,560 runs. Compared with prompting using the task alone, neither the full generated skill nor any ablated skill variant significantly improves performance; all p-values are at least 0.396, and the total spread across variants is only 1.2 pp. A supplemental token-matched control adds 1,512 runs and finds that Full skills perform similarly to task-irrelevant skill-formatted content. The results caution against using one LLM-generated skill per data-science workflow as a default single-shot prompting strategy.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.