2607.21482v1 Jul 23, 2026 cs.AI

Agentic coding without the cloud: evaluating open-weight large language models on longitudinal data preparation tasks

Mack Nixon
Mack Nixon
Citations: 0
h-index: 0
Liam Wright
Liam Wright
Citations: 33
h-index: 4
Yevgeniya Kovalchuk
Yevgeniya Kovalchuk
Citations: 48
h-index: 3
Alison Fang-Wei Wu
Alison Fang-Wei Wu
Citations: 11
h-index: 1
Martin N. Danka
Martin N. Danka
Citations: 5
h-index: 1
Andy Boyd
Andy Boyd
Citations: 1
h-index: 1
D. Bann
D. Bann
Citations: 3,445
h-index: 31

Large language models (LLMs) and agents are now widely used tools in code development, with data typically sent to third-party cloud-based models. Their adoption in research using personal data is constrained by governance requirements that typically prohibit data transmission to external services. Locally deployable open-weight models offer an alternative since sensitive data never leave the local environment. We introduce an open-source framework for evaluating the efficacy of AI agents powered by open-weight LLMs on one of the most persistent bottlenecks in research on longitudinal population studies: data preparation. The framework comprises: a curated ground-truth dataset (cleaning scripts preparing six sweeps of data from a British cohort study), task definitions encompassing tasks such as category harmonization and multi-wave merging, and automated routines for evaluating the LLM-produced R code and outputted data. We benchmark LLMs across the (consumer grade) deployment spectrum to assess their efficacy in 20 data preparation tasks (creation of 102 variables). Current state-of-the-art, 31-35B parameter models almost saturated our benchmark ("average task completion" up to 87.9%). The performance of open-weight LLMs running on consumer-grade hardware shows promise of a viable path toward AI-assisted data preparation in governance-restricted research settings. Our framework is publicly available at: https://github.com/UCL-ARC/RRBench.

0 Citations
0 Influential
40.993061443341 Altmetric
205.0 Score
Original PDF
2

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!