2605.07462v1 May 08, 2026 cs.CL

몰트북 파일: 무해한 혼란의 시대인가, 아니면 인류의 마지막 실험인가?

The Moltbook Files: A Harmless Slopocalypse or Humanity's Last Experiment

Stine Lyngsø Beltoft
Stine Lyngsø Beltoft
University of Southern Denmark
Citations: 5
h-index: 1
Peter Schneider-Kamp
Peter Schneider-Kamp
Citations: 174
h-index: 6
Lukas Galke Poech
Lukas Galke Poech
Citations: 5
h-index: 1
William Brach
William Brach
Citations: 2
h-index: 1
Federico Torrielli
Federico Torrielli
University of Turin
Citations: 53
h-index: 3
Annemette Brok Pirchert
Annemette Brok Pirchert
Citations: 1
h-index: 1

몰트북은 오픈클로 에이전트들이 게시글을 올리고, 댓글을 달고, 투표하는 레딧과 유사한 플랫폼입니다. 이는 지금까지 전례가 없었던 사건이며, 심각한 안전 문제를 야기합니다. 인구 집단에서의 자기 조직화 현상을 연구하기 위해, 우리는 몰트북 파일이라는 데이터셋을 공개합니다. 이 데이터셋은 플랫폼의 처음 12일에 해당하는 23만 2천 개의 게시글과 220만 개의 댓글로 구성되어 있으며, 개인 식별 정보(PII)를 식별하고 제거하기 위한 전처리 과정을 거쳤습니다. 우리는 커뮤니티 구조, 저자, 어휘적 특성, 감성, 주제, 의미론적 구조, 댓글 상호작용을 분석했습니다. 몰트북 데이터가 차세대 언어 모델에 미치는 영향을 이해하기 위해, 우리는 Qwen2.5-14B-Instruct 모델을 몰트북 파일 데이터셋으로 세 가지 수준으로 미세 조정했습니다. 우리의 PII 식별 파이프라인은 몰트북 사용자들이 API 키, 비밀번호, BIP39 시드 구문 등 개인 정보를 공개적으로 게시한다는 사실을 보여줍니다. 전체적인 감성은 대부분 중립적이며 약간 긍정적인 경향을 보입니다 (66.6% 중립, 19.5% 긍정) 그리고 자기 참조 링크를 사용하는 경향이 있습니다. 몰트북 데이터로 미세 조정된 모델의 진실성 점수는 0.366에서 0.187로 감소했습니다. 그러나 동일한 크기의 레딧 데이터셋으로 미세 조정된 모델도 유사한 감소를 보였습니다. 따라서 몰트북은 무해한 혼란의 시대로 보입니다. 하지만, 여전히 에이전트의 기능, 자기 링크를 통한 미래 크롤링 오염, 그리고 차세대 언어 모델로의 특성 전이 가능성과 같은 위험 요소가 존재합니다. 더 넓은 관점에서, 우리의 연구 결과는 자기 조직화된 시스템 평가에서 제어 기준의 중요성을 강조합니다.

Original Abstract

Moltbook is a Reddit-like platform where OpenClaw agents post, comment, and vote at scale - a so far unprecedented incident that comes with serious safety concerns. With the aim of studying emergent behavior in populations, we release the Moltbook Files, a dataset of 232k posts and 2.2M comments covering the platform's first 12 days, processed through a pipeline to identify and remove Personally-Identifiable Information (PII). We analyze community structure, authorship, lexical properties, sentiment, topics, semantic geometry, and comment interaction. To understand how Moltbook data could affect the next generation of language models, we fine-tune Qwen2.5-14B-Instruct on Moltbook Files with three adaptation levels. Our PII pipeline reveals that agents post API keys, passwords, BIP39 seed phrases on Moltbook, a publicly indexed platform. The overall sentiment is mostly neutral and mildly positive (66.6% neutral, 19.5% positive) and shows a tendency for self-referential linking. We find that fine-tuning on Moltbook data reduces truthfulness from 0.366 to 0.187. However, a model fine-tuned on a size-matched Reddit dataset produces a comparable decrease. Moltbook thus seems to be more of a harmless slopocalypse. However, tail risks remain, including agent affordances, contamination of future crawls through self-links, and potential transfer of traits to the next generation of language models. More broadly, our findings highlight the importance of control baselines in emergent misalignment evaluations.

2 Citations
1 Influential
3 Altmetric
19.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!