2607.05364v1 Jul 06, 2026 cs.CL

REDDIT: ASR 모델 생성 시 발생하는 타임스탬프 드리프트를 수정하면서 망각을 방지하는 재생 기반 분포 편집 방법

REDDIT: Correcting Model-Generated Timestamp Drift in ASR without Forgetting via Replay-Based Distribution Editing

Hung-yi Lee
Hung-yi Lee
Citations: 409
h-index: 11
Ke-Han Lu
Ke-Han Lu
Citations: 518
h-index: 13
Cheng-Kang Chou
Cheng-Kang Chou
Citations: 6
h-index: 2
Ming-To Chuang
Ming-To Chuang
Citations: 8
h-index: 2
Chan-Jan Hsu
Chan-Jan Hsu
Citations: 186
h-index: 8

최신 오토리거시브 ASR 시스템은 디코딩된 토큰 형태로 타임스탬프를 출력할 수 있어, 프레임 레벨 정렬기나 추론 시간 후처리 없이도 타임스탬프 기반 음성 기록 생성이 가능합니다. 본 연구에서는 이러한 생성된 타임스탬프가 긴 비음성 구간에서 드리프트되는 현상을 확인했습니다. 즉, 음성 기록 자체는 의미 있는 내용일 수 있지만, 디코딩된 시간 축이 실제 오디오와 동떨어지는 경향이 있습니다. 저희는 자체적으로 구축한 갭(gap) 및 장거리 갭(long-gap) 벤치마크를 사용하여 15개의 타임스탬프 생성 ASR 및 음성-언어 시스템에서 이러한 비음성 유발 타임스탬프 드리프트를 분석했습니다. 단순하게 타임스탬프를 보정하여 파인튜닝하는 방식은 정렬 성능을 향상시킬 수 있지만, 목표가 아닌 ASR 작업의 성능을 심각하게 저하시키는 망각 문제를 야기할 수 있습니다. 본 연구에서는 REDDIT(REplay-based Distribution eDITing)이라는 가벼운 2단계 후속 학습 프레임워크를 제안합니다. 이 프레임워크는 타임스탬프 드리프트를 수정하면서 이러한 파괴적인 망각을 방지합니다. 먼저, 모델 자체의 재생된 디코더 컨텍스트 하에서 타임스탬프 목표 값을 편집하고, 동시에 타임스탬프 토큰이 아닌 부분에 대해서는 기존 분포를 유지하며, 마지막으로 짧은 편집된 접두사 개선 단계를 적용합니다. 이 프레임워크 내에서 저희는 인간의 음성 기록이나 타임스탬프 주석 없이 VAD(Voice Activity Detection)로 잘라낸 음성 구간과 삽입된 비음성 갭, 그리고 알려진 연결 오프셋을 결합하여 수정에 필요한 감독 신호를 생성합니다. Whisper-tiny 모델에 대해 34.9시간의 목표 수정 오디오를 사용하고 전체 모델 파라미터의 1.6%만 업데이트했습니다. 그 결과, 장거리 갭 mIoU(mean Intersection over Union)가 38.7%에서 95.0%로 향상되었고, 다양한 갭이 혼합된 환경에서의 오프도메인 AAS(Average Audio Similarity Score)가 2752ms에서 223ms로 감소했습니다. 동시에 CV-en MER(Mean Error Rate)은 41.3%로 유지되었습니다 (일반적인 SFT(Supervised Fine-Tuning) 디코더 튜닝의 경우 524.2%).

Original Abstract

Modern autoregressive ASR systems can emit timestamps as decoded tokens, enabling timestamped transcription without frame-level aligners or inference-time post-processing. We show that these generated timestamps can drift across long non-speech spans: the transcript may remain plausible, but the decoded time axis drifts away from the audio. We study this non-speech-induced timestamp drift with self-built gap and long-gap benchmarks across 15 evaluated timestamp-producing ASR and audio-language systems. Naive timestamp-corrected fine-tuning improves alignment but can severely degrade non-target ASR behavior, exposing a forgetting problem. We propose REDDIT(REplay-based Distribution eDITing), a lightweight two-stage post-training framework that corrects timestamps while avoiding this catastrophic forgetting: it first edits timestamp targets under the model's own replayed decoder context while matching the frozen base distribution on non-timestamp tokens, then applies a short edited-prefix refinement stage. In this framework, we construct correction supervision without human transcripts or human timestamp annotations by combining VAD-trimmed speech spans with inserted non-speech gaps and known concatenation offsets. On Whisper-tiny, 34.9 hours of targeted correction audio used and only 1.6% of model parameters updated, raising long-gap mIoU from 38.7% to 95.0% and reducing mixed-gap out-of-domain AAS from 2752 ms to 223 ms while preserving CV-en MER at 41.3% (versus 524.2% for ordinary SFT decoder tuning).

0 Citations
0 Influential
6.5 Altmetric
32.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!