요약
- faster-whisper와 openai-whisper 모두
transcribe()에 파일 경로를 넘길 때만 16kHz로 리샘플한다. numpy 배열을 넘기면 샘플레이트를 묻지 않고 16kHz로 읽고, 경고도 띄우지 않는다. - 48kHz 배열을 그대로 넣으면 실제 1초(48,000샘플)가 3초로 읽힌다. 소리가 3배 느려지고 음정도 낮아져서 전사가 망가진다. 3배 빨라지는 게 아니다.
- 배열을 직접 넘길 때는 호출하는 쪽이 먼저 16kHz 모노 float32로 바꿔야 한다.
본문
faster-whisper
WhisperModel.transcribe(audio: Union[str, BinaryIO, np.ndarray], ...)transcribe.py:875-876:배열이면 이 분기를 건너뛴다. 디코딩과 리샘플을 모두 하지 않는다.if not isinstance(audio, np.ndarray): audio = decode_audio(audio, sampling_rate=sampling_rate)- 리샘플은
decode_audio가 PyAVAudioResampler(format="s16", layout="mono", rate=sampling_rate)로 한다(audio.py:37). sampling_rate는FeatureExtractor의 기본값 16000으로 정해진다(feature_extractor.py:8).transcribe()에는 입력 샘플레이트를 받는 인자가 없다.- 길이 계산도 16kHz로 한다.
duration = audio.shape[0] / sampling_rate(transcribe.py:878)라서 48kHz 배열이면info.duration이 실제의 3배로 나온다.info.duration을 실제 오디오 길이와 비교하면 바로 알아챌 수 있다.
openai-whisper
transcribe(model, audio: Union[str, np.ndarray, torch.Tensor], ...)log_mel_spectrogram은 입력이str일 때만load_audio를 부른다(audio.py:138-140).load_audio는 ffmpeg에-ar 16000을 줘서 리샘플한다(audio.py:53).- docstring에도 배열은 16kHz라고 가정한다고 적혀 있다(
audio.py:122): "a NumPy array or Tensor containing the audio waveform in 16 kHz".
샘플레이트가 다를 때 일어나는 일
- 16kHz로 읽으면 재생 속도가
16000 / 원래 샘플레이트배가 된다.- 48kHz: 1/3 속도로 느려지고 음정이 낮아진다.
- 8kHz(전화 음성): 2배 빨라지고 음정이 높아진다.
- 에러가 나지 않고 결과도 그럴듯하게 나온다. 누락이나 이상한 문장이 생겨도 모델이나 오디오 품질을 먼저 의심하게 된다.
배열을 넘겨야 할 때 16kHz로 바꾸는 방법
- 파일에서 읽는다면 라이브러리 함수로 16kHz 배열을 만든다.
- faster-whisper:
from faster_whisper.audio import decode_audio후decode_audio(path, sampling_rate=16000) - openai-whisper:
whisper.load_audio(path)
- faster-whisper:
- 이미 48kHz 배열이 있다면 직접 리샘플한다. 예:
scipy.signal.resample_poly(x, up=1, down=3)또는librosa.resample(x, orig_sr=48000, target_sr=16000) - 파일 경로를 그대로 넘기는 코드는 라이브러리가 리샘플하므로 이 문제가 없다.
검증 범위
- 소스로 확인: 위 분기와 기본값. faster-whisper
ed9a06c, openai-whisper8609812기준이다. - 계산으로 도출(실행해 보지는 않음): 1/3 속도,
info.duration3배.
관련 노트
- 음성 AI 학습 입력은 인식·분석은 16kHz, 합성·대화는 24kHz 모노로 수렴한다
- Whisper 음성 처리와 최적화 방식
- faster-whisper BatchedInferencePipeline은 production blocker급 미해결 이슈가 다수 있다
참고
- https://github.com/SYSTRAN/faster-whisper/blob/ed9a06cd89a93e47838f564998a6c09b655d7f43/faster_whisper/transcribe.py#L875-L878
- https://github.com/SYSTRAN/faster-whisper/blob/ed9a06cd89a93e47838f564998a6c09b655d7f43/faster_whisper/audio.py#L19-L41
- https://github.com/SYSTRAN/faster-whisper/blob/ed9a06cd89a93e47838f564998a6c09b655d7f43/faster_whisper/feature_extractor.py#L8
- https://github.com/openai/whisper/blob/86098128c0b4f24f0e2aa2994de830614b474227/whisper/audio.py#L110-L140
- https://github.com/openai/whisper/blob/86098128c0b4f24f0e2aa2994de830614b474227/whisper/audio.py#L25-L53