Zettelkasten
Search
⌘K
Graph
Tags
#
ctranslate2
4개
Flash Attention은 어텐션을 칩 안의 메모리에서 블록 단위로 계산해 큰 메모리 왕복을 없앤다
1,504자
flash-attention
gpu-memory-hierarchy
online-softmax
tiling
ctranslate2
faster-whisper BatchedInferencePipeline은 production blocker급 미해결 이슈가 다수 있다
3,334자
faster-whisper
ctranslate2
batched-inference
production-readiness
library-evaluation
ctranslate2 num_workers는 가중치를 공유하고 worker별 메모리는 호출 중에만 늘어난다
2,330자
ctranslate2
faster-whisper
memory-management
inference-parallelism
benchmarking
ctranslate2의 GIL 해제와 num_workers는 독립적인 두 메커니즘이다
2,428자
ctranslate2
faster-whisper
gil
python-concurrency
inference-parallelism
threading