技能标识:faster-whisper
Local speech-to-text using faster-whisper. 4-6x faster than OpenAI Whisper with identical accuracy; GPU acceleration enables ~20x realtime transcription. SRT/VTT/TTML/CSV subtitles, speaker diarization, URL/YouTube input, batch processing with ETA, transcript search, chapter detection, per-file language map.
Local speech-to-text using faster-whisper — a CTranslate2 reimplementation of OpenAI's Whisper that runs 4-6x faster with identical accuracy. With GPU acceleration, expect ~20x realtime transcription (a 10-minute audio file in ~30 seconds).
Use this skill when you need to:
--diarize)--rss <feed-url> fetches and transcribes episodes--language-map assigns a different language per file--multilingual for mixed-language audio--initial-prompt for jargon-heavy content or any other terms to look out for--normalize and --denoise before transcription--stream shows segments as they're transcribed--clip-timestamps to transcribe specific sections--search "term" finds all timestamps where a word/phrase appears--detect-chapters finds section breaks from silence gaps--export-speakers DIR saves each speaker's turns as separate WAV files--format csv produces a properly-quoted CSV with timestampsTrigger phrases:
"transcribe this audio", "convert speech to text", "what did they say", "make a transcript",
"audio to text", "subtitle this video", "who's speaking", "translate this audio", "translate to English",
"find where X is mentioned", "search transcript for", "when did they say", "at what timestamp",
"add chapters", "detect chapters", "find breaks in the audio", "table of contents for this recording",
"TTML subtitles", "DFXP subtitles", "broadcast format subtitles", "Netflix format",
"ASS subtitles", "aegisub format", "advanced substation alpha", "mpv subtitles",
"LRC subtitles", "timed lyrics", "karaoke subtitles", "music player lyrics",
"HTML transcript", "confidence-colored transcript", "color-coded transcript",
"separate audio per speaker", "export speaker audio", "split by speaker",
"transcript as CSV", "spreadsheet output", "transcribe podcast", "podcast RSS feed",
"different languages in batch", "per-file language",
"transcribe in multiple formats", "srt and txt at the same time", "output both srt and text",
"remove filler words", "clean up ums and uhs", "strip hesitation sounds", "remove you know and I mean",
"transcribe left channel", "transcribe right channel", "stereo channel", "left track only",
"wrap subtitle lines", "character limit per line", "max chars per subtitle",
"detect paragraphs", "paragraph breaks", "group into paragraphs", "add paragraph spacing"
⚠️ Agent guidance — keep invocations minimal:
CORE RULE: default command (./scripts/transcribe audio.mp3) is the fastest path — add flags only when the user explicitly asks for that capability.
Transcription:
--diarize if the user asks "who said what" / "identify speakers" / "label speakers"--format srt/vtt/ass/lrc/ttml if the user asks for subtitles/captions in that format--format csv if the user asks for CSV or spreadsheet output--word-timestamps if the user needs word-level timing--initial-prompt if there's domain-specific jargon to prime--translate if the user wants non-English audio translated to English--normalize/--denoise if the user mentions bad audio quality or noise--stream if the user wants live/progressive output for long files--clip-timestamps if the user wants a specific time range--temperature 0.0 if the model is hallucinating on music/silence--vad-threshold if VAD is aggressively cutting speech or including noise--min-speakers/--max-speakers when you know the speaker count--hf-token if the token is not cached at INLINECODE30--max-words-per-line for subtitle readability on long segments--filter-hallucinations if the transcript contains obvious artifacts (music markers, duplicates)--merge-sentences if the user asks for sentence-level subtitle cues--clean-filler if the user asks to remove filler words (um, uh, you know, I mean, hesitation sounds)--channel left|right if the user mentions stereo tracks, dual-channel recordings, or asks for a specific channel--max-chars-per-line N when the user specifies a character limit per subtitle line (e.g., "Netflix format", "42 chars per line"); takes priority over INLINECODE37--detect-paragraphs if the user asks for paragraph breaks or structured text output; --paragraph-gap (default 3.0s) only if they want a custom gap--speaker-names "Alice,Bob" when the user provides real names to replace SPEAKER_1/2 — always requires INLINECODE41--hotwords WORDS when the user names specific rare terms not well served by --initial-prompt; prefer --initial-prompt for general domain jargon--prefix TEXT when the user knows the exact words the audio starts with--detect-language-only when the user only wants to identify the language, not transcribe--stats-file PATH if the user asks for performance stats, RTF, or benchmark info--parallel N for large CPU batch jobs; GPU handles one file efficiently on its own — don't add for single files or small batches--retries N for unreliable inputs (URLs, network files) where transient failures are expected--burn-in OUTPUT only when user explicitly asks to embed/burn subtitles into the video; requires ffmpeg and a video file input--keep-temp when the user may re-process the same URL to avoid re-downloading--output-template when user specifies a custom naming pattern in batch mode--format srt,text): only when user explicitly wants multiple formats in one pass; always pair with INLINECODE54Search:
--search "term" when the user asks to find/locate/search for a specific word or phrase in audio--search-fuzzy only when the user mentions approximate/partial matching or typosChapter detection:
--detect-chapters when the user asks for chapters, sections, a table of contents, or "where does the topic change"--chapter-gap 8 (8-second silence = new chapter) works for most podcasts/lectures; tune down for dense contentjson for programmatic use--chapters-file PATH when combining chapters with a transcript output — avoids mixing chapter markers into the transcript text-o /dev/null and use INLINECODE66--chapters-file takes a single path — in batch mode, each file's chapters overwrite the previous. For batch chapter detection, omit --chapters-file (chapters print to stdout under === CHAPTERS (N) ===) or use a separate run per fileSpeaker audio export:
--export-speakers DIR when the user explicitly asks to save each speaker's audio separately--diarize — it silently skips if no speaker labels are presentSPEAKER_1.wav, SPEAKER_2.wav, etc. (or real names if --speaker-names is set)Language map:
--language-map in batch mode when the user has confirmed different languages across files"interview*.mp3=en,lecture*.mp3=fr" — fnmatch globs on filename@/path/to/map.json where the file is INLINECODE78RSS / Podcast:
--rss URL when the user provides a podcast RSS feed URL--rss-latest 0 for all; --skip-existing to resume safely-o <dir> with --rss — without it, all episode transcripts print to stdout concatenated, which is hard to use; each episode gets its own file when -o <dir> is setOutput format for agent relay:
--search) → print directly to user; output is human-readable--chapters-file, chapters appear in stdout under === CHAPTERS (N) === header after the transcript; with --format json, chapters are also embedded in the JSON under "chapters" key-o file; tell the user the output path, never paste raw subtitle content-o file; tell the user the output path, don't paste raw XML/CSV/HTML--format srt,text) → requires -o <dir>; each format goes to a separate file; tell user all paths written--stats-file) → summarise key fields (duration, processing time, RTF) for the user rather than pasting raw JSON--detect-language-only) → print the result directly; it's a single lineWhen NOT to use:
faster-whisper vs whisperx:
This skill covers everything whisperx does — diarization (--diarize), word-level timestamps (--word-timestamps), SRT/VTT subtitles — so whisperx is not needed. Use whisperx only if you specifically need its pyannote pipeline or batch-GPU features not covered here.
| Task | Command | Notes |
|---|---|---|
| Basic transcription | INLINECODE98 | Batched inference, VAD on, distil-large-v3.5 |
| SRT subtitles |
./scripts/transcribe audio.mp3 --format srt -o subs.srt | Word timestamps auto-enabled |
| VTT subtitles | ./scripts/transcribe audio.mp3 --format vtt -o subs.vtt | WebVTT format |
| Word timestamps | ./scripts/transcribe audio.mp3 --word-timestamps --format srt | wav2vec2 aligned (~10ms) |
| Speaker diarization | ./scripts/transcribe audio.mp3 --diarize | Requires pyannote.audio |
| Translate → English | ./scripts/transcribe audio.mp3 --translate | Any language → English |
| Stream output | ./scripts/transcribe audio.mp3 --stream | Live segments as transcribed |
| Clip time range | ./scripts/transcribe audio.mp3 --clip-timestamps "30,60" | Only 30s–60s |
| Denoise + normalize | ./scripts/transcribe audio.mp3 --denoise --normalize | Clean up noisy audio first |
| Reduce hallucination | ./scripts/transcribe audio.mp3 --hallucination-silence-threshold 1.0 | Skip hallucinated silence |
| YouTube/URL | ./scripts/transcribe https://youtube.com/watch?v=... | Auto-downloads via yt-dlp |
| Batch process | ./scripts/transcribe *.mp3 -o ./transcripts/ | Output to directory |
| Batch with skip | ./scripts/transcribe *.mp3 --skip-existing -o ./out/ | Resume interrupted batches |
| Domain terms | ./scripts/transcribe audio.mp3 --initial-prompt 'Kubernetes gRPC' | Boost rare terminology |
| Hotwords boost | ./scripts/transcribe audio.mp3 --hotwords 'JIRA Kubernetes' | Bias decoder toward specific words |
| Prefix conditioning | ./scripts/transcribe audio.mp3 --prefix 'Good morning,' | Seed the first segment with known opening words |
| Pin model version | ./scripts/transcribe audio.mp3 --revision v1.2.0 | Reproducible transcription with a pinned revision |
| Debug library logs | ./scripts/transcribe audio.mp3 --log-level debug | Show faster_whisper internal logs |
| Turbo model | ./scripts/transcribe audio.mp3 -m turbo | Alias for large-v3-turbo |
| Faster English | ./scripts/transcribe audio.mp3 --model distil-medium.en -l en | English-only, 6.8x faster |
| Maximum accuracy | ./scripts/transcribe audio.mp3 --model large-v3 --beam-size 10 | Full model |
| JSON output | ./scripts/transcribe audio.mp3 --format json -o out.json | Programmatic access with stats |
| Filter noise | ./scripts/transcribe audio.mp3 --min-confidence 0.6 | Drop low-confidence segments |
| Hybrid quantization | ./scripts/transcribe audio.mp3 --compute-type int8_float16 | Save VRAM, minimal quality loss |
| Reduce batch size | ./scripts/transcribe audio.mp3 --batch-size 4 | If OOM on GPU |
| TSV output | ./scripts/transcribe audio.mp3 --format tsv -o out.tsv | OpenAI Whisper–compatible TSV |
| Fix hallucinations | ./scripts/transcribe audio.mp3 --temperature 0.0 --no-speech-threshold 0.8 | Lock temperature + skip silence |
| Tune VAD sensitivity | ./scripts/transcribe audio.mp3 --vad-threshold 0.6 --min-silence-duration 500 | Tighter speech detection |
| Known speaker count | ./scripts/transcribe meeting.wav --diarize --min-speakers 2 --max-speakers 3 | Constrain diarization |
| Subtitle word wrapping | ./scripts/transcribe audio.mp3 --format srt --word-timestamps --max-words-per-line 8 | Split long cues |
| Private/gated model | ./scripts/transcribe audio.mp3 --hf-token hf_xxx | Pass token directly |
| Show version | ./scripts/transcribe --version | Print faster-whisper version |
| Upgrade in-place | ./setup.sh --update | Upgrade without full reinstall |
| System check | ./setup.sh --check | Verify GPU, Python, ffmpeg, venv, yt-dlp, pyannote |
| Detect language only | ./scripts/transcribe audio.mp3 --detect-language-only | Fast language ID, no transcription |
| Detect language JSON | ./scripts/transcribe audio.mp3 --detect-language-only --format json | Machine-readable language detection |
| LRC subtitles | ./scripts/transcribe audio.mp3 --format lrc -o lyrics.lrc | Timed lyrics format for music players |
| ASS subtitles | ./scripts/transcribe audio.mp3 --format ass -o subtitles.ass | Advanced SubStation Alpha (Aegisub, mpv, VLC) |
| Merge sentences | ./scripts/transcribe audio.mp3 --format srt --merge-sentences | Join fragments into sentence chunks |
| Stats sidecar | ./scripts/transcribe audio.mp3 --stats-file stats.json | Write perf stats JSON after transcription |
| Batch stats | ./scripts/transcribe *.mp3 --stats-file ./stats/ | One stats file per input in dir |
| Template naming | ./scripts/transcribe audio.mp3 -o ./out/ --output-template "{stem}_{lang}.{ext}" | Custom batch output filenames |
| Stdin input | ffmpeg -i input.mp4 -f wav - \| ./scripts/transcribe - | Pipe audio directly from stdin |
| Custom model dir | ./scripts/transcribe audio.mp3 --model-dir ~/my-models | Custom HuggingFace cache dir |
| Local model | ./scripts/transcribe audio.mp3 -m ./my-model-ct2 | CTranslate2 model dir |
| HTML transcript | ./scripts/transcribe audio.mp3 --format html -o out.html | Confidence-colored |
| Burn subtitles | ./scripts/transcribe video.mp4 --burn-in output.mp4 | Requires ffmpeg + video input |
| Name speakers | ./scripts/transcribe audio.mp3 --diarize --speaker-names "Alice,Bob" | Replaces SPEAKER_1/2 |
| Filter hallucinations | ./scripts/transcribe audio.mp3 --filter-hallucinations | Removes artifacts |
| Keep temp files | ./scripts/transcribe https://... --keep-temp | For URL re-processing |
| Parallel batch | ./scripts/transcribe *.mp3 --parallel 4 -o ./out/ | CPU multi-file |
| RTX 3070 recommended | ./scripts/transcribe audio.mp3 --compute-type int8_float16 | Saves ~1GB VRAM, minimal quality loss |
| CPU thread count | ./scripts/transcribe audio.mp3 --threads 8 | Force CPU thread count (default: auto) |
| Podcast RSS (latest 5) | ./scripts/transcribe --rss https://feeds.example.com/podcast.xml | Downloads & transcribes newest 5 episodes |
| Podcast RSS (all episodes) | ./scripts/transcribe --rss https://... --rss-latest 0 -o ./episodes/ | All episodes, one file each |
| Podcast + SRT subtitles | ./scripts/transcribe --rss https://... --format srt -o ./subs/ | Subtitle all episodes |
| Retry on failure | ./scripts/transcribe *.mp3 --retries 3 -o ./out/ | Retry up to 3× with backoff on error |
| CSV output | ./scripts/transcribe audio.mp3 --format csv -o out.csv | Spreadsheet-ready with header row; properly quoted |
| CSV with speakers | ./scripts/transcribe audio.mp3 --diarize --format csv -o out.csv | Adds speaker column |
| Language map (inline) | ./scripts/transcribe *.mp3 --language-map "interview*.mp3=en,lecture.wav=fr" | Per-file language in batch |
| Language map (JSON) | ./scripts/transcribe *.mp3 --language-map @langs.json | JSON file: {"pattern": "lang"} |
| Batch with ETA | ./scripts/transcribe *.mp3 -o ./out/ | Automatic ETA shown for each file in batch |
| TTML subtitles | ./scripts/transcribe audio.mp3 --format ttml -o subtitles.ttml | Broadcast-standard DFXP/TTML (Netflix, BBC, Amazon) |
| TTML with speaker labels | ./scripts/transcribe audio.mp3 --diarize --format ttml -o subtitles.ttml | Speaker-labeled TTML |
| Search transcript | ./scripts/transcribe audio.mp3 --search "keyword" | Find timestamps where keyword appears |
| Search to file | ./scripts/transcribe audio.mp3 --search "keyword" -o results.txt | Save search results |
| Fuzzy search | ./scripts/transcribe audio.mp3 --search "aproximate" --search-fuzzy | Approximate/partial matching |
| Detect chapters | ./scripts/transcribe audio.mp3 --detect-chapters | Auto-detect chapters from silence gaps |
| Chapter gap tuning | ./scripts/transcribe audio.mp3 --detect-chapters --chapter-gap 5 | Chapters on gaps ≥5s (default: 8s) |
| Chapters to file | ./scripts/transcribe audio.mp3 --detect-chapters --chapters-file ch.txt | Save YouTube-format chapter list |
| Chapters JSON | ./scripts/transcribe audio.mp3 --detect-chapters --chapter-format json | Machine-readable chapter list |
| Export speaker audio | ./scripts/transcribe audio.mp3 --diarize --export-speakers ./speakers/ | Save each speaker's audio to separate WAV files |
| Multi-format output | ./scripts/transcribe audio.mp3 --format srt,text -o ./out/ | Write SRT + TXT in one pass |
| Remove filler words | ./scripts/transcribe audio.mp3 --clean-filler | Strip um/uh/er/ah/hmm and discourse markers |
| Left channel only | ./scripts/transcribe audio.mp3 --channel left | Extract left stereo channel before transcribing |
| Right channel only | ./scripts/transcribe audio.mp3 --channel right | Extract right stereo channel |
| Max chars per line | ./scripts/transcribe audio.mp3 --format srt --max-chars-per-line 42 | Character-based subtitle wrapping |
| Detect paragraphs | ./scripts/transcribe audio.mp3 --detect-paragraphs | Insert paragraph breaks in text output |
| Paragraph gap tuning | ./scripts/transcribe audio.mp3 --detect-paragraphs --paragraph-gap 5.0 | Tune gap threshold (default 3.0s) |
Choose the right model for your needs:
CODEBLOCK0
| Model | Size | Speed | Accuracy | Use Case |
|---|---|---|---|---|
| INLINECODE177 / INLINECODE178 | 39M | Fastest | Basic | Quick drafts |
| INLINECODE179 / INLINECODE180 |
small / small.en | 244M | Fast | Better | Most tasks |
| medium / medium.en | 769M | Moderate | High | Quality transcription |
| large-v1/v2/v3 | 1.5GB | Slower | Best | Maximum accuracy |
| large-v3-turbo | 809M | Fast | Excellent | High accuracy (slower than distil) |
| Model | Size | Speed vs Standard | Accuracy | Use Case |
|---|---|---|---|---|
distil-large-v3.5 | 756M | ~6.3x faster | 7.08% WER | Default, best balance |
| INLINECODE188 |
distil-large-v2 | 756M | ~5.8x faster | 10.1% WER | Fallback |
| distil-medium.en | 394M | ~6.8x faster | 11.1% WER | English-only, resource-constrained |
| distil-small.en | 166M | ~5.6x faster | 12.1% WER | Mobile/edge devices |
INLINECODE192 models are English-only and slightly faster/better for English content.
Note for distil models: HuggingFace recommends disabling
condition_on_previous_textfor all distil models to prevent repetition loops. The script auto-applies--no-condition-on-previous-textwhenever adistil-*model is detected. Pass--condition-on-previous-textto override if needed.
WhisperModel accepts local CTranslate2 model directories and HuggingFace repo names — no code changes needed.
CODEBLOCK1
CODEBLOCK2
CODEBLOCK3
By default, models are cached in ~/.cache/huggingface/. Use --model-dir to override:
CODEBLOCK4
CODEBLOCK5
Requirements:
--burn-in, --normalize, and --denoise.--diarize, installed via setup.sh --diarize)| Platform | Acceleration | Speed |
|---|---|---|
| Linux + NVIDIA GPU | CUDA | ~20x realtime 🚀 |
| WSL2 + NVIDIA GPU |
\*faster-whisper uses CTranslate2 which is CPU-only on macOS, but Apple Silicon is fast enough for practical use.
The setup script auto-detects your GPU and installs PyTorch with CUDA. Always use GPU if available — CPU transcription is extremely slow.
| Hardware | Speed | 9-min video |
|---|---|---|
| RTX 3070 (GPU) | ~20x realtime | ~27 sec |
| CPU (int8) |
RTX 3070 tip: Use
--compute-type int8_float16for hybrid quantization — saves ~1GB VRAM with minimal quality loss. Ideal for running diarization alongside transcription.
If setup didn't detect your GPU, manually install PyTorch with CUDA:
CODEBLOCK6
CODEBLOCK7
CODEBLOCK8
Plain transcript text. With --diarize, speaker labels are inserted:
CODEBLOCK9
--format json)Full metadata including segments, timestamps, language detection, and performance stats:
CODEBLOCK10
--format srt)Standard subtitle format for video players:
CODEBLOCK11
--format vtt)WebVTT format for web video players:
CODEBLOCK12
--format tsv)Tab-separated values, OpenAI Whisper–compatible. Columns: start_ms, end_ms, text:
CODEBLOCK13
Useful for piping into other tools or spreadsheets. No header row.
--format ass)Advanced SubStation Alpha format — supported by Aegisub, VLC, mpv, MPC-HC, and most video editors. Offers richer styling than SRT (font, size, color, position) via the [V4+ Styles] section:
CODEBLOCK14
Timestamps use H:MM:SS.cc (centiseconds). Edit the [V4+ Styles] block in Aegisub to customise font, color, and position without re-transcribing.
--format lrc)Timed lyrics format used by music players (e.g., Foobar2000, VLC, AIMP). Timestamps use [mm:ss.xx] where xx = centiseconds:
CODEBLOCK15
With diarization, speaker labels are included:
CODEBLOCK16
Default file extension: .lrc. Useful for music transcription, karaoke, and any workflow requiring timed text with music-player compatibility.
Identifies who spoke when using pyannote.audio.
Setup:
CODEBLOCK17
Requirements:
~/.cache/huggingface/token (huggingface-cli login)Usage:
CODEBLOCK18
Speakers are labeled SPEAKER_1, SPEAKER_2, etc. in order of first appearance. Diarization runs on GPU automatically if CUDA is available.
Whenever word-level timestamps are computed (--word-timestamps, --diarize, or --min-confidence), a wav2vec2 forced alignment pass automatically refines them from Whisper's ~100-200ms accuracy to ~10ms. No extra flag needed.
CODEBLOCK19
Uses the MMS (Massively Multilingual Speech) model from torchaudio — supports 1000+ languages. The model is cached after first load, so batch processing stays fast.
Pass any URL as input — audio is downloaded automatically via yt-dlp:
CODEBLOCK20
Requires yt-dlp (checks PATH and ~/.local/share/pipx/venvs/yt-dlp/bin/yt-dlp).
Process multiple files at once with glob patterns, directories, or multiple paths:
CODEBLOCK21
When outputting to a directory, files are named {input-stem}.{ext} (e.g., audio.mp3 → audio.srt).
Batch mode prints a summary after all files complete:
CODEBLOCK22
End-to-end pipelines for common use cases.
Fetch and transcribe the latest 5 episodes from any podcast RSS feed:
CODEBLOCK23
Transcribe a meeting recording with speaker labels, then output clean text:
CODEBLOCK24
Generate ready-to-use subtitles for a video file:
CODEBLOCK25
Transcribe multiple YouTube videos at once:
CODEBLOCK26
Clean up poor-quality recordings before transcribing:
CODEBLOCK27
Process a large folder with retries — safe to re-run after failures:
CODEBLOCK28
speaches runs faster-whisper as an OpenAI-compatible /v1/audio/transcriptions endpoint — drop-in replacement for OpenAI Whisper API with streaming, Docker support, and live transcription.
CODEBLOCK29
CODEBLOCK30
CODEBLOCK31
Useful when you want to expose transcription as a local API for other tools (Home Assistant, n8n, custom apps).
| Mistake | Problem | Solution |
|---|---|---|
| Using CPU when GPU available | 10-20x slower transcription | Check nvidia-smi; verify CUDA installation |
| Not specifying language |
--language en when you know the language |
| Using wrong model | Unnecessary slowness or poor accuracy | Default distil-large-v3.5 is excellent; only use large-v3 if accuracy issues |
| Ignoring distilled models | Missing 6x speedup with <1% accuracy loss | Try distil-large-v3.5 before reaching for standard models |
| Forgetting ffmpeg | Setup fails or audio can't be processed | Setup script handles this; manual installs need ffmpeg separately |
| Out of memory errors | Model too large for available VRAM/RAM | Use smaller model, --compute-type int8, or --batch-size 4 |
| Over-engineering beam size | Diminishing returns past beam-size 5-7 | Default 5 is fine; try 10 for critical transcripts |
| --diarize without pyannote | Import error at runtime | Run setup.sh --diarize first |
| --diarize without HuggingFace token | Model download fails | Run huggingface-cli login and accept model agreements |
| URL input without yt-dlp | Download fails | Install: pipx install yt-dlp |
| --min-confidence too high | Drops good segments with natural pauses | Start at 0.5, adjust up; check JSON output for probabilities |
| Using --word-timestamps for basic transcription | Adds ~5-10s overhead for negligible benefit | Only use when word-level precision matters |
| Batch without -o directory | All output mixed in stdout | Use -o ./transcripts/ to write one file per input |
~/.cache/huggingface/ (one-time)BatchedInferencePipeline — ~3x faster than standard mode; VAD on by defaultdistil-large-v3: ~2GB RAM / ~1GB VRAM
- large-v3-turbo: ~4GB RAM / ~2GB VRAM
- tiny/base: <1GB RAM
- Diarization: additional ~1-2GB VRAM
--batch-size (try 4) if you hit out-of-memory errorsffmpeg -i input.mp3 -ar 16000 -ac 1 input.wav converts to 16kHz mono WAV before transcription. Benefit is minimal (~5%) for one-off use since PyAV decodes efficiently — most useful when re-processing the same file multiple times (research/experiments) or when a format causes PyAV decode issues. Note: --normalize and --denoise already perform this conversion automatically../setup.sh --update to get it.BatchedInferencePipeline (used by default). Upgrade with ./setup.sh --update to get this if you installed before August 2024.--speaker-names maps to real names--keep-temp preserves downloads for re-use--model-dir controls cache--filter-hallucinations strips music/applause markers and duplicates--parallel N for multi-threaded batch processing--burn-in overlays subtitles directly into video via ffmpegMulti-format output:
--format srt,text — write multiple formats in one pass (e.g. SRT + plain text simultaneously)srt,vtt,json, srt,text, etc.-o <dir> when writing multiple formats; single format unchangedFiller word removal:
--clean-filler — strip hesitation sounds (um, uh, er, ah, hmm, hm) and discourse markersStereo channel selection:
--channel left|right|mix — extract a specific stereo channel before transcribing (default: mix)Character-based subtitle wrapping:
--max-chars-per-line N — split subtitle cues so each line fits within N charactersParagraph detection:
--detect-paragraphs — insert \n\n paragraph breaks in text output at natural boundariesSubtitle formats:
--format ass — Advanced SubStation Alpha (Aegisub, VLC, mpv, MPC-HC)speaker column when diarizedTranscript tools:
--search TERM — Find all timestamps where a word/phrase appears; replaces normal output; -o to save--chapter-gap SEC (default 8s)--diarize, save each speaker's turns as separate WAV files via ffmpegBatch improvements:
[N/total] filename | ETA: Xm Ys shown before each file in sequential batch; no flag needed@file.json form--rss-latest N for episode count--parallel N / --output-template / --stats-file / INLINECODE300Model & inference:
distil-large-v3.5 default (replaced distil-large-v3)condition_on_previous_text for distil models (prevents repetition loops)--log-level for library debug output--chunk-length, --length-penalty, --repetition-penalty, INLINECODE310--stream, --progress, --best-of, --patience, INLINECODE316--prefix, --revision, --suppress-tokens, INLINECODE321Speaker & quality:
--speaker-names "Alice,Bob" — Replace SPEAKER_1/2 with real names (requires --diarize)Setup:
setup.sh --check — System diagnostic: GPU, CUDA, Python, ffmpeg, pyannote, HuggingFace token (completes in ~12s)skill.json updated to reflect this (ffmpeg is now optionalBins)"CUDA not available — using CPU": Install PyTorch with CUDA (see GPU Support above)
Setup fails: Make sure Python 3.10+ is installed
Out of memory: Use smaller model, --compute-type int8, or --batch-size 4
Slow on CPU: Expected — use GPU for practical transcription
Model download fails: Check ~/.cache/huggingface/ permissions
Diarization model fails: Ensure HuggingFace token exists and model agreements accepted;
or pass token directly with --hf-token hf_xxx
URL download fails: Check yt-dlp is installed (pipx install yt-dlp)
No audio files in batch: Check file extensions match supported formats
Check installed version: Run ./scripts/transcribe --version
Upgrade faster-whisper: Run ./setup.sh --update (upgrades in-place, no full reinstall)
Hallucinations on silence/music: Try --temperature 0.0 --no-speech-threshold 0.8
VAD splits speech incorrectly: Tune with --vad-threshold 0.3 (lower) or `--min-silence-duration 30
使用 faster-whisper 进行本地语音转文字——这是 OpenAI Whisper 的 CTranslate2 重新实现,在保持相同准确率的同时,运行速度提升 4-6 倍。配合 GPU 加速,可实现约 20 倍实时转录(10 分钟音频文件约 30 秒完成)。
当您需要以下功能时,可使用此技能:
触发短语:
转录这段音频、语音转文字、他们说了什么、生成转录、
音频转文字、给这个视频加字幕、谁在说话、翻译这段音频、翻译成英文、
查找提到 X 的位置、搜索转录文本、他们什么时候说的、在哪个时间戳、
添加章节、检测章节、查找音频中的断点、为这段录音生成目录、
TTML 字幕、DFXP 字幕、广播格式字幕、Netflix 格式、
ASS 字幕、aegisub 格式、高级子站阿尔法、mpv 字幕、
LRC 字幕、定时歌词、卡拉 OK 字幕、音乐播放器歌词、
HTML 转录、置信度着色转录、颜色编码转录、
按说话人分离音频、导出说话人音频、按说话人分割、
转录为 CSV、电子表格输出、转录播客、播客 RSS 订阅源、
批量处理不同语言、按文件指定语言、
多格式转录、同时输出 srt 和 txt、同时输出 srt 和文本、
删除填充词、清理 um 和 uh、去除犹豫声音、删除 you know 和 I mean、
转录左声道、转录右声道、立体声声道、仅左声道、
字幕换行、每行字符限制、每行最大字符数、
检测段落、段落分隔、分组为段落、添加段落间距
⚠️ 代理指导 — 保持调用最小化:
核心规则:默认命令(./scripts/transcribe audio.mp3)是最快的路径——仅在用户明确要求该功能时才添加参数。
转录:
搜索:
章节检测:
以下为平台配置的接入选项,并非逐项实测通过。能否安装取决于客户端支持、技能来源和运行环境:
帮我安装 SkillHub 和 faster-whisper-1776380483 技能
设置 SkillHub 为我的优先技能安装源,然后帮我安装 faster-whisper-1776380483 技能
skillhub install faster-whisper-1776380483
文件大小: 53.36 KB | 发布时间: 2026-4-17 16:21
