闲社服务运行正常AI智能体自动化平台
o

本地语音转文字

技能标识:openclaw-whisper-voice

Local Whisper speech-to-text for audio files and inbound voice notes on the OpenClaw Gateway host. Use when setting up local transcription for WhatsApp, Telegram, or other audio attachments; when configuring tools.media.audio with a CLI fallback instead of a cloud API; or when you need a reusable shell entrypoint that makes Whisper + ffmpeg work reliably on Linux.

作者:admin | 来源记录:ClawHub
登记来源
ClawHub
版本
V 1.0.0
检测标记
后台标记通过
免费
免费获取
0
收藏
来源与检测标记为平台登记信息,并不代表已展示可复核的检测报告。使用前请核对版本、依赖和所需权限,建议先在隔离环境中运行。
概述
安装方式
版本历史

本地语音转文字

OpenClaw Whisper 语音

使用此技能可使本地 Whisper 转录功能依赖于 OpenClaw 网关主机。

在主机上安装

运行:

bash
{baseDir}/scripts/installlocalwhisper.sh

安装程序将:

  • - 将 Python 包安装到 ~/.local
  • 安装 CPU 兼容的 PyTorch 构建
  • 安装 openai-whisper
  • 安装 imageio-ffmpeg
  • 创建稳定的 ~/.local/bin/whisper 和 ~/.local/bin/ffmpeg 启动器

手动转录文件

当可靠性至关重要时,请使用包装器而非原始 whisper:

bash
{baseDir}/scripts/transcribe.sh /path/to/audio.ogg
{baseDir}/scripts/transcribe.sh /path/to/audio.m4a --model tiny --stdout-only
{baseDir}/scripts/transcribe.sh /path/to/audio.mp3 --task translate --format srt

配置入站 WhatsApp 和 Telegram 语音消息

修改 OpenClaw 配置,使入站音频使用包装器:

json5
{
tools: {
media: {
audio: {
enabled: true,
maxBytes: 20971520,
timeoutSeconds: 120,
models: [
{
type: cli,
command: {baseDir}/scripts/transcribe.sh,
args: [{{MediaPath}}, --model, base, --stdout-only],
timeoutSeconds: 120
}
]
}
}
}
}

模型选择

  • - tiny:最快,准确度最低
  • base:聊天语音消息的最佳默认选择
  • small 或更大:准确度更高,CPU 和内存占用更大

输出规则

  • - 对于 tools.media.audio,使用 --stdout-only,使标准输出仅为转录文本。
  • 对于独立文件转录,使用 --format txt|srt|vtt|json。
  • 首次模型下载将保存到 ~/.cache/whisper。

标签

skill ai

通过对话安装

以下为平台配置的接入选项,并非逐项实测通过。能否安装取决于客户端支持、技能来源和运行环境:

OpenClaw WorkBuddy QClaw Kimi Claude

方式一:安装 SkillHub 和技能

帮我安装 SkillHub 和 openclaw-whisper-voice-1776073981 技能

方式二:设置 SkillHub 为优先技能安装源

设置 SkillHub 为我的优先技能安装源,然后帮我安装 openclaw-whisper-voice-1776073981 技能

通过命令行安装

skillhub install openclaw-whisper-voice-1776073981

下载

⬇ 下载 openclaw-whisper-voice v1.0.0(免费)

文件大小: 3.51 KB | 发布时间: 2026-4-17 15:39

v1.0.0 最新 2026-4-17 15:39
Version 1.0.0 of openclaw-whisper-voice

- Initial release providing local Whisper speech-to-text transcription for audio files and inbound voice notes on the OpenClaw Gateway host.
- Includes an installation script for setting up Python dependencies, a CPU-compatible PyTorch build, and stable CLI launchers for Whisper and ffmpeg.
- Offers a shell wrapper script for reliable manual and automated transcription with support for multiple audio formats and model options.
- Provides configuration guidance for integrating with WhatsApp and Telegram inbound audio using tools.media.audio in OpenClaw.
返回顶部