一个使用Windows原生语音服务提供文本转语音和语音转文本功能的服务器,无需外部依赖。
speech-processing
89 查看 · 2026-07-07 更新
简介
一个使用Windows原生语音服务提供文本转语音和语音转文本功能的服务器,无需外部依赖。
简介
一个使用Windows原生语音服务提供文本转语音和语音转文本功能的服务器,无需外部依赖。
MS-Lucidia-Voice-Gateway-MCP
一个使用 Windows 内置语音服务提供文本转语音和语音转文本功能的模型上下文协议(MCP)服务器。该服务器通过 PowerShell 命令利用本地 Windows 语音 API (SAPI),无需外部 API 或服务。
特性
- 使用 Windows SAPI 语音进行文本转语音 (TTS)
- 使用 Windows 语音识别进行语音转文本 (STT)
- 简单的网页界面用于测试
- 无外部 API 依赖
- 使用 Windows 本机功能
先决条件
- 启用了语音识别的 Windows 10/11
- Node.js 16+
- PowerShell
安装
- 克隆仓库:
git clone https://github.com/ExpressionsBot/MS-Lucidia-Voice-Gateway-MCP.git cd MS-Lucidia-Voice-Gateway-MCP
- 安装依赖项:
npm install
- 构建项目:
npm run build
使用
测试界面
- 启动测试服务器:
npm run test
- 在浏览器中打开
http://localhost:3000 - 使用网页界面测试 TTS 和 STT 功能
可用工具
text_to_speech
使用 Windows SAPI 将文本转换为语音。
参数:
text(必需): 要转换为语音的文本voice(可选): 要使用的语音(例如:"Microsoft David Desktop")speed(可选): 语速从 0.5 到 2.0(默认:1.0)
示例:
fetch('http://localhost:3000/tts', { method: 'POST', headers: { 'Content-Type': 'application/json' }, body: JSON.stringify({ text: "Hello, this is a test", voice: "Microsoft David Desktop", speed: 1.0 }) });
speech_to_text
录制音频并使用 Windows 语音识别将其转换为文本。
参数:
duration(可选): 录制时长(秒)(默认:5,最大值:60)
示例:
fetch('http://localhost:3000/stt', { method: 'POST', headers: { 'Content-Type': 'application/json' }, body: JSON.stringify({ duration: 5 }) }).then(response => response.json()) .then(data => console.log(data.text));
故障排除
-
确保启用了 Windows 语音识别:
- 打开 Windows 设置
- 转到时间和语言 > 语音
- 启用语音识别
-
检查可用的语音:
- 打开 PowerShell 并运行:
Add-Type -AssemblyName System.Speech (New-Object System.Speech.Synthesis.SpeechSynthesizer).GetInstalledVoices().VoiceInfo.Name -
测试语音识别:
- 在 Windows 设置中打开语音识别
- 如果尚未完成,请运行设置向导
- 测试 Windows 是否可以识别您的声音
贡献
- 分叉仓库
- 创建你的功能分支
- 提交你的更改
- 推送到分支
- 创建一个新的 Pull Request
许可证
MIT