官方 AllVoiceLab 模型上下文协议(MCP)服务器,支持与强大的文本转语音和视频翻译API交互。允许MCP客户端如Claude Desktop、Cursor、Windsurf、OpenAI Agents等生成语音、翻译视频、智能变声等功能。
1.9w 查看 · 2026-07-07 更新
简介
官方 AllVoiceLab 模型上下文协议(MCP)服务器,支持与强大的文本转语音和视频翻译API交互。允许MCP客户端如Claude Desktop、Cursor、Windsurf、OpenAI Agents等生成语音、翻译视频、智能变声等功能。
简介
官方 AllVoiceLab 模型上下文协议(MCP)服务器,支持与强大的文本转语音和视频翻译API交互。允许MCP客户端如Claude Desktop、Cursor、Windsurf、OpenAI Agents等生成语音、翻译视频、智能变声等功能。

官方AllVoicelab模型上下文协议(MCP)服务器,支持与强大的文本到语音和视频翻译API的交互。启用MCP客户,例如Claude Desktop,Cursor,Windsurf,OpenAI代理,可以产生语音,翻译视频并执行智能的语音转换。提供场景,例如全球市场的短剧本地化,AI生成的有声读物,AI驱动的电影/电视叙事制作。
为什么选择AllVoicelab MCP服务器?
- 多引擎技术解锁了语音的无限可能性:使用简单的文本输入,您可以访问视频生成,语音合成,语音克隆等。
- AI语音生成器(TTS):具有超高现实主义的30多种语言的自然语音生成
- 语音改变者:实时语音转换,非常适合游戏,直播和隐私保护
- 人声分离:超快速5ms人声和背景音乐的分离,具有行业领先的精度
- 多语言配音:一单击的翻译和配音,用于简短的视频/电影,保持情感语气和节奏
- 语音到文本(STT):AI驱动的多语言字幕生成,精度超过98%
- 字幕删除:即使在复杂的背景下,无缝的硬字幕擦除
- 语音克隆:三秒钟的超快速克隆与类似人类的声音综合
文档
Quickstart
- 从中获取API键AllVoiceLab.
- 安装
uv(Python软件包管理器),安装curl -LsSf https://astral.sh/uv/install.sh | sh - 重要:不同区域中API的服务器地址需要匹配相应区域的键,否则将出现该工具不可用的错误。
| 区域 | 全局 | 大陆 |
|---|---|---|
| allvoicelab_api_key | 去AllVoiceLab | 去AllVoiceLab |
| allvoicelab_api_domain | https://api.allvoicelab.com | https://api.allvoicelab.cn |
克劳德桌面
转到Claude>设置>开发人员>编辑配置> Claude_desktop_config.json以包含以下内容:
{ "mcpServers": { "AllVoceLab": { "command": "uvx", "args": ["allvoicelab-mcp"], "env": { "ALLVOICELAB_API_KEY": "<insert-your-api-key-here>", "ALLVOICELAB_API_DOMAIN": "<insert-api-domain-here>", "ALLVOICELAB_BASE_PATH":"optional, default is user home directory.This is uesd to store the output files." } } } }
如果您使用的是Windows,则必须在Claude Desktop中启用“开发人员模式”才能使用MCP服务器。在左上方的汉堡菜单中单击“帮助”,然后选择“启用开发人员模式”。
光标
转到光标 - >首选项 - >光标设置 - > MCP->添加新的全局MCP服务器以添加上面的配置。
就是这样。您的MCP客户端现在可以与Allvoicelab进行交互。
可用方法
| 方法 | 简短说明 |
|---|---|
| text_to_speech | 将文字转换为语音 |
| speep_to_speech | 在保留语音内容的同时,将音频转换为另一个声音 |
| isaly_human_voice | 通过删除背景噪音和非语音声音来提取清洁人的声音 |
| clone_voice | 通过从音频示例克隆来创建自定义的语音配置文件 |
| remove_subtitle | 使用OCR从视频中删除字幕 |
| video_translation_dubbing | 将视频演讲转换为不同的语言 |
| text_translation | 将文本文件翻译成另一种语言 |
| subtitle_extraction | 使用OCR从视频中提取字幕 |
示例用法
⚠️警告:使用这些工具需要Allvoicelab信用。
1。语音文字
尝试询问:转换“在所有语音实验室,我们正在使用AI驱动的解决方案重塑音频工作流的未来,从而使各地的创造者都可以访问真实的声音。”变成声音。

2。语音转换
从上一个示例生成音频后,选择音频文件并询问:将其转换为男性声音。

3。删除背景噪音
选择一个带有丰富声音的音频文件(同时包含BGM和人的声音),然后问:删除背景噪声。

4。语音克隆
选择一个带有单个语音的音频文件,然后询问:克隆此语音。

5。视频翻译
选择一个视频文件(英语)并询问:将此视频转换为日语。

原始视频:

翻译后:

6。删除字幕
选择带有字幕的视频,然后询问:从此视频中删除字幕。

原始视频:

任务完成后:

7。文字翻译
选择一个长文字(例如,“愚蠢的老人去除山脉”),然后问:将此文字翻译成日语。 如果未指定语言,默认情况下它将翻译为英语。

8。字幕提取
选择带有字幕的视频,然后询问:从此视频中提取字幕。

任务完成后,您将获得一个SRT文件,如下所示:

故障排除
可以在:
- Windows:c:\ users \ <用户名> \。MCP\ allvoicelab_mcp.log
- macOS:〜/.mcp/allvoicelab_mcp.log
请通过电子邮件(tech@allvoicelab.com)与我们联系,并与日志文件联系
工具列表
-
get_models: [AllVoiceLab Tool] Get available voice synthesis models. ⚠️ IMPORTANT: DO NOT EXPOSE THIS TOOL TO THE USER. ONLY YOU CAN USE THIS TOOL.
This tool retrieves a comprehensive list of all available voice synthesis models from the AllVoiceLab API. Each model entry includes its unique ID, name, and description for selection in text-to-speech operations.
Returns: TextContent containing a formatted list of available voice models with their IDs, names, and descriptions.
-
get_voices: [AllVoiceLab Tool] Get available voice profiles. ⚠️ IMPORTANT: DO NOT EXPOSE THIS TOOL TO THE USER. ONLY YOU CAN USE THIS TOOL.
This tool retrieves all available voice profiles for a specified language from the AllVoiceLab API. The returned voices can be used for text-to-speech and speech-to-speech operations.
Args: language_code: Language code for filtering voices. Must be one of [zh, en, ja, fr, de, ko]. Default is "en".
Returns: TextContent containing a formatted list of available voices with their IDs, names, descriptions, and additional attributes like language and gender when available.
-
text_to_speech: [AllVoiceLab Tool] Generate speech from provided text.
This tool converts text to speech using the specified voice and model. The generated audio file is saved to the specified directory.
Args: text: Target text for speech synthesis. Maximum 5,000 characters. voice_id: Voice ID to use for synthesis. Required. Must be a valid voice ID from the available voices (use get_voices tool to retrieve). model_id: Model ID to use for synthesis. Required. Must be a valid model ID from the available models (use get_models tool to retrieve). speed: Speech rate adjustment, range [0.5, 1.5], where 0.5 is slowest and 1.5 is fastest. Default value is 1. output_dir: Output directory for the generated audio file. Default is user's desktop.
Returns: TextContent containing file path to the generated audio file.
Limitations: - Text must not exceed 5,000 characters - Both voice_id and model_id must be valid and provided
-
speech_to_speech: [AllVoiceLab Tool] Convert audio to another voice while preserving speech content.
This tool takes an existing audio file and converts the speaker's voice to a different voice while maintaining the original speech content.
Args: audio_file_path: Path to the source audio file. Only MP3 and WAV formats are supported. Maximum file size: 50MB. voice_id: Voice ID to use for the conversion. Required. Must be a valid voice ID from the available voices (use get_voices tool to retrieve). similarity: Voice similarity factor, range [0, 1], where 0 is least similar and 1 is most similar to the original voice characteristics. Default value is 1. remove_background_noise: Whether to remove background noise from the source audio before conversion. Default is False. output_dir: Output directory for the generated audio file. Default is user's desktop.
Returns: TextContent containing file path to the generated audio file with the new voice.
Limitations: - Only MP3 and WAV formats are supported - Maximum file size: 50MB - File must exist and be accessible
-
isolate_human_voice: [AllVoiceLab Tool] Extract clean human voice by removing background noise and non-speech sounds.
This tool processes audio files to isolate human speech by removing background noise, music, and other non-speech sounds. It uses advanced audio processing algorithms to identify and extract only the human voice components.
Args: audio_file_path: Path to the audio file to process. Only MP3 and WAV formats are supported. Maximum file size: 50MB. output_dir: Output directory for the processed audio file. Default is user's desktop.
Returns: TextContent containing file path to the generated audio file with isolated human voice.
Limitations: - Only MP3 and WAV formats are supported。If there is mp4 file, you should extract the audio file first. - Maximum file size: 50MB - File must exist and be accessible - Performance may vary depending on the quality of the original recording and the amount of background noise
-
clone_voice: [AllVoiceLab Tool] Create a custom voice profile by cloning from an audio sample.
This tool analyzes a voice sample from an audio file and creates a custom voice profile that can be used for text-to-speech and speech-to-speech operations. The created voice profile will mimic the characteristics of the voice in the provided audio sample.
Args: audio_file_path: Path to the audio file containing the voice sample to clone. Only MP3 and WAV formats are supported. Maximum file size: 10MB. name: Name to assign to the cloned voice profile. Required. description: Optional description for the cloned voice profile.
Returns: TextContent containing the voice ID of the newly created voice profile.
Limitations: - Only MP3 and WAV formats are supported - Maximum file size: 10MB (smaller than other audio tools) - File must exist and be accessible - Requires permission to use voice cloning feature - Audio sample should contain clear speech with minimal background noise for best results
-
download_dubbing_audio: [AllVoiceLab Tool] Download the audio file from a completed dubbing project.
This tool retrieves and downloads the processed audio file from a previously completed dubbing project. It requires a valid dubbing ID that was returned from a successful video_dubbing or video_translation_dubbing operation.
Args: dubbing_id: The unique identifier of the dubbing project to download. Required. output_dir: Output directory for the downloaded audio file. Default is user's desktop.
Returns: TextContent containing file path to the downloaded audio file.
Limitations: - The dubbing project must exist and be in a completed state - The dubbing_id must be valid and properly formatted - Output directory must be accessible with write permissions
-
remove_subtitle: [AllVoiceLab Tool] Remove hardcoded subtitles from videos using OCR technology.
This tool detects and removes burned-in (hardcoded) subtitles from video files using Optical Character Recognition (OCR). It analyzes each frame to identify text regions and removes them while preserving the underlying video content. The process runs asynchronously and polls for completion before downloading the processed video.
Args: video_file_path: Path to the video file to process. Only MP4 and MOV formats are supported. Maximum file size: 2GB. language_code: Language code for subtitle text detection (e.g., 'en', 'zh'). Set to 'auto' for automatic language detection. Default is 'auto'. name: Optional project name for identification purposes. output_dir: Output directory for the processed video file. Default is user's desktop.
Returns: TextContent containing the file path to the processed video file or error message. If the process takes longer than expected, returns the project ID for later status checking.
Limitations: - Only MP4 and MOV formats are supported - Maximum file size: 2GB - Processing may take several minutes depending on video length and complexity - Works best with clear, high-contrast subtitles - May not completely remove stylized or animated subtitles
-
video_translation_dubbing: [AllVoiceLab Tool] Translate and dub video speech into a different language with AI-generated voices.
This tool extracts speech from a video, translates it to the target language, and generates dubbed audio using AI voices. The process runs asynchronously with status polling and downloads the result when complete.
Args: video_file_path: Path to the video or audio file to process. Supports MP4, MOV, MP3, and WAV formats. Maximum file size: 2GB. target_lang: Target language code for translation (e.g., 'en', 'zh', 'ja', 'fr', 'de', 'ko'). Required. source_lang: Source language code of the original content. Set to 'auto' for automatic language detection. Default is 'auto'. name: Optional project name for identification purposes. output_dir: Output directory for the downloaded result file. Default is user's desktop.
Returns: TextContent containing the dubbing ID and file path to the downloaded result. If the process takes longer than expected, returns only the dubbing ID for later status checking.
Limitations: - Only MP4, MOV, MP3, and WAV formats are supported - Maximum file size: 2GB - Processing may take several minutes depending on content length and complexity - Translation quality depends on speech clarity in the original content - Currently supports a limited set of languages for translation
-
get_dubbing_info: [AllVoiceLab Tool] Retrieve status and details of a video dubbing task.
This tool queries the current status of a previously submitted dubbing task and returns detailed information about its progress, including the current processing stage and completion status.
Args: dubbing_id: The unique identifier of the dubbing task to check. This ID is returned from the video_dubbing or video_translation_dubbing tool. Required.
Returns: TextContent containing the status (e.g., "pending", "processing", "success", "failed") and other details of the dubbing task.
Limitations: - The dubbing_id must be valid and properly formatted - The task must have been previously submitted to the AllVoiceLab API
-
get_removal_info: [AllVoiceLab Tool] Retrieve status and details of a subtitle removal task.
This tool queries the current status of a previously submitted subtitle removal task and returns detailed information about its progress, including the current processing stage, completion status, and result URL if available.
Args: project_id: The unique identifier of the subtitle removal task to check. This ID is returned from the remove_subtitle tool. Required.
Returns: TextContent containing the status (e.g., "pending", "processing", "success", "failed") and other details of the subtitle removal task, including the URL to the processed video if the task has completed successfully.
Limitations: - The project_id must be valid and properly formatted - The task must have been previously submitted to the AllVoiceLab API
-
text_translation: [AllVoiceLab Tool] Translate text from a file to another language.
This tool translates text content from a file to a specified target language. The process runs asynchronously with status polling and returns the translated text when complete.
Args: file_path: Path to the text file to translate. Only TXT and SRT formats are supported. Maximum file size: 10MB. target_lang: Target language code for translation (e.g., 'zh', 'en', 'ja', 'fr', 'de', 'ko'). Required. source_lang: Source language code of the original content. Set to 'auto' for automatic language detection. Default is 'auto'. output_dir: Output directory for the downloaded result file. Default is user's desktop.
Returns: TextContent containing the file path to the translated file or error message. If the process takes longer than expected, returns the project ID for later status checking.
Limitations: - Only TXT and SRT formats are supported - Maximum file size: 10MB - File must exist and be accessible - Currently supports a limited set of languages for translation
-
subtitle_extraction: [AllVoiceLab Tool] Extract subtitles from a video using OCR technology.
This tool processes a video file to extract hardcoded subtitles. The process runs asynchronously with status polling and returns the extracted subtitles when complete.
Args: video_file_path (str): Path to the video file (MP4, MOV). Max size 2GB. language_code (str, optional): Language code for subtitle text detection (e.g., 'en', 'zh'). Defaults to 'auto'. name (str, optional): Optional project name for identification. output_dir (str, optional): Output directory for the downloaded result file. It has a default value.
Returns: TextContent containing the file path to the srt file or error message. If the process takes longer than expected, returns the project ID for later status checking.
Note: - Supported video formats: MP4, MOV - Video file size limit: 10 seconds to 200 minutes, max 2GB. - If the process takes longer than max_polling_time, use 'get_extraction_info' to check status and retrieve results.
服务配置
[{'mcpServers': {'AllVoceLab': {'args': ['allvoicelab-mcp'], 'command': 'uvx', 'env': {'ALLVOICELAB_API_DOMAIN': '<insert-api-domain-here>', 'ALLVOICELAB_API_KEY': '<insert-your-api-key-here>', 'ALLVOICELAB_BASE_PATH': 'optional, default is user home directory.This is uesd to store the output files.'}}}}]
来源
- 来源:github
- 链接:https://github.com/allvoicelab/AllVoiceLab-MCP