
Elevenlabs MCP
elevenlabs/elevenlabs-mcp
ElevenLabs official MCP server providing text-to-speech and audio processing API interaction.
Voice Processing · Developer Tools
AIO-2030/mcp-audio
mcp-audio is an AIO-2030 compliant MCP plugin that performs voice-to-text transcription using the Audio speech recognition API.
docker run --env-file .env -p 8080:8080 mcp-audioAUDIO_URLAPI_KEYIt exposes the identify_voice method via both multipart/form-data and base64 formats, supports the AIO tools.call protocol, and returns JSON-RPC structured outputs.
Fully AIO-compliant MCP plugin (/tools.call, /help)
Converts.wav/.mp3 audio files to transcripts using SiliconFlow
API key managed securely via.env file
Docker-compatible and minimal dependencies
Registration-ready for AIO endpoint registry
Details on this page are taken from the project's README. Open README

elevenlabs/elevenlabs-mcp
ElevenLabs official MCP server providing text-to-speech and audio processing API interaction.

MiniMax-AI/MiniMax-MCP
Official MiniMax Model Context Protocol (MCP) server that enables interaction with powerful Text to Speech and video/image generation APIs. This server allows MCP clients like Claude Desktop, Cursor, Windsurf, OpenAI Agents and others to generate speech, clone voices, generate video, generate image and more.

mamertofabian/elevenlabs-mcp-server
A Model Context Protocol (MCP) server that integrates with ElevenLabs text-to-speech API, featuring both a server component and a sample web-based MCP Client (SvelteKit) for managing voice generation tasks.

shijianzhong/smart-pet-with-mcp
An intelligent pet companion application based on the MCP protocol, enabling interaction with virtual pets through speech recognition and natural language processing, with support for multiple platforms.

PsychArch/minimax-mcp-tools
A Model Context Protocol (MCP) server for Minimax AI integration, providing async image generation and text-to-speech with advanced rate limiting and error handling.

kvadratni/speech-mcp
Speech MCP provides a voice interface for Goose, allowing users to interact through speech rather than text. It includes.