Voice Processing · Developer Tools

Audio Interface logo

Audio Interface

gongrzhe/audio-mcp-server

An MCP (Model Context Protocol) server that provides audio input/output capabilities for AI assistants like Claude. This server enables Claude to interact with your computer's audio system, including recording from microphones and playing audio through speakers.

Install

npx -y @smithery/cli install @GongRzhe/Audio-MCP-Server --client claude

Client configuration

{
  "mcpServers": {
    "audio-interface": {
      "command": "/path/to/your/.venv/bin/python",
      "args": [
        "/path/to/your/audio_server.py"
      ],
      "env": {
        "PYTHONPATH": "<PYTHONPATH>"
      }
    }
  }
}
GitHub stars
2
Author
@gongrzhe
Category
Voice Processing, Developer Tools
License
MIT
Updated
Oct 6, 2026

Features

  • List Audio Devices: View all available microphones and speakers on your system

  • Record Audio: Capture audio from any microphone with customizable duration and quality

  • Playback Recordings: Play back your most recent recording

  • Audio File Playback: Play audio files through your speakers

  • Text-to-Speech: (Placeholder for future implementation)

Tools (5)

  • list_audio_devices

    Lists all available audio input and output devices on your system.

  • record_audio

    Records audio from your microphone.

  • play_latest_recording

    Plays back the most recently recorded audio.

  • play_audio

    Placeholder for text-to-speech functionality.

  • play_audio_file

    Plays an audio file through your speakers.

Details on this page are taken from the project's README. Open README

Supported clients

Clients mentioned in this server's README:

View all
Claude Desktop logo

Claude Desktop

Desktop · Freemium · Proprietary

Anthropic's official Claude AI desktop application. Supports MCP servers to extend functionality.

WindowsMacOS

Related MCP servers

More servers
Elevenlabs MCP logo

Elevenlabs MCP

elevenlabs/elevenlabs-mcp

620

ElevenLabs official MCP server providing text-to-speech and audio processing API interaction.

Voice Processing
MiniMax MCP Server logo

MiniMax MCP Server

MiniMax-AI/MiniMax-MCP

302

Official MiniMax Model Context Protocol (MCP) server that enables interaction with powerful Text to Speech and video/image generation APIs. This server allows MCP clients like Claude Desktop, Cursor, Windsurf, OpenAI Agents and others to generate speech, clone voices, generate video, generate image and more.

Voice Processing
Elevenlabs MCP Server logo

Elevenlabs MCP Server

mamertofabian/elevenlabs-mcp-server

76

A Model Context Protocol (MCP) server that integrates with ElevenLabs text-to-speech API, featuring both a server component and a sample web-based MCP Client (SvelteKit) for managing voice generation tasks.

Voice Processing
Smart Pet with MCP logo

Smart Pet with MCP

shijianzhong/smart-pet-with-mcp

46

An intelligent pet companion application based on the MCP protocol, enabling interaction with virtual pets through speech recognition and natural language processing, with support for multiple platforms.

Voice Processing
Minimax MCP Tools logo

Minimax MCP Tools

PsychArch/minimax-mcp-tools

45

A Model Context Protocol (MCP) server for Minimax AI integration, providing async image generation and text-to-speech with advanced rate limiting and error handling.

Voice Processing
Speech Interface (Faster Whisper) logo

Speech Interface (Faster Whisper)

kvadratni/speech-mcp

33

Speech MCP provides a voice interface for Goose, allowing users to interact through speech rather than text. It includes.

Voice Processing