Voice Processing · Education & Learning

MCP Server Whisper logo

MCP Server Whisper

arcaputo3/mcp-server-whisper

MCP Server Whisper provides a standardized way to process audio files through OpenAI's latest transcription and speech services. By implementing the Model Context Protocol, it enables AI assistants like Claude to seamlessly interact with audio processing capabilities.

Install

uvx mcp-server-whisper

Client configuration

{
  "mcpServers": {
    "whisper": {
      "command": "uvx",
      "args": [
        "mcp-server-whisper"
      ],
      "env": {
        "OPENAI_API_KEY": "<OPENAI_API_KEY>",
        "AUDIO_FILES_PATH": "<AUDIO_FILES_PATH>"
      }
    }
  }
}

Environment variables

OPENAI_API_KEYAUDIO_FILES_PATH
Category
Voice Processing, Education & Learning
License
MIT
Updated
Oct 6, 2026

About MCP Server Whisper

A Model Context Protocol (MCP) server for advanced audio transcription and processing using OpenAI's Whisper and GPT-4o models.

Tools (8)

  • list_audio_files

    Lists audio files with comprehensive filtering and sorting options:

  • get_latest_audio

    Gets the most recently modified audio file with model support info

  • convert_audio

    Converts audio files to supported formats (mp3 or wav)

  • compress_audio

    Compresses audio files that exceed size limits

  • transcribe_audio

    Advanced transcription using OpenAI's models:

  • chat_with_audio

    Interactive audio analysis using GPT-4o audio models:

  • transcribe_with_enhancement

    Enhanced transcription with specialized templates:

  • Text-to-Speech

Details on this page are taken from the project's README. Open README

Supported clients

Clients mentioned in this server's README:

View all
Claude Desktop logo

Claude Desktop

Desktop · Freemium · Proprietary

Anthropic's official Claude AI desktop application. Supports MCP servers to extend functionality.

WindowsMacOS

Related MCP servers

More servers
Elevenlabs MCP logo

Elevenlabs MCP

elevenlabs/elevenlabs-mcp

620

ElevenLabs official MCP server providing text-to-speech and audio processing API interaction.

Voice Processing
MiniMax MCP Server logo

MiniMax MCP Server

MiniMax-AI/MiniMax-MCP

302

Official MiniMax Model Context Protocol (MCP) server that enables interaction with powerful Text to Speech and video/image generation APIs. This server allows MCP clients like Claude Desktop, Cursor, Windsurf, OpenAI Agents and others to generate speech, clone voices, generate video, generate image and more.

Voice Processing
Elevenlabs MCP Server logo

Elevenlabs MCP Server

mamertofabian/elevenlabs-mcp-server

76

A Model Context Protocol (MCP) server that integrates with ElevenLabs text-to-speech API, featuring both a server component and a sample web-based MCP Client (SvelteKit) for managing voice generation tasks.

Voice Processing
Smart Pet with MCP logo

Smart Pet with MCP

shijianzhong/smart-pet-with-mcp

46

An intelligent pet companion application based on the MCP protocol, enabling interaction with virtual pets through speech recognition and natural language processing, with support for multiple platforms.

Voice Processing
Minimax MCP Tools logo

Minimax MCP Tools

PsychArch/minimax-mcp-tools

45

A Model Context Protocol (MCP) server for Minimax AI integration, providing async image generation and text-to-speech with advanced rate limiting and error handling.

Voice Processing
Speech Interface (Faster Whisper) logo

Speech Interface (Faster Whisper)

kvadratni/speech-mcp

33

Speech MCP provides a voice interface for Goose, allowing users to interact through speech rather than text. It includes.

Voice Processing