Provider-agnostic MCP server for secure, self-hosted audio transcription with Microsoft MAI, ElevenLabs Scribe, or Groq Whisper.
- ✓Open-source license (Apache-2.0)
- ✓Actively maintained (<30d)
- ✓Clear description
- ✓Topics declared
- ✓Documented (README)
claude mcp add open-transcribe-mcp -- uvx open-transcribe-mcp{
"mcpServers": {
"open-transcribe-mcp": {
"command": "uvx",
"args": ["open-transcribe-mcp"],
"env": {
"OT_HOST": "<ot_host>",
"OT_MICROSOFT__API_KEY": "<ot_microsoft__api_key>",
"OT_SECURITY__BEARER_TOKEN": "<ot_security__bearer_token>"
}
}
}
}OT_HOSTOT_MICROSOFT__API_KEYOT_SECURITY__BEARER_TOKENMCP Servers overview
# OpenTranscribe MCP
<!-- mcp-name: io.github.fbossiere/open-transcribe-mcp -->
[](https://github.com/fbossiere/open-transcribe-mcp/actions/workflows/ci.yml)
[](https://github.com/fbossiere/open-transcribe-mcp/actions/workflows/codeql.yml)
[](LICENSE)
<p align="center">
<img src="docs/assets/open-transcribe-hero.jpg" alt="OpenTranscribe MCP turns recorder audio into a provider-independent transcript" width="100%">
</p>
**Own the recorder. Choose the intelligence.**
OpenTranscribe is an open-source MCP server that routes audio to the speech-to-text model of your choice and returns one provider-independent transcript schema.
It is built around a simple idea: buying a great recorder should not lock you into one transcription subscription. Use Plaud, a phone, an open-source wearable, or any other audio source you are authorized to access, then choose Microsoft MAI, ElevenLabs Scribe, or Groq Whisper without changing the downstream workflow.
OpenTranscribe does not jailbreak hardware or bypass access controls. It works only with audio the operator is authorized to access.
## How it works
<picture>
<source media="(prefers-color-scheme: dark)" srcset="docs/assets/transcription-pipeline-dark.jpg">
<source media="(prefers-color-scheme: light)" srcset="docs/assets/transcription-pipeline-light.jpg">
<img src="docs/assets/transcription-pipeline-light.jpg" alt="Audio flows from a recorder through OpenTranscribe MCP and a selectable speech-to-text provider into one canonical transcript schema for downstream workflows" width="100%">
</picture>
OpenTranscribe keeps the integration boundary stable: the recorder supplies an audio file or HTTPS URL, the server selects or calls the requested speech-to-text provider, and downstream tools receive the same canonical transcript shape.
## What v1 ships
- MCP Streamable HTTP at `/mcp`, with stateless operation and bearer or external OIDC authentication
- `transcribe_audio`, `list_transcription_models`, `estimate_transcription_cost`, `get_transcript_chunk`, and `delete_transcript`
- Microsoft `MAI-Transcribe-2`, ElevenLabs `scribe-v2`, and Groq Whisper adapters
- explicit capability negotiation, routing policies, retries, and observable fallbacks
- HTTPS URL passthrough when supported and a bounded streaming proxy otherwise
- SSRF controls, signed-URL redaction, zero content logging, and cost/resource limits
- disabled-by-default retention; optional memory or S3-compatible temporary result storage
- Docker, production-oriented Scaleway Serverless Containers Terraform, tests, and GitHub Actions CI
## On the Linux desktop
An optional Debian package, **OpenTranscribe Setup**, bundles the Python runtime, the engine, and
a native setup application. Install it, add a provider key, and connect a local MCP client — no
Python, no `uv`, no configuration files. The engine then runs as one client-owned STDIO process
that opens no network socket.
The hosted and command-line paths below are unaffected by it. See
[OpenTranscribe Setup on Linux](docs/desktop.md); its release evidence lives in
[docs/desktop-acceptance.md](docs/desktop-acceptance.md), where every scenario starts at *Not
tested*.
## Ten-minute quickstart
Requirements: Python 3.12 and [uv](https://docs.astral.sh/uv/), or Docker.
```bash
git clone https://github.com/fbossiere/open-transcribe-mcp.git
cd open-transcribe-mcp
```
Create a private `.env` containing only the settings for your selected provider and a strong
random MCP bearer token. For Microsoft:
```dotenv
OT_ENVIRONMENT=prod
OT_HOST=127.0.0.1
OT_DEFAULT_PROVIDER=microsoft
OT_DEFAULT_MODEL=MAI-Transcribe-2
OT_MICROSOFT__ENDPOINT=https://YOUR-RESOURCE.cognitiveservices.azure.com
OT_MICROSOFT__API_KEY=YOUR-KEY
OT_SECURITY__AUTH_MODE=bearer
OT_SECURITY__BEARER_TOKEN=YOUR-RANDOM-TOKEN
```
Use the complete [Groq, ElevenLabs or multi-provider recipes](docs/configuration.md#choose-one-or-more-transcription-providers)
for other services. Omit unused settings rather than leaving empty URLs or numbers from
[.env.example](.env.example), which is a configuration inventory. Protect your file with
`chmod 600 .env`.
Start the server:
```bash
uv sync
uv run open-transcribe-mcp
```
The MCP endpoint is `http://localhost:8000/mcp`; probes are available at `/healthz` and `/readyz`.
Run the included bilingual, two-voice synthetic recording through the connected provider:
```bash
uv run python examples/transcribe.py --token YOUR-RANDOM-TOKEN
```
The [fixture and reference transcript](tests/fixtures/README.md) are non-sensitive and
redistributable. Pass another authorized public HTTPS audio URL as the first argument to use your
own source.
Connect a FastMCP client:
```python
import asyncio
from fastmcp import Client
async def main() -> None:
async with Client("http://localhost:8000/mcp", auth="YOUR-RANDOM-TOKEN") as client:
models = await client.call_tool("list_transcription_models", {})
print(models)
result = await client.call_tool(
"transcribe_audio",
{
"request": {
"source": {"type": "url", "url": "https://example.org/authorized-audio.mp3"},
"provider": "auto",
"routing_policy": "quality",
}
},
)
print(result)
asyncio.run(main())
```
Provider choice does not alter the response contract. Set `provider` and `model` to switch explicitly, or use `auto` with `default`, `quality`, `cost`, or `latency` routing.
`diarization`, `timestamps`, and `transcript_style` are unset above on purpose: an unset capability is not requested, so the request routes to any configured model and the response metadata reports what that model applied. Stating one makes it a requirement — `"diarization": True` excludes every model that cannot diarize rather than quietly returning a single-speaker transcript. See [providers and capabilities](docs/providers.md).
## Docker
Use the published image and the private environment file from the quickstart:
```bash
docker run --rm --env-file .env \
-e OT_ENVIRONMENT=prod -e OT_HOST=0.0.0.0 \
-p 127.0.0.1:8000:8000 \
ghcr.io/fbossiere/open-transcribe-mcp:1.2.1
```
This binds the host port locally; put an HTTPS proxy in front when hosting remotely. You can also
build from source with `docker build -t open-transcribe-mcp:1.2.1 .`. Inject provider credentials
at runtime. The image includes the `s3` extra, but result storage stays disabled unless configured.
## Deploy on Scaleway
The reference [`infra/scaleway`](infra/scaleway/README.md) module provisions a private image
registry, a Serverless Container with scale-to-zero (0–3 instances), HTTPS ingress and health
probes. It can optionally add a TTL-bound result bucket and runtime identity.
1. Create a dedicated Scaleway Project and a deployment API key with the
[documented IAM permissions](docs/deploy-scaleway.md#1-create-the-deployment-api-key).
Supply `SCW_ACCESS_KEY` and `SCW_SECRET_KEY` to the deployment shell.
2. Copy `infra/scaleway/terraform.tfvars.example` to `terraform.tfvars` in that directory.
Fill your project ID, provider keys, authentication settings and `image_tag = "1.2.1"`.
The template configures ElevenLabs and Groq; review the ElevenLabs retention choice before use.
3. Bootstrap the registry, copy the public release image into it, then review and apply the
complete deployment. In `infra/scaleway`:
```bash
terraform init
terraform apply -target=scaleway_registry_namespace.this
IMAGE_REFERENCE="$(terraform output -raw image_reference)"
REGISTRY_ENDPOINT="$(terraform output -raw registry_endpoint)"
printf '%s' "$SCW_SECRET_KEY" | docker login "${REGISTRY_ENDPOINT%%/*}" \
--username nologin --password-stdin
docker buildx imagetools create --tag "$IMAGE_REFERENCE" \
ghcr.io/fbossiere/open-transcribe-mcp:1.2.1
terraform plan -out=open-transcribe.tfplan
terraform apply open-transcribe.tfplan
terraform output -raw mcp_endpoint
```
4. Check `/healthz`, `/readyz`, MCP authentication and one short authorized transcription.
The [full Scaleway guide](docs/deploy-scaleway.md) covers release digest verification,
adoption of an existing deployment, OIDC setup, upgrades, rollback and result-store permissions.
Bootstrap applies are for new infrastructure. Import an existing registry, namespace and
container into state first to preserve their IDs and endpoint. Keep real `terraform.tfvars`,
state and saved plans private: state and plans can contain secrets even when output is redacted.
Publishing a release does not automatically update your Scaleway deployment.
## Configuration and service combinations
Choose the provider, MCP authentication and result storage independently. Provider keys stay
on the server and are never accepted as tool arguments. See the full
[configuration guide](docs/configuration.md) for complete environment/Terraform examples,
settings precedence, key permissions and troubleshooting.
| Transcription services | Configuration | Use and limitations |
| --- | --- | --- |
| Microsoft only | Microsoft endpoint + key; defaults `microsoft` / `MAI-Transcribe-2` | Diarization and clean/verbatim output |
| ElevenLabs only | ElevenLabs key; defaults `elevenlabs` / `scribe-v2` | Diarization; choose eligible zero retention or explicitly accepted standard mode |
| Groq only | Groq key; defaults `groq` / a Whisper model | Verbatim transcription; no speaker diarization |
| Multiple providers | Each provider's credentials; one matching default pair | Requests can select any configured provider; compatibility is checked before ranking |
`OT_DEFAULT_PROVIDER` and `OT_DEFAULT_MODEL` are preferencesWhat people ask about open-transcribe-mcp
What is fbossiere/open-transcribe-mcp?
+
fbossiere/open-transcribe-mcp is mcp servers for the Claude AI ecosystem. Provider-agnostic MCP server for secure, self-hosted audio transcription with Microsoft MAI, ElevenLabs Scribe, or Groq Whisper. It has 1 GitHub stars and its last recorded update is dated 2026-10-04.
How do I install open-transcribe-mcp?
+
You can install open-transcribe-mcp by cloning the repository (https://github.com/fbossiere/open-transcribe-mcp) or following the README instructions on GitHub. ClaudeWave also provides quick install blocks on this page.
Is fbossiere/open-transcribe-mcp safe to use?
+
Our security agent has analyzed fbossiere/open-transcribe-mcp and assigned a Trust Score of 95/100 (tier: Verified). See the full breakdown of passed checks and flags on this page.
Who maintains fbossiere/open-transcribe-mcp?
+
fbossiere/open-transcribe-mcp is maintained by fbossiere. The last recorded GitHub activity is dated 2026-10-04, with 7 open issues.
Are there alternatives to open-transcribe-mcp?
+
Yes. On ClaudeWave you can browse similar mcp servers at /categories/mcp, sorted by popularity or recent activity.
Deploy open-transcribe-mcp to your cloud
Ship this repo to production in minutes. Each platform spins up its own environment with editable env vars.
Maintain this repo? Add a badge to your README
Drop the badge into your GitHub README to show it's tracked on ClaudeWave. Each badge links back to this page and reflects the live Trust Score.
[](https://claudewave.com/repo/fbossiere-open-transcribe-mcp)<a href="https://claudewave.com/repo/fbossiere-open-transcribe-mcp"><img src="https://claudewave.com/api/badge/fbossiere-open-transcribe-mcp" alt="Featured on ClaudeWave: fbossiere/open-transcribe-mcp" width="320" height="64" /></a>More MCP Servers
Fair-code workflow automation platform with native AI capabilities. Combine visual building with custom code, self-host or cloud, 400+ integrations.
User-friendly AI Interface (Supports Ollama, OpenAI API, ...)
An open-source AI agent that brings the power of Gemini directly into your terminal.
Real-time global intelligence dashboard. AI-powered news aggregation, geopolitical monitoring, and infrastructure tracking in a unified situational awareness interface
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
The fastest path to AI-powered full stack observability, even for lean teams.