Unified MCP server for DevOps engineers — query and manage Kubernetes, ArgoCD, Prometheus, and PagerDuty from any MCP-compatible AI agent.
- ✓Open-source license (MIT)
- ✓Actively maintained (<30d)
- ✓Clear description
- ✓Topics declared
- ✓Documented (README)
claude mcp add devops-mcp -- npx -y @notharshhaa/devops-mcp{
"mcpServers": {
"devops-mcp": {
"command": "npx",
"args": ["-y", "@notharshhaa/devops-mcp"],
"env": {
"ARGOCD_TOKEN": "<argocd_token>",
"PROMETHEUS_URL": "<prometheus_url>",
"PAGERDUTY_TOKEN": "<pagerduty_token>",
"LOKI_URL": "<loki_url>",
"LOKI_TOKEN": "<loki_token>",
"MCP_HTTP_HOST": "<mcp_http_host>"
}
}
}
}ARGOCD_TOKENPROMETHEUS_URLPAGERDUTY_TOKENLOKI_URLLOKI_TOKENMCP_HTTP_HOSTMCP Servers overview
# devops-mcp
> Unified MCP server for DevOps engineers — query and manage Kubernetes, ArgoCD, Prometheus, and PagerDuty from any MCP-compatible AI agent.
[](https://www.npmjs.com/package/@notharshhaa/devops-mcp)
[](LICENSE)
[](https://modelcontextprotocol.io)
---
## What is this?
`devops-mcp` is an open source [Model Context Protocol](https://modelcontextprotocol.io) server that gives AI agents (Claude, etc.) real-time read and write access to your infrastructure stack — all from a single install.
Instead of copy-pasting `kubectl` output into a chat window, you can ask:
> *"Why is the payments deployment in CrashLoopBackOff?"*
> *"What changed in the last ArgoCD sync for the auth app?"*
> *"Show me the p99 latency for the API gateway over the last hour."*
> *"Who's on call right now and what incidents are open?"*
> *"Debug the payments service - what's wrong with it?"*
...and get live answers, sourced directly from your cluster and tooling.
**Providers included:**
| Prefix | Provider | Transport |
|---|---|---|
| `k8s__*` | Kubernetes (via kubeconfig or in-cluster SA) | client-go |
| `argo__*` | ArgoCD | REST API |
| `prom__*` | Prometheus | HTTP API (PromQL) |
| `pd__*` | PagerDuty | REST API v2 |
| `helm__*` | Helm | CLI (helm binary) |
| `devops__*` | Cross-provider incident debugging | Aggregates all providers |
| `logs__*` | Loki | HTTP API (LogQL) |
---
## Quick start
### Claude Desktop (stdio — recommended)
Add this to `~/.config/claude/claude_desktop_config.json` (macOS: `~/Library/Application Support/Claude/claude_desktop_config.json`):
```json
{
"mcpServers": {
"devops": {
"command": "npx",
"args": ["-y", "@notharshhaa/devops-mcp@latest"],
"env": {
"KUBECONFIG": "/home/you/.kube/config",
"ARGOCD_SERVER": "https://argocd.company.com",
"ARGOCD_TOKEN": "your-argocd-token",
"PROMETHEUS_URL": "http://prometheus.monitoring:9090",
"PAGERDUTY_TOKEN": "your-pd-api-token",
"PAGERDUTY_USER_EMAIL": "you@company.com",
"LOKI_URL": "http://loki.monitoring:3100",
"LOKI_TOKEN": "your-loki-token"
}
}
}
}
```
Restart Claude Desktop. The `devops` server will appear in the tools list.
### Claude Code (CLI)
```bash
claude mcp add devops-mcp -e KUBECONFIG=$HOME/.kube/config \
-e ARGOCD_SERVER=https://argocd.company.com \
-e ARGOCD_TOKEN=... \
-e PROMETHEUS_URL=http://prometheus:9090 \
-e PAGERDUTY_TOKEN=... \
-e LOKI_URL=http://loki.monitoring:3100 \
-e LOKI_TOKEN=... \
-- npx -y @notharshhaa/devops-mcp@latest
```
### Local dev / test
Requires Node.js 20 or newer.
```bash
npx @notharshhaa/devops-mcp
# or clone and run:
git clone https://github.com/NotHarshhaa/devops-mcp
cd devops-mcp
npm install
cp .env.example .env # fill in your values
npm run dev
```
---
## Configuration
All config is via environment variables. Only set the ones for providers you actually use — providers with missing config are silently skipped.
```env
# ── Kubernetes ────────────────────────────────────────────────
KUBECONFIG=/home/user/.kube/config # omit to use in-cluster service account
K8S_CONTEXT=my-prod-context # optional: pin a specific context
K8S_ALLOWED_NAMESPACES=default,backend # optional: restrict namespace access
# ── ArgoCD ───────────────────────────────────────────────────
ARGOCD_SERVER=https://argocd.company.com
ARGOCD_TOKEN=eyJhbGci... # argocd account generate-token
# ── Prometheus ───────────────────────────────────────────────
PROMETHEUS_URL=http://prometheus:9090
PROMETHEUS_BEARER_TOKEN= # optional: for authenticated Prometheus
# ── PagerDuty ────────────────────────────────────────────────
PAGERDUTY_TOKEN=your-api-v2-token
PAGERDUTY_USER_EMAIL=you@company.com # required for PagerDuty incident mutations (From header)
# ── Loki ───────────────────────────────────────────────────
LOKI_URL=http://loki.monitoring:3100
LOKI_TOKEN=your-loki-token
# ── Stateless Streamable HTTP ────────────────────────────────
# For stdio mode (default): no transport config needed
MCP_HTTP_HOST=127.0.0.1 # use 0.0.0.0 inside a container
PORT=3000
MCP_AUTH_TOKEN=shared-secret # optional static Bearer token
MCP_REQUEST_STATE_SECRET=32+-byte-secret # optional MRTR signing key shared by all replicas
MCP_ALLOWED_HOSTS=localhost,127.0.0.1 # required with non-loopback binding
MCP_ALLOWED_ORIGINS= # optional browser Origin hostname allowlist
MCP_CACHE_TTL_MS=60000 # discovery/tools catalog TTL; 0 disables caching
# ── Safety ───────────────────────────────────────────────────
DEVOPS_MCP_DRY_RUN=false # true = block all mutations globally
DEVOPS_MCP_AUDIT_LOG=/var/log/devops-mcp-audit.jsonl
```
---
## Tool reference
All tools follow a three-tier safety model:
- **Read** — safe, no side effects, no confirmation needed
- **Mutate** — defaults to `dry_run: true`; set `dry_run: false` to execute
- **Destructive** — requires `confirm: true`, or a 2026-07-28 client can complete the server's interactive MRTR confirmation
### Kubernetes (`k8s__*`)
| Tool | Tier | Description |
|---|---|---|
| `k8s__list_pods` | read | List pods with status, restarts, node, age |
| `k8s__get_pod_logs` | read | Tail or stream logs from a pod container |
| `k8s__describe_resource` | read | Full describe for any resource type (`pod`, `service`, `configmap`, `secret` [values auto-redacted], `deployment`, `statefulset`, `daemonset`, `job`, `cronjob`) |
| `k8s__get_events` | read | Cluster or namespace events, filterable by reason |
| `k8s__list_deployments` | read | Deployments with replica counts, rollout health, and container images |
| `k8s__get_resource_usage` | read | CPU/mem usage per pod via metrics-server |
| `k8s__get_node_status` | read | Node health, conditions, capacity, allocatable resources, taints |
| `k8s__get_network_policies` | read | Network policies with pod selectors and ingress/egress rules |
| `k8s__get_ingresses` | read | Ingress resources with hosts, paths, backends, TLS config |
| `k8s__list_cronjobs` | read | CronJobs with schedule, last run, active jobs, suspend status |
| `k8s__get_cronjob_status` | read | Detailed CronJob status with recent job history |
| `k8s__diff_resource` | read | Compare current resource state vs last-applied-configuration |
| `k8s__get_hpa` | read | HorizontalPodAutoscaler with current/target metrics and scaling status |
| `k8s__list_pvcs` | read | PersistentVolumeClaims with status, capacity, storage class |
| `k8s__list_services` | read | Services with type, ports, selectors, clusterIP, endpoints |
| `k8s__list_contexts` | read | All kubeconfig contexts and the active one |
| `k8s__switch_context` | mutate | Preview a context selection; set `K8S_CONTEXT` and restart to apply it safely |
| `k8s__scale_deployment` | mutate | Scale replicas with dry-run diff preview |
| `k8s__apply_manifest` | mutate | Apply a manifest string with server-side dry-run |
| `k8s__rollout_restart` | mutate | Trigger rolling restart of a deployment or statefulset |
| `k8s__delete_resource` | destructive | Delete a named resource (`pod`, `service`, `configmap`, `secret`, `deployment`, `statefulset`, `daemonset`, `job`, `cronjob`) — requires direct or interactive confirmation |
### ArgoCD (`argo__*`)
| Tool | Tier | Description |
|---|---|---|
| `argo__list_apps` | read | All apps with health, sync status, source repo |
| `argo__get_app` | read | Full spec and status for one application |
| `argo__get_app_diff` | read | Live diff between git and cluster state |
| `argo__get_app_history` | read | Deployment history with git SHAs and timestamps |
| `argo__get_resource_tree` | read | Full owned resource tree for an app |
| `argo__sync_app` | mutate | Trigger sync — supports dry-run, prune, force |
| `argo__rollback_app` | mutate | Preview rollback to a history revision; set `dry_run: false` to execute |
| `argo__terminate_op` | mutate | Preview cancellation of an in-progress sync; set `dry_run: false` to execute |
### Prometheus (`prom__*`)
| Tool | Tier | Description |
|---|---|---|
| `prom__query` | read | Instant PromQL query with label + value output |
| `prom__query_range` | read | Range query with step, returns time-series data |
| `prom__list_alerts` | read | All alert rules with state (firing / pending / inactive) |
| `prom__get_firing_alerts` | read | Only currently firing alerts with duration |
| `prom__list_targets` | read | All scrape targets with health and last scrape |
| `prom__label_values` | read | Enumerate values for a given label name |
| `prom__metric_metadata` | read | Type, help text, and unit for a metric |
| `prom__compare_periods` | read | 📈 **Compare metrics** between two time windows — detect before/after deployment changes |
| `prom__slo_status` | read | 🎯 **SLO compliance** — error budget remaining, burn rate, time to exhaustion |
| `prom__summarize_service_health` | read | 📊 **Smart summary** - human-readable service health metrics including latency changes, error rate vs SLO, and traffic patterns |
**Example usage:**
```bash
# Get a human-readable health summary
prom__summarize_service_health(service="payments", timeframeMinutes=30, sloThreshold=0.05)
```
**What it outputs:**
- **Latency**: "Latency increased: 120ms → 480ms (+300%)" or "Latency stable: 125ms"
- **Error rate**: "Error rate crossed SLO (5%): 7.2%" or "Error rate within SLO: 2.1%"
- **Traffic**: "Traffic dropped: 500 → 350 req/s (-30%)" or "Traffic spike detected (+150%)"
- **Overall assessment**: Summary of issues and positive indicators
**Why this matters:**
Instead of raw PromQL numbers that require interpretation, this tool provides actionable insights that AI agents can use directly What people ask about devops-mcp
What is NotHarshhaa/devops-mcp?
+
NotHarshhaa/devops-mcp is mcp servers for the Claude AI ecosystem. Unified MCP server for DevOps engineers — query and manage Kubernetes, ArgoCD, Prometheus, and PagerDuty from any MCP-compatible AI agent. It has 3 GitHub stars and its last recorded update is dated 2026-10-05.
How do I install devops-mcp?
+
You can install devops-mcp by cloning the repository (https://github.com/NotHarshhaa/devops-mcp) or following the README instructions on GitHub. ClaudeWave also provides quick install blocks on this page.
Is NotHarshhaa/devops-mcp safe to use?
+
Our security agent has analyzed NotHarshhaa/devops-mcp and assigned a Trust Score of 95/100 (tier: Verified). See the full breakdown of passed checks and flags on this page.
Who maintains NotHarshhaa/devops-mcp?
+
NotHarshhaa/devops-mcp is maintained by NotHarshhaa. The last recorded GitHub activity is dated 2026-10-05, with 0 open issues.
Are there alternatives to devops-mcp?
+
Yes. On ClaudeWave you can browse similar mcp servers at /categories/mcp, sorted by popularity or recent activity.
Deploy devops-mcp to your cloud
Ship this repo to production in minutes. Each platform spins up its own environment with editable env vars.
Maintain this repo? Add a badge to your README
Drop the badge into your GitHub README to show it's tracked on ClaudeWave. Each badge links back to this page and reflects the live Trust Score.
[](https://claudewave.com/repo/notharshhaa-devops-mcp)<a href="https://claudewave.com/repo/notharshhaa-devops-mcp"><img src="https://claudewave.com/api/badge/notharshhaa-devops-mcp" alt="Featured on ClaudeWave: NotHarshhaa/devops-mcp" width="320" height="64" /></a>More MCP Servers
Fair-code workflow automation platform with native AI capabilities. Combine visual building with custom code, self-host or cloud, 400+ integrations.
User-friendly AI Interface (Supports Ollama, OpenAI API, ...)
An open-source AI agent that brings the power of Gemini directly into your terminal.
Real-time global intelligence dashboard. AI-powered news aggregation, geopolitical monitoring, and infrastructure tracking in a unified situational awareness interface
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
The fastest path to AI-powered full stack observability, even for lean teams.