Skip to main content
ClaudeWave

Governed GPU inference ops (vLLM + Ray Serve): latency RCA, scaling, drain, 30 MCP tools (preview)

SkillsOfficial Registry0 stars0 forksPythonMITUpdated today
ClaudeWave Trust Score
95/100
Verified
Passed
  • Open-source license (MIT)
  • Actively maintained (<30d)
  • Clear description
  • Topics declared
  • Documented (README)
Last scanned: 9/12/2026
Install as a Claude Code skill
Method: Clone
Terminal
git clone https://github.com/AIops-tools/Inference-AIops ~/.claude/skills/inference-aiops
1. Clone the repository into your ~/.claude/skills directory (or copy the skill folder containing SKILL.md).
2. Start a new Claude Code session so the skill registry reloads.
3. Invoke it by name, or let Claude trigger it automatically when the task matches.
💡 If the repo bundles several skills, copy only the folders you need.
Use cases

Skills overview

<!-- mcp-name: io.github.AIops-tools/inference-aiops -->

# Inference AIops

> **Disclaimer**: Community-maintained open-source project. **Not affiliated with, endorsed by, or sponsored by the vLLM or Ray projects or any inference-serving vendor.** Product and trademark names belong to their owners. MIT licensed.

Governed AI-ops for **GPU inference clusters** — **vLLM** (OpenAI API + Prometheus
`/metrics`) and **Ray Serve / Ray Jobs** (Ray dashboard), plus the single-process
serving engines **SGLang** and **TGI (Text Generation Inference)** — with a
**built-in governance harness**: unified audit log, policy engine, token/runaway
budget guard, undo-token recording, and descriptive risk-tier labels on every
audit row. It parses each engine's Prometheus `/metrics` directly (no Prometheus
server required) and
probes the Ray dashboard independently. A bearer token is **optional** (many
stacks run open).

**Serving engines.** vLLM is the flagship (full Ray Serve control plane: scale,
drain, autoscale, LoRA, hot-swap). **SGLang** and **TGI** are supported for
engine-agnostic observability — health, running-model identity, request-latency
metrics, queue depth, and latency RCA — read from each engine's own endpoints and
metric names. Being single-process servers, they have no Ray-shaped scale/drain
API: those writes return a teaching error pointing you at a real horizontal-scale
layer (Ray Serve / Kubernetes / a load balancer).

## What it does

The flagship value is **root-cause analysis**, wrapped in guarded reads and writes:

- **`diagnose_latency_spike`** (flagship RCA) — when TTFT/TPOT/e2e latency
  climbs, it correlates **queue depth** (running vs waiting), **KV-cache
  pressure / preemptions**, and **prefix-cache locality** into a *ranked* cause
  plus the **specific knob to turn** (add replicas, raise `max-num-seqs`, fix
  routing, enlarge KV cache). Every flag is a number, not a black-box verdict.
- **`diagnose_low_utilization`** — the inverse: idle GPUs, over-provisioned
  replicas, or routing that strands a cache-warm replica → what to scale down.
- **Prometheus-native** — reads vLLM's `/metrics` endpoint directly; no
  Prometheus/Grafana deployment needed.
- **Governance-grade** — the **first governance-grade entrant** in this niche:
  audit + budget + risk-tier approval + undo-token + prompt-injection sanitize,
  with **dry-run + double-confirm** on the fragile prod ops (scale-down,
  scale-to-zero, drain, redeploy, hot-swap) the community reports as dangerous.
- **Laptop self-test** — ~80% of the tool self-tests free: vLLM on a single GPU
  or CPU-mock + Ray in one local container (`ray start --head`).

## What this tool does, and does not, decide

It delivers inference-cluster operations — reads and writes — accurately and
efficiently, and records every one of them. It does **not** decide whether a
write is allowed to happen. That is the agent's judgement, or the permission of
the environment you connect it with: restrict the network path so it can only
reach the read/metrics endpoints, or run the Ray dashboard without its
job-submission API, and the writes fail at the server — the place that actually
owns the permission.

So there is no read-only switch, no policy file, no approval gate to configure.
The one thing the tool guarantees is that nothing is silent: **every call, over
MCP and over the CLI alike, lands an audit row** in
`~/.inference-aiops/audit.db`, and destructive writes still capture their
before-state and record an inverse where one exists.

> Each tool declares a `risk_level`, kept in agreement with its `[READ]`/`[WRITE]`
> documentation tag by a test, and carried into the audit row as a descriptive
> tier — so a reviewer can see at a glance that a row was a high-risk
> scale-to-zero. It is a label, not a gate.

Running a smaller / local model? See
[agent-guardrails.md](skills/inference-aiops/references/agent-guardrails.md) — it lists
the guardrails this tool enforces for you (so you don't spend prompt budget
restating them) and gives a ready-made system prompt for what's left.

## Capability matrix (39 MCP tools)

| Group | Tools | Count | R/W (risk) |
|-------|-------|:-----:|:-----------|
| **Metrics & RCA** (vLLM) | `request_metrics`, `queue_depth`, `kv_cache_stats`, `diagnose_latency_spike`, `diagnose_low_utilization` | 5 | read |
| **Engine-agnostic** (vLLM / SGLang / TGI) | `engine_health`, `engine_inventory`, `engine_request_metrics`, `engine_queue_depth`, `diagnose_engine_latency` | 5 | read |
| **Ray Serve (read)** | `serve_deployment_list`, `deployment_status`, `replica_list`, `autoscale_config_get` | 4 | read |
| **Ray Serve (write)** | `scale_replicas_up`, `scale_replicas_down`, `scale_to_zero`, `autoscale_config_update`, `drain_replica` | 5 | write (med / **high**) |
| **Models / vLLM** | `model_list`, `model_info`, `model_is_sleeping`, `lora_load`, `lora_unload` | 5 | read + write (med) |
| **Sleep Mode / vLLM** (needs `VLLM_SERVER_DEV_MODE=1`) | `model_sleep`, `model_wake` | 2 | write (**high** / med) |
| **Ray cluster / jobs / GPU** | `ray_cluster_resources`, `ray_dashboard_status`, `ray_job_list`, `gpu_utilization`, `ray_job_cancel`, `replica_restart` | 6 | read + write (med / **high**) |
| **Deploy lifecycle** | `model_deploy`, `model_undeploy`, `deployment_redeploy`, `routing_policy_update` | 4 | write (med / **high**) |
| **Cost** | `cost_per_token` | 1 | read |

The engine-agnostic group works against **any** supported engine (including
vLLM); use it for SGLang/TGI targets or a uniform view across a mixed fleet. The
Ray Serve / cluster / deploy write groups are vLLM-only (Ray control plane) — they
teach-and-refuse on a SGLang/TGI target.

**23 read, 16 write.** High-risk writes (`scale_replicas_down`,
`scale_to_zero`, `drain_replica`, `lora_unload`, `model_sleep`,
`replica_restart`, `model_undeploy`, `deployment_redeploy`) all support
`dry_run` + double-confirm; reversible writes record an undo descriptor.

> **Sleep Mode requires a dev-mode server.** vLLM registers `/sleep`,
> `/wake_up` and `/is_sleeping` **only** when started with
> `VLLM_SERVER_DEV_MODE=1`. Against any other server these three tools
> report that the route is absent and why, rather than failing vaguely.
> Sleep Mode suspends the **same** model; it does not swap base models —
> serving a different base model means restarting vLLM with a different
> `--model`.

## Install

```bash
uv tool install inference-aiops          # or: pipx install inference-aiops
```

## Quick start

### As a Claude Code plugin

One install gives an agent both the skill and the MCP server:

```
/plugin marketplace add AIops-tools/marketplace
/plugin install inference-aiops@aiops-tools
```

The MCP server is fetched with [uv](https://docs.astral.sh/uv/) and pinned to the
package version this plugin declares, so an audit row can be traced back to the
code that wrote it. Credentials are still configured with `inference-aiops init` — see below.

### As a CLI or standalone MCP server

```bash
inference-aiops init                     # wizard: engine (vllm/sglang/tgi) + host + port + scheme
inference-aiops doctor                   # vLLM: probes Ray + vLLM; SGLang/TGI: engine health + inventory
inference-aiops overview                 # deployments + total replicas + queue backpressure
inference-aiops metrics diagnose         # why is inference slow? ranked RCA + the knob to turn
inference-aiops serve list               # Ray Serve deployments + replica counts
```

Run as an MCP server (stdio) for the full 39-tool surface:

```bash
export INFERENCE_AIOPS_MASTER_PASSWORD=...   # only if a bearer token is stored
inference-aiops mcp
```

The CLI is a convenience subset (`init`, `overview`, `serve …`, `metrics …`,
`secret …`, `doctor`, `mcp`); the full 39 tools are exposed via the MCP server.

## Governance

Every MCP tool passes through the bundled `@governed_tool` harness. It does not
decide whether a write is permitted — see *What this tool does, and does not,
decide* above — but it records every call:

- **Audit** — every call (params, result, status, duration, risk tier, and any
  approver/rationale annotation) logged to `~/.inference-aiops/audit.db`
  (relocatable via `INFERENCE_AIOPS_HOME`).
- **Budget / runaway guard** — a safety backstop, not authorization: token and
  call budgets trip a circuit breaker on tight poll/retry loops.
- **Risk tier** — each audit row carries a descriptive tier derived from the
  tool's `risk_level`; it is a label, not a gate. `INFERENCE_AUDIT_APPROVED_BY`
  / `INFERENCE_AUDIT_RATIONALE` are optional annotations recorded when set,
  never required.
- **Undo recording** — reversible writes (scale, autoscale-config, routing,
  hot-swap, LoRA load) record an inverse descriptor.

## Supported scope + limitations

Behaviour is exercised by the test suite against mocked vLLM `/metrics`, vLLM
OpenAI API, and Ray dashboard responses. **~80% of the tool self-tests on a
laptop** — vLLM on a single GPU or CPU-mock plus a local one-node Ray head. It
has not been run against a live production cluster; see
[`docs/VERIFICATION.md`](docs/VERIFICATION.md) for the live-verification
checklist.

Unverified against real hardware / topology:

- multi-GPU **tensor-parallel / pipeline-parallel** deployments,
- real GPU **thermal / throttle** telemetry (utilisation is best-effort from
  the Ray dashboard's `/api/nodes`),
- **multi-node drain** and node-reboot orchestration.

The fastest live check is `inference-aiops doctor`; the full checklist lives in
[`docs/VERIFICATION.md`](docs/VERIFICATION.md).

## Missing a capability?

This is the GPU-inference member of the AIops-tools family (governed AI-ops with
audit + budget + undo + risk tiers). If a vLLM or Ray capability you need is
missing, or your stack speaks a dialect these tools don't yet handle — open an
issue or a PR. Contributions welcome.
agent-skillsai-opsgovernancellm-inferencemcprayvllm

What people ask about Inference-AIops

What is AIops-tools/Inference-AIops?

+

AIops-tools/Inference-AIops is skills for the Claude AI ecosystem. Governed GPU inference ops (vLLM + Ray Serve): latency RCA, scaling, drain, 30 MCP tools (preview) It has 0 GitHub stars and its last recorded update is dated 2026-09-12.

How do I install Inference-AIops?

+

You can install Inference-AIops by cloning the repository (https://github.com/AIops-tools/Inference-AIops) or following the README instructions on GitHub. ClaudeWave also provides quick install blocks on this page.

Is AIops-tools/Inference-AIops safe to use?

+

Our security agent has analyzed AIops-tools/Inference-AIops and assigned a Trust Score of 95/100 (tier: Verified). See the full breakdown of passed checks and flags on this page.

Who maintains AIops-tools/Inference-AIops?

+

AIops-tools/Inference-AIops is maintained by AIops-tools. The last recorded GitHub activity is dated 2026-09-12, with 0 open issues.

Are there alternatives to Inference-AIops?

+

Yes. On ClaudeWave you can browse similar skills at /categories/skills, sorted by popularity or recent activity.

Deploy Inference-AIops to your cloud

Ship this repo to production in minutes. Each platform spins up its own environment with editable env vars.

Maintain this repo? Add a badge to your README

Drop the badge into your GitHub README to show it's tracked on ClaudeWave. Each badge links back to this page and reflects the live Trust Score.

Featured on ClaudeWave: AIops-tools/Inference-AIops
[![Featured on ClaudeWave](https://claudewave.com/api/badge/aiops-tools-inference-aiops)](https://claudewave.com/repo/aiops-tools-inference-aiops)
<a href="https://claudewave.com/repo/aiops-tools-inference-aiops"><img src="https://claudewave.com/api/badge/aiops-tools-inference-aiops" alt="Featured on ClaudeWave: AIops-tools/Inference-AIops" width="320" height="64" /></a>
farion1231
cc-switch
today

A cross-platform desktop All-in-One assistant for Claude Code, Codex, OpenCode, OpenClaw, Grok Build & Hermes Agent. Only official website: ccswitch.io

132.4k9.1kRust
Skillsai-toolsclaude-codeInstall
Egonex-AI
Understand-Anything
today

Graphs that teach > graphs that impress. Turn any code into an interactive knowledge graph you can explore, search, and ask questions about. Works with Claude Code, Codex, Cursor, Copilot, Gemini CLI, and more.

82.1k6.9kTypeScript
Skillsantigravity-skillsbusiness-knowledgeInstall
code-yeongyu
oh-my-openagent
today

OmO: Just type "mass ulw" keyword with your prompt. Now you are the master of graph engineering.

69k5.7kTypeScript
Skillsaiai-agentsInstall
tt-a1i
archify
today

Agent skill for beautiful, verifiable architecture, workflow, sequence, data-flow, and lifecycle diagrams—self-contained HTML with motion and crisp export.

58.7k3.8kJavaScript
Skillsagent-skillsarchitecture-as-codeInstall
K-Dense-AI
scientific-agent-skills
today

Turn any AI agent into an AI Scientist. The #1 Agent Skills library for science, used by 190,000+ scientists worldwide. 165 ready-to-use validated skills plus 100+ scientific databases covering biology, chemistry, medicine, and drug discovery. Compatible with Cursor, Claude Code, Codex, Pi, Antigravity, and the open Agent Skills standard.

44.5k4kPython
Skillsagent-skillsai-scientistInstall
VoltAgent
awesome-agent-skills
4d ago

A curated collection of 1000+ agent skills from official dev teams and the community, compatible with Claude Code, Codex, Gemini CLI, Cursor, and more.

34.1k3.6k
Skillsagent-skillsai-agentsInstall