Skip to main content
ClaudeWave

Proves software work: triages failing tests to a cause with evidence, says whether a test could ever fail, breaks your code to see if the tests notice, reads what a migration does to a live database, and gives one verdict on whether a change is done.

MCP ServersOfficial Registry0 stars0 forks● TypeScriptApache-2.0Updated today
ClaudeWave Trust Score
95/100
✓ Verified
Passed
  • ✓Open-source license (Apache-2.0)
  • ✓Actively maintained (<30d)
  • ✓Clear description
  • ✓Topics declared
  • ✓Documented (README)
Last scanned: 10/11/2026
Install in Claude Code / Claude Desktop
Method: NPX · playwright
Claude Code CLI
claude mcp add tewip -- npx -y playwright
claude_desktop_config.json (Claude Desktop)
{
  "mcpServers": {
    "tewip": {
      "command": "npx",
      "args": ["-y", "playwright"]
    }
  }
}
1. Run the command above in your terminal (Claude Code), or paste the JSON config into claude_desktop_config.json (Claude Desktop).
2. Replace any <placeholder> values with your API keys or paths.
3. Restart Claude. The MCP server and its tools appear automatically.
Use cases

MCP Servers overview

# Tewip

**Tewip** (Uyghur *téwip*, تېۋىپ: healer, doctor) is the layer that **proves software work** - written by a person or by an agent.

It finds why each failing test failed (a real bug, a flaky test, the environment, test data, or a changed selector), says whether a test could ever fail at all, breaks your code to find out whether the tests would notice, reads what a migration will do to a live database, and gives one verdict on whether a change is done. Every answer carries its evidence and names what decided it: a deterministic rule, the AI model, or a person's correction.

What makes it worth installing is what it refuses to say. It answers *"cannot say"* rather than guessing, it will not narrow a test run it cannot prove safe to narrow, and every report ends with what it did **not** check. Each number below was measured on strangers' real code, never on fixtures.

It works for every stack: Node.js (node:test, Jest, Vitest, Mocha), Python (pytest, unittest), Java and Kotlin (JUnit, TestNG), Go, .NET, Ruby, PHP and Rust, each checked against its real runner's report ([the full list](QUICKSTART.md#not-using-playwright-every-stack-works)), plus Playwright in depth with traces, selector drift and verified fixes. It also reads the failure wording of the browser and component-test tools on top of those runners (Selenium, WebdriverIO, Cypress, Testing Library for React, Vue and Angular, Vue Test Utils, jest-dom and Jasmine), checked against more than 500 real messages from public projects.

| Where | How |
|---|---|
| **Terminal** | `npm install -g tewip-ai`, then `tewip explain report.xml`, `tewip ci owner/repo --pr 12`, `tewip mcp` |
| **VS Code, Cursor, Windsurf, VSCodium** | the Tewip extension ([download](https://github.com/hmamut39/tewip/releases/tag/vscode-0.5.0)): registers Tewip with the editor's AI assistant and adds six commands - explain a report, why CI failed on this branch, the risk review of your change, which flaky tests to quarantine, the security check, and what your corrections add up to |
| **JetBrains, Claude, other AI assistants** | add the MCP server `tewip mcp`: fifteen tools, including *has this happened here before?*, *is this requirement covered and passing?* and *what should be done next about this goal?* (see [editor setup](#use-it-from-your-editor-any-ide-that-speaks-mcp)) |
| **Any model** | OpenAI, Claude, Gemini, Azure, or a model on your own machine with Ollama (no account, nothing leaves the computer): `TEWIP_PROVIDER` |
| **GitHub CI** | the Tewip Action: a comment on every pull request, an optional merge gate, verified fix PRs ([QUICKSTART](QUICKSTART.md)) |

Formerly *triage-agent*: the repository moved to `hmamut39/tewip`, and GitHub redirects the old address. See [ROADMAP.md](ROADMAP.md) for the plan.

**Status: phases 1 to 14 are built; the two auto-fix phases (5 and 8) are the ones still in progress.** Tewip collects failures from any framework, diagnoses them with rules first and a model only for what rules cannot settle, and turns that into work for each role: a CI comment and merge gate, test plans and verified tests, a flake dossier, a PR-risk review, security scanning, tickets in Jira, Rally, GitHub Issues, Azure Boards or Linear, checks generated from a Figma design, specifications written for people, and - opt in, and only when proved - fixes to test or application code. It runs live in GitHub CI and answers questions from AI assistants through MCP. On fresh real Gutenberg CI failures, with the pipeline frozen before labelling, accuracy is 93.4%; across six real datasets it ranges from 86.1% to 99.7%, and the caveats are in [eval/README.md](eval/README.md).

## What it costs

Most failures never reach a model. On 3,556 real CI failures from six public projects, the
deterministic rules answered **75.1%** on their own - and on the labelled ones they answered
correctly **1,979 times out of 2,012 (98.4%)**. A rule-answered failure costs nothing, explains
itself, and gives the same answer every time.

Only the remaining quarter is sent to a model, so triaging **1,000 real failures costs about
$1.12** at gpt-5.4-mini's published price. That is measured, not estimated: three real uncached
calls on held-out failures billed 3,624 prompt and 499 completion tokens each. A real week costs
less, because a failure that repeats is answered from cache, and a failure someone has corrected
never goes to the model again.

Tewip prints what every command costs, caps its own spend, and works with no key at all - the rules
run on their own and everything they cannot settle is left as UNKNOWN rather than guessed.

## Use it in your project

**Giving it to your team (or another company): [TEAM-GUIDE.md](TEAM-GUIDE.md)** — what to install, what each role uses it for, and what it never does.

**Putting it on a real repository: [QUICKSTART.md](QUICKSTART.md)** — ten minutes, copy-paste, and what to check on the first failing run.

### 1. In GitHub CI (developers, SDETs, DevOps)

Tell Playwright to write a blob report in CI. It's a built-in reporter, so there's no triage-agent code in your config:

```ts
// playwright.config.ts
reporter: process.env.CI ? [['list'], ['blob']] : 'list',
use: { trace: 'retain-on-failure' },
```

Then add one step after your tests:

```yaml
permissions:
  contents: write        # only needed for autofix
  pull-requests: write
  issues: write
  actions: read          # lets it read the failed job's log when no test report was written

steps:
  # ... checkout, install, npx playwright test ...
  - name: triage-agent
    if: always()
    uses: hmamut39/tewip@main
    with:
      openai-api-key: ${{ secrets.OPENAI_API_KEY }}   # optional: without it only the rules run
      autofix: 'true'                                  # optional: verified locator-fix PRs
```

On every run, the action:
- **Developers:** posts one diagnosis comment on the pull request and keeps it updated. For each failure it says whether the change likely caused it, why, and what to do next.
- **SDETs:** writes the triage board into the job summary.
- **DevOps:** collects environment failures (server errors, connection resets, outages) in **one** GitHub issue, and each new run adds a single grouped comment to it.
- **The tests in the pull request:** on a PR it reads the test files the change touched and says which
  of them could never fail - an unfinished expectation, a value compared with itself, a test that
  discards its own failure, one that never runs. It says nothing when they are all sound, and it says
  it **even on a green run**, because "all tests passed" is the exact sentence that needs qualifying
  when one of them cannot fail. Needs `fetch-depth: 0` on `actions/checkout` so the diff can be read;
  without it the section is skipped rather than guessed at. Any such test also enters the attestation
  as a refusal, so a release decision taken from the bundle cannot rest on a test that proves nothing.
- **What the change would slip past (`prove: true`):** breaks the source files the pull request touched
  in small deliberate ways, runs the tests that cover them, and reports every change none of them
  noticed - "these 2 changes to your code would have slipped past your tests", with the lines. Off by
  default because each change costs a real test run; `prove-budget` (default 300s) caps the whole
  thing, and it says which files it did not reach rather than implying they were clean. Every file is
  put back exactly as it was. Each unnoticed change also enters the attestation as a refusal: a line
  nothing checks is not a line this bundle can claim was verified.
- **Evidence, every run:** the Action writes an attestation into its artifact - the claims with
  their evidence and which layer decided each, what Tewip **refused** to decide, and what it did
  not check at all. Nothing enters it that Tewip did not verify itself, so a release decision can
  be taken from it without opening CI.
- **Fixes (`autofix: true`):** for drifted locators that the deterministic rule matched to a renamed element, it rewrites the locator, re-runs that test, and opens a separate pull request with the changes that passed. It never merges and never edits application code.
- **Every role:** uploads all role views as the `tewip-report` artifact.
- **When the job dies before any test runs** (an image that would not pull, a runner out of disk, a service container that never came up): it reads the failed job's log and names the cause, instead of leaving a red X with no explanation.
- **History:** keeps the run history in the Actions cache, never in git.
- **Corrections:** reply to the agent's comment with `/triage CATEGORY #n reason` (e.g. `/triage TIMING #2 flaky on CI`). It stores that with the history and, next time that test fails the same way, gives the person's answer instead of its own, naming them and their reason. A corrected failure also costs nothing: it never goes to the model again.

For **senior engineers, architects and leads**, add a scheduled workflow with `mode: report`, for example weekly. It builds the engineering-quality, architecture and risk-digest views from the stored CI history, reusing earlier diagnoses so it makes no repeat AI calls. The views are kept in one GitHub issue that each report updates:

```yaml
on:
  schedule: [{ cron: '0 6 * * 1' }]
  workflow_dispatch:
permissions: { contents: read, issues: write }
jobs:
  report:
    runs-on: ubuntu-latest
    steps:
      - uses: hmamut39/tewip@main
        with: { mode: report }
```

The report reads the history cache of the branch it runs on, usually the default branch.

The repository needs "Allow GitHub Actions to create and approve pull requests" enabled for fix PRs. Pull requests opened by the workflow token don't trigger CI themselves; the agent verified each change by re-running the test before opening the PR. It was first proved end to end on a private practice repository, which is why there is no link to click here. If this re
ai-agentci-cdcode-reviewdatabase-migrationsdeveloper-toolsflaky-testsjestmcpmcp-servermutation-testingplaywrightpytestqatest-automationtesting

What people ask about tewip

What is hmamut39/tewip?

+

hmamut39/tewip is mcp servers for the Claude AI ecosystem. Proves software work: triages failing tests to a cause with evidence, says whether a test could ever fail, breaks your code to see if the tests notice, reads what a migration does to a live database, and gives one verdict on whether a change is done. It has 0 GitHub stars and its last recorded update is dated 2026-10-10.

How do I install tewip?

+

You can install tewip by cloning the repository (https://github.com/hmamut39/tewip) or following the README instructions on GitHub. ClaudeWave also provides quick install blocks on this page.

Is hmamut39/tewip safe to use?

+

Our security agent has analyzed hmamut39/tewip and assigned a Trust Score of 95/100 (tier: Verified). See the full breakdown of passed checks and flags on this page.

Who maintains hmamut39/tewip?

+

hmamut39/tewip is maintained by hmamut39. The last recorded GitHub activity is dated 2026-10-10, with 2 open issues.

Are there alternatives to tewip?

+

Yes. On ClaudeWave you can browse similar mcp servers at /categories/mcp, sorted by popularity or recent activity.

Deploy tewip to your cloud

Ship this repo to production in minutes. Each platform spins up its own environment with editable env vars.

Maintain this repo? Add a badge to your README

Drop the badge into your GitHub README to show it's tracked on ClaudeWave. Each badge links back to this page and reflects the live Trust Score.

Featured on ClaudeWave: hmamut39/tewip
[![Featured on ClaudeWave](https://claudewave.com/api/badge/hmamut39-tewip)](https://claudewave.com/repo/hmamut39-tewip)
<a href="https://claudewave.com/repo/hmamut39-tewip"><img src="https://claudewave.com/api/badge/hmamut39-tewip" alt="Featured on ClaudeWave: hmamut39/tewip" width="320" height="64" /></a>

More MCP Servers

tewip alternatives