Proves software work: triages failing tests to a cause with evidence, says whether a test could ever fail, breaks your code to see if the tests notice, reads what a migration does to a live database, and gives one verdict on whether a change is done.
- ✓Open-source license (Apache-2.0)
- ✓Actively maintained (<30d)
- ✓Clear description
- ✓Topics declared
- ✓Documented (README)
claude mcp add tewip -- npx -y playwright{
"mcpServers": {
"tewip": {
"command": "npx",
"args": ["-y", "playwright"]
}
}
}MCP Servers overview
# Tewip
**Tewip** (Uyghur *téwip*, تېۋىپ: healer, doctor) is the layer that **proves software work** - written by a person or by an agent.
It finds why each failing test failed (a real bug, a flaky test, the environment, test data, or a changed selector), says whether a test could ever fail at all, breaks your code to find out whether the tests would notice, reads what a migration will do to a live database, and gives one verdict on whether a change is done. Every answer carries its evidence and names what decided it: a deterministic rule, the AI model, or a person's correction.
What makes it worth installing is what it refuses to say. It answers *"cannot say"* rather than guessing, it will not narrow a test run it cannot prove safe to narrow, and every report ends with what it did **not** check. Each number below was measured on strangers' real code, never on fixtures.
It works for every stack: Node.js (node:test, Jest, Vitest, Mocha), Python (pytest, unittest), Java and Kotlin (JUnit, TestNG), Go, .NET, Ruby, PHP and Rust, each checked against its real runner's report ([the full list](QUICKSTART.md#not-using-playwright-every-stack-works)), plus Playwright in depth with traces, selector drift and verified fixes. It also reads the failure wording of the browser and component-test tools on top of those runners (Selenium, WebdriverIO, Cypress, Testing Library for React, Vue and Angular, Vue Test Utils, jest-dom and Jasmine), checked against more than 500 real messages from public projects.
| Where | How |
|---|---|
| **Terminal** | `npm install -g tewip-ai`, then `tewip explain report.xml`, `tewip ci owner/repo --pr 12`, `tewip mcp` |
| **VS Code, Cursor, Windsurf, VSCodium** | the Tewip extension ([download](https://github.com/hmamut39/tewip/releases/tag/vscode-0.5.0)): registers Tewip with the editor's AI assistant and adds six commands - explain a report, why CI failed on this branch, the risk review of your change, which flaky tests to quarantine, the security check, and what your corrections add up to |
| **JetBrains, Claude, other AI assistants** | add the MCP server `tewip mcp`: fifteen tools, including *has this happened here before?*, *is this requirement covered and passing?* and *what should be done next about this goal?* (see [editor setup](#use-it-from-your-editor-any-ide-that-speaks-mcp)) |
| **Any model** | OpenAI, Claude, Gemini, Azure, or a model on your own machine with Ollama (no account, nothing leaves the computer): `TEWIP_PROVIDER` |
| **GitHub CI** | the Tewip Action: a comment on every pull request, an optional merge gate, verified fix PRs ([QUICKSTART](QUICKSTART.md)) |
Formerly *triage-agent*: the repository moved to `hmamut39/tewip`, and GitHub redirects the old address. See [ROADMAP.md](ROADMAP.md) for the plan.
**Status: phases 1 to 14 are built; the two auto-fix phases (5 and 8) are the ones still in progress.** Tewip collects failures from any framework, diagnoses them with rules first and a model only for what rules cannot settle, and turns that into work for each role: a CI comment and merge gate, test plans and verified tests, a flake dossier, a PR-risk review, security scanning, tickets in Jira, Rally, GitHub Issues, Azure Boards or Linear, checks generated from a Figma design, specifications written for people, and - opt in, and only when proved - fixes to test or application code. It runs live in GitHub CI and answers questions from AI assistants through MCP. On fresh real Gutenberg CI failures, with the pipeline frozen before labelling, accuracy is 93.4%; across six real datasets it ranges from 86.1% to 99.7%, and the caveats are in [eval/README.md](eval/README.md).
## What it costs
Most failures never reach a model. On 3,556 real CI failures from six public projects, the
deterministic rules answered **75.1%** on their own - and on the labelled ones they answered
correctly **1,979 times out of 2,012 (98.4%)**. A rule-answered failure costs nothing, explains
itself, and gives the same answer every time.
Only the remaining quarter is sent to a model, so triaging **1,000 real failures costs about
$1.12** at gpt-5.4-mini's published price. That is measured, not estimated: three real uncached
calls on held-out failures billed 3,624 prompt and 499 completion tokens each. A real week costs
less, because a failure that repeats is answered from cache, and a failure someone has corrected
never goes to the model again.
Tewip prints what every command costs, caps its own spend, and works with no key at all - the rules
run on their own and everything they cannot settle is left as UNKNOWN rather than guessed.
## Use it in your project
**Giving it to your team (or another company): [TEAM-GUIDE.md](TEAM-GUIDE.md)** — what to install, what each role uses it for, and what it never does.
**Putting it on a real repository: [QUICKSTART.md](QUICKSTART.md)** — ten minutes, copy-paste, and what to check on the first failing run.
### 1. In GitHub CI (developers, SDETs, DevOps)
Tell Playwright to write a blob report in CI. It's a built-in reporter, so there's no triage-agent code in your config:
```ts
// playwright.config.ts
reporter: process.env.CI ? [['list'], ['blob']] : 'list',
use: { trace: 'retain-on-failure' },
```
Then add one step after your tests:
```yaml
permissions:
contents: write # only needed for autofix
pull-requests: write
issues: write
actions: read # lets it read the failed job's log when no test report was written
steps:
# ... checkout, install, npx playwright test ...
- name: triage-agent
if: always()
uses: hmamut39/tewip@main
with:
openai-api-key: ${{ secrets.OPENAI_API_KEY }} # optional: without it only the rules run
autofix: 'true' # optional: verified locator-fix PRs
```
On every run, the action:
- **Developers:** posts one diagnosis comment on the pull request and keeps it updated. For each failure it says whether the change likely caused it, why, and what to do next.
- **SDETs:** writes the triage board into the job summary.
- **DevOps:** collects environment failures (server errors, connection resets, outages) in **one** GitHub issue, and each new run adds a single grouped comment to it.
- **The tests in the pull request:** on a PR it reads the test files the change touched and says which
of them could never fail - an unfinished expectation, a value compared with itself, a test that
discards its own failure, one that never runs. It says nothing when they are all sound, and it says
it **even on a green run**, because "all tests passed" is the exact sentence that needs qualifying
when one of them cannot fail. Needs `fetch-depth: 0` on `actions/checkout` so the diff can be read;
without it the section is skipped rather than guessed at. Any such test also enters the attestation
as a refusal, so a release decision taken from the bundle cannot rest on a test that proves nothing.
- **What the change would slip past (`prove: true`):** breaks the source files the pull request touched
in small deliberate ways, runs the tests that cover them, and reports every change none of them
noticed - "these 2 changes to your code would have slipped past your tests", with the lines. Off by
default because each change costs a real test run; `prove-budget` (default 300s) caps the whole
thing, and it says which files it did not reach rather than implying they were clean. Every file is
put back exactly as it was. Each unnoticed change also enters the attestation as a refusal: a line
nothing checks is not a line this bundle can claim was verified.
- **Evidence, every run:** the Action writes an attestation into its artifact - the claims with
their evidence and which layer decided each, what Tewip **refused** to decide, and what it did
not check at all. Nothing enters it that Tewip did not verify itself, so a release decision can
be taken from it without opening CI.
- **Fixes (`autofix: true`):** for drifted locators that the deterministic rule matched to a renamed element, it rewrites the locator, re-runs that test, and opens a separate pull request with the changes that passed. It never merges and never edits application code.
- **Every role:** uploads all role views as the `tewip-report` artifact.
- **When the job dies before any test runs** (an image that would not pull, a runner out of disk, a service container that never came up): it reads the failed job's log and names the cause, instead of leaving a red X with no explanation.
- **History:** keeps the run history in the Actions cache, never in git.
- **Corrections:** reply to the agent's comment with `/triage CATEGORY #n reason` (e.g. `/triage TIMING #2 flaky on CI`). It stores that with the history and, next time that test fails the same way, gives the person's answer instead of its own, naming them and their reason. A corrected failure also costs nothing: it never goes to the model again.
For **senior engineers, architects and leads**, add a scheduled workflow with `mode: report`, for example weekly. It builds the engineering-quality, architecture and risk-digest views from the stored CI history, reusing earlier diagnoses so it makes no repeat AI calls. The views are kept in one GitHub issue that each report updates:
```yaml
on:
schedule: [{ cron: '0 6 * * 1' }]
workflow_dispatch:
permissions: { contents: read, issues: write }
jobs:
report:
runs-on: ubuntu-latest
steps:
- uses: hmamut39/tewip@main
with: { mode: report }
```
The report reads the history cache of the branch it runs on, usually the default branch.
The repository needs "Allow GitHub Actions to create and approve pull requests" enabled for fix PRs. Pull requests opened by the workflow token don't trigger CI themselves; the agent verified each change by re-running the test before opening the PR. It was first proved end to end on a private practice repository, which is why there is no link to click here. If this reWhat people ask about tewip
What is hmamut39/tewip?
+
hmamut39/tewip is mcp servers for the Claude AI ecosystem. Proves software work: triages failing tests to a cause with evidence, says whether a test could ever fail, breaks your code to see if the tests notice, reads what a migration does to a live database, and gives one verdict on whether a change is done. It has 0 GitHub stars and its last recorded update is dated 2026-10-10.
How do I install tewip?
+
You can install tewip by cloning the repository (https://github.com/hmamut39/tewip) or following the README instructions on GitHub. ClaudeWave also provides quick install blocks on this page.
Is hmamut39/tewip safe to use?
+
Our security agent has analyzed hmamut39/tewip and assigned a Trust Score of 95/100 (tier: Verified). See the full breakdown of passed checks and flags on this page.
Who maintains hmamut39/tewip?
+
hmamut39/tewip is maintained by hmamut39. The last recorded GitHub activity is dated 2026-10-10, with 2 open issues.
Are there alternatives to tewip?
+
Yes. On ClaudeWave you can browse similar mcp servers at /categories/mcp, sorted by popularity or recent activity.
Deploy tewip to your cloud
Ship this repo to production in minutes. Each platform spins up its own environment with editable env vars.
Maintain this repo? Add a badge to your README
Drop the badge into your GitHub README to show it's tracked on ClaudeWave. Each badge links back to this page and reflects the live Trust Score.
[](https://claudewave.com/repo/hmamut39-tewip)<a href="https://claudewave.com/repo/hmamut39-tewip"><img src="https://claudewave.com/api/badge/hmamut39-tewip" alt="Featured on ClaudeWave: hmamut39/tewip" width="320" height="64" /></a>More MCP Servers
Fair-code workflow automation platform with native AI capabilities. Combine visual building with custom code, self-host or cloud, 400+ integrations.
User-friendly AI Interface (Supports Ollama, OpenAI API, ...)
An open-source AI agent that brings the power of Gemini directly into your terminal.
Real-time global intelligence dashboard. AI-powered news aggregation, geopolitical monitoring, and infrastructure tracking in a unified situational awareness interface
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
The fastest path to AI-powered full stack observability, even for lean teams.