Semantic archaeology for Git. Find the history that explains the code.
- ✓Open-source license (Apache-2.0)
- ✓Actively maintained (<30d)
- ✓Clear description
- ✓Topics declared
- ✓Documented (README)
claude mcp add git-why -- npx -y @alliecatowo/git-why{
"mcpServers": {
"git-why": {
"command": "npx",
"args": ["-y", "@alliecatowo/git-why"]
}
}
}Resumen de MCP Servers
# Git Why
[](LICENSE)
[](https://nodejs.org/)
[](https://github.com/alliecatowo/git-why/actions/workflows/ci.yml)
**`git blame` tells you who changed the code. `git why` finds the history
that explains it.**
Local-first semantic + full-text search over Git history. It returns the
actual commit, its author's real words, and the relevant diff. It retrieves
evidence; it does not generate an explanation of its own.
Real output from `mise run demo`, which rebuilds a small fixture repository
from `bench/fixtures/demo/build.mjs` and runs these two queries against it.
Nothing here is hand-written, and you can reproduce it in one command.
For output against **real repositories** — curl, zod, redis — including a
worked case where this is the wrong tool, see
[`docs/examples.md`](docs/examples.md) or the
[recorded terminal sessions](https://alliecatowo.github.io/git-why/guide/examples).
```text
$ git why "that bizarre bug where reconnecting subscribed twice"
1. 7deab42 Stop duplicate subscriptions after reconnect
2025-11-20 · Maya Chen
Reconnecting re-ran the subscribe handler without clearing the previous
registration, so every reconnect doubled the delivered events.
src/net/socket.ts
-export function connect(url) { return new Socket(url); }
+export function connect(url) {
+ const s = new Socket(url);
+ s.on('reconnect', () => resubscribeOnce(s));
+ return s;
+}
```
```text
$ git why "why do we keep the session when the refresh token is empty?"
1. f17db20 Fix infinite token-refresh loop
2025-11-03 · Maya Chen
Provider X can return an empty refresh token while the current access
token remains valid. Retrying here puts clients into an infinite loop.
src/auth/refresh.ts
export function refresh(session, refreshToken) {
- if (!refreshToken) throw new InvalidTokenError();
+ if (!refreshToken) return session;
return exchange(refreshToken);
}
```
## Install
Install from npm (Node >= 22.12):
```sh
npm install -g @alliecatowo/git-why
```
Or with Homebrew (macOS and Linux; pulls in `node`):
```sh
brew install alliecatowo/tap/git-why
```
Or build from source:
```sh
git clone https://github.com/alliecatowo/git-why.git
cd git-why
npm ci
npm run build
npm install -g .
```
The package installs a `git-why` executable, which Git dispatches as the
subcommand `git why`. No alias setup needed. More install paths — pinned
versions, a GitHub release tarball, building from source, uninstalling —
are in [`docs/install.md`](docs/install.md).
## First run
```sh
cd your-repo
git why "why do we retry twice before giving up"
```
The first query downloads a small embedding model (about 32 MB, checksum-verified,
cached under `~/.cache/git-why/models`) and indexes the repository's history into
`.git/why/`; later queries are incremental. Behind a proxy or offline, see
`HTTPS_PROXY`, `GIT_WHY_MODEL_BASE_URL` (mirror) and `GIT_WHY_OFFLINE` in
[`docs/operations.md`](docs/operations.md). `git why status` shows whether the
index is current.
## Use it with an agent
Two plugins ship in `plugins/`:
| plugin | what you get |
| -------------- | -------------------------------------------------------------------------------------------------------------------------------------- |
| `git-why` | The MCP server, a routing skill, and an index-management skill. |
| `git-why-full` | The same, plus [`zg`](https://zvec.org) for current-code search and a history-explorer agent for questions that need several searches. |
OpenCode users get `opencode/` — an `opencode.json` with the MCP registration
and an `AGENTS.md` fragment.
What the skill actually teaches is **when not to reach for history**. It tells
an agent to use `git log -S` when it can name the symbol, because that is a
case this tool measurably loses (Hit@10 0.950 against 0.350), and to use code
search rather than history for the current state of the code. A skill that
claims its own tool is always best makes an agent worse at its job.
Any other MCP client can run the server directly: `npx -y @alliecatowo/git-why mcp`
(or `git why mcp` once installed). It is also listed in the
[MCP registry](https://registry.modelcontextprotocol.io) as
`io.github.alliecatowo/git-why`.
See [`docs/plugin.md`](docs/plugin.md).
## Make it faster (optional)
```sh
git why server on
```
Holds the index, the embedding model and the lineage table open between
queries. On curl — 30,000 commits — that is 569 ms a query down to 286 ms.
Searches use it automatically once it is running.
It can never be the reason a search fails: if no daemon is running, or it is
unreachable, or your index moved under it, the query runs directly instead.
`--daemon=direct` opts out, `--daemon=server` requires one. See
[`docs/daemon.md`](docs/daemon.md).
## Why
`git log --grep` only matches words you already know. The reason code
changed usually lives in a commit message, but you rarely remember its exact
wording — you remember the _problem_, in your own words, months later.
Look again at the first example above: the query is "reconnecting subscribed
twice" and the commit is titled "Stop duplicate subscriptions after
reconnect." They share almost no vocabulary. That's not a coincidence Git
Why is showing off — it's the actual point of a semantic index. It carries
the meaning of the change, not just its words, alongside an ordinary keyword
index for when you _do_ know the exact term.
Git Why retrieves historical evidence: the commit, the author's actual
words, and the relevant diff. Historical commit messages are assertions by
their authors, not infallible accounts of intent, and some reasons were
never committed at all — Git Why will not manufacture those.
## Measured, not claimed
Every number below is written into this file by `bench/report.mjs` from raw
run data, and CI fails if it drifts — no figure here was typed by hand. Full
methodology, limits, and the negative results are in
[`docs/report.md`](docs/report.md).
### Against the tools you would otherwise use
<!-- generated:corpus-table -->
174 **recall** questions — you remember a problem but cannot name anything in the
commit that fixed it — derived mechanically from 6 pinned real repositories
(curl, redis, requests, ripgrep, caddy, zod). Every question is verified **unanswerable by
keyword search** before it enters the set: if `git log --grep` or `git log -S` finds the
answer from the question's own words, the case is discarded.
| strategy | Hit@1 | Hit@5 | MRR | returned nothing |
| -------------------------- | --------- | --------- | --------- | ---------------- |
| **git why** | **0.201** | **0.374** | **0.266** | 0 |
| zg (semantic code search) | 0.017 | 0.040 | 0.026 | 2 |
| git log -G | 0.006 | 0.017 | 0.011 | 15 |
| git log --grep | 0.000 | 0.011 | 0.003 | 0 |
| git log --grep --all-match | 0.000 | 0.000 | 0.000 | **109** |
| git log -S | 0.000 | 0.000 | 0.000 | 15 |
**10.2x `zg` and 23x the best Git-native strategy** — and the only approach that answers nearly every question rather than returning an empty set.
<!-- /generated:corpus-table -->
### Where it loses
<!-- generated:crossfile -->
When you can name the symbol, use pickaxe search instead. On cross-file causal
questions, `git log -S` scores Hit@10 **0.950** against `git why`'s 0.350.
Semantic search has no advantage over a tool you can hand the exact literal.
<!-- /generated:crossfile -->
That boundary is the honest positioning, and the shipped
[skill](plugins/git-why/skills/history-archaeology/SKILL.md) tells agents both halves:
- **cannot name the term** → `git why`
- **can name the term** → `git log -S`
- **current code, not history** → `zg`
### Scale and cost
### Does it help an agent?
<!-- generated:agent -->
Retrieval quality is not the product. The question is whether an agent answering a real
question does it more accurately, or in fewer turns, with the tool than without. Four arms
over the same frozen tasks, paired per task, with token counts reconciled against the
provider's own accounting database.
| model | paired n | accuracy W-L | median tool calls saved |
| --------------------- | -------: | -----------: | ----------------------: |
| gemini-3.1-flash-lite | 6 | 1-2 | 2.5 |
| claude-haiku-4-5 | 9 | 1-1 | 1 |
| claude-sonnet-5 | 8 | 0-0 | 1 |
| deepseek-v4-flash | 8 | 0-0 | 3.5 |
| gemini-2.5-flash-lite | 6 | 2-1 | 1 |
| gemini-3.1-flash-lite | 7 | 2-0 | 1 |
| gemini-3.5-flash | 6 | 0-0 | 0.5 more |
Results are mixed across models. At single-digit paired n per model this is descriptive, not significant, and it is reported that way
deliberately — the direction is consistent, the magnitude is not established. Full method and per-arm figures in the benchmark report.
<!-- /generated:agent -->
Method, per-arm figures and the registered hypothesis: [`docs/report.md`](docs/report.md#5-agent-benchmark-does-this-help-an-agent-and-which-agents).
<!-- generated:scale -->
| measurement | result Lo que la gente pregunta sobre git-why
¿Qué es alliecatowo/git-why?
+
alliecatowo/git-why es mcp servers para el ecosistema de Claude AI. Semantic archaeology for Git. Find the history that explains the code. Tiene 0 estrellas en GitHub y su última actualización registrada es del 2026-10-05.
¿Cómo se instala git-why?
+
Puedes instalar git-why clonando el repositorio (https://github.com/alliecatowo/git-why) o siguiendo las instrucciones del README en GitHub. ClaudeWave también te ofrece bloques de instalación rápida en esta misma página.
¿Es seguro usar alliecatowo/git-why?
+
Nuestro agente de seguridad ha analizado alliecatowo/git-why y le ha asignado un Trust Score de 95/100 (tier: Verified). Revisa el desglose completo de comprobaciones superadas y flags en esta página.
¿Quién mantiene alliecatowo/git-why?
+
alliecatowo/git-why es mantenido por alliecatowo. La última actividad registrada en GitHub es del 2026-10-05, con 0 issues abiertos.
¿Hay alternativas a git-why?
+
Sí. En ClaudeWave puedes explorar mcp servers similares en /categories/mcp, ordenados por popularidad o actividad reciente.
Despliega git-why en tu cloud
Lleva este repo a producción en minutos. Cada plataforma genera su propio entorno con variables de entorno editables.
¿Mantienes este repo? Añade un badge a tu README
Pega el badge en tu README de GitHub para mostrar que está auditado por ClaudeWave. Cada badge enlaza de vuelta a esta página y muestra el Trust Score actual.
[](https://claudewave.com/repo/alliecatowo-git-why)<a href="https://claudewave.com/repo/alliecatowo-git-why"><img src="https://claudewave.com/api/badge/alliecatowo-git-why" alt="Featured on ClaudeWave: alliecatowo/git-why" width="320" height="64" /></a>Más MCP Servers
Fair-code workflow automation platform with native AI capabilities. Combine visual building with custom code, self-host or cloud, 400+ integrations.
User-friendly AI Interface (Supports Ollama, OpenAI API, ...)
An open-source AI agent that brings the power of Gemini directly into your terminal.
Real-time global intelligence dashboard. AI-powered news aggregation, geopolitical monitoring, and infrastructure tracking in a unified situational awareness interface
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
The fastest path to AI-powered full stack observability, even for lean teams.