Adaptive memory for AI agents & teams — beyond RAG. Self-hosted MCP server that gets smarter every time you search: hybrid search + a neural memory graph that learns. Works with Claude, ChatGPT & any MCP client.
- ✓Open-source license (Apache-2.0)
- ✓Actively maintained (<30d)
- ✓Clear description
- ✓Topics declared
- ✓Documented (README)
claude mcp add memory-cloud -- uvx memory-cloud{
"mcpServers": {
"memory-cloud": {
"command": "uvx",
"args": ["memory-cloud"]
}
}
}MCP Servers overview
<p align="center">
<a href="https://www.kagura-ai.com">
<picture>
<source media="(prefers-color-scheme: dark)" srcset="docs/assets/social-preview.png">
<img src="docs/assets/readme-banner.png" alt="Kagura Memory Cloud — adaptive memory for AI agents and teams, beyond RAG" width="820">
</picture>
</a>
</p>
<p align="center">
English · <a href="README.ja.md">日本語</a>
</p>
<p align="center">
<strong>Adaptive memory for AI agents and teams</strong> — self-hosted, beyond RAG.<br>
An MCP server that gets smarter every time you search:<br>
hybrid search + a neural memory graph that learns which memories belong together.
</p>
<p align="center">
<a href="LICENSE"><img src="https://img.shields.io/badge/License-Apache_2.0-blue.svg" alt="License"></a>
<a href="https://github.com/kagura-ai/memory-cloud/actions/workflows/ci.yml"><img src="https://github.com/kagura-ai/memory-cloud/actions/workflows/ci.yml/badge.svg" alt="CI"></a>
<a href="https://codecov.io/gh/kagura-ai/memory-cloud"><img src="https://codecov.io/gh/kagura-ai/memory-cloud/graph/badge.svg" alt="codecov"></a>
<a href="https://www.python.org/downloads/"><img src="https://img.shields.io/badge/python-3.11+-blue.svg" alt="Python 3.11+"></a>
<a href="https://nodejs.org/"><img src="https://img.shields.io/badge/node.js-20+-green.svg" alt="Node.js 20+"></a>
<a href="https://modelcontextprotocol.io/"><img src="https://img.shields.io/badge/MCP-Streamable_HTTP-purple.svg" alt="MCP"></a>
<a href="https://safeskill.dev/scan/kagura-ai-memory-cloud"><img src="https://img.shields.io/badge/SafeSkill-90%2F100_Verified%20Safe-brightgreen" alt="SafeSkill 90/100"></a>
</p>
<p align="center">
Works with Claude, ChatGPT, Gemini, and any MCP-compatible client.<br>
<a href="https://github.com/kagura-ai/kagura-memory-python-sdk"><strong>Python SDK (KaguraClient, REST clients & FileIngestor)</strong></a>
</p>
<p align="center">
<a href="https://www.kagura-ai.com/demo/terminal-en-cli-2x.mp4">
<img src="docs/assets/cli-demo.gif" alt="Claude Code CLI recalling from Kagura Memory over MCP" width="760">
</a>
<br>
<em>Claude Code CLI recalling memories from Kagura over MCP — <a href="https://www.kagura-ai.com/demo/terminal-en-cli-2x.mp4">▶ watch the demo</a></em>
</p>
## Why Kagura Memory Cloud?
> **Your AI forgets everything after each conversation. Kagura fixes that — and gets smarter every time you search.**
Most AI memory tools are just vector databases with a chat wrapper. Kagura is different — it implements the full **LLM Knowledge Base** pattern (Karpathy's [LLM Wiki](https://gist.github.com/karpathy/442a6bf555914893e9891c11519de94f)) at team scale:
| Approach | Storage | Compounding | Scale |
|---|---|---|---|
| Vector DB / RAG | Embedded chunks | None — retrieve-only | Any |
| Karpathy's LLM Wiki | Markdown files | LLM rewrites pages | Personal (~100 pages) |
| **Kagura Memory Cloud** | **PostgreSQL + Qdrant + Neural graph** | **Hebbian + Sleep Maintenance** | **Team / org** |
| Feature | Description |
|---------|-------------|
| **Adaptive Memory** | Every search automatically strengthens connections between related memories. The more you use it, the better `explore()` discovers hidden relationships. |
| **Hybrid Search** | Semantic (OpenAI / self-hosted) + BM25 keyword — 96% top-1 accuracy |
| **AI Reranking** | Self-hosted (Ollama/vLLM — local, free), Voyage AI, or Cohere — cross-encoder reranking for precision |
| **Neural Memory Graph** | Hebbian learning builds a knowledge graph in the background. `explore()` traverses it for serendipitous discovery. |
| **Agent Memory Substrate** | Beyond a knowledge store: delivery modes (pinned / time-triggered), a server-stamped trust boundary, an agent state lane, and a retrieval-feedback signal — the primitives an autonomous agent loop needs. |
| **Agent Control Plane (preview)** | Workspace-scoped Agent Registry, subtractive context bindings, agent-bound member keys, lifecycle kill switches, and one-call session bootstrap. Introduced in v0.49.0. |
| **64 MCP Tools** | Memory, Agent Substrate, Agent Control Plane, Neural edges, Contexts, Tags, Files (R2), Analyses (Memory Analysis), Resources, Secrets, Sleep Maintenance, Usage, API-Key Bindings |
| **Multi-Provider** | OpenAI or self-hosted (Ollama, vLLM — local, private, zero cost) for embeddings |
| **Team Ready** | Workspaces, RBAC, context isolation, shared memory |
| **Web UI** | Next.js dashboard — contexts, search settings, member management |
| **5-Minute Setup** | `./setup.sh` and you're done |
## Architecture
```
Workspace (team/org)
├── Context A ("my-project") ← like a folder
│ ├── Memory 1 ← 3-layer: summary / context / content
│ ├── Memory 2
│ └── Neural edges (Hebbian) ← automatic connections
├── Context B ("learning-notes")
│ └── ...
└── Members (Owner/Admin/Member/Viewer)
```
### LLM Knowledge Base — 5-Layer Implementation
Karpathy's [LLM Wiki pattern](https://gist.github.com/karpathy/442a6bf555914893e9891c11519de94f) describes a 5-layer "living knowledge base" — beyond traditional RAG. Kagura implements all 5 layers at team scale:
| Layer | Kagura Implementation | Difference from Karpathy's pattern |
|---|---|---|
| **Ingest** | REST `/api/v1/memory`, MCP `remember`, R2 file storage, resource tokens | + binary blobs, + multi-tenant |
| **Compile** | **MCP-as-compile-API** — chat agent compiles via structured tool calls (`remember(summary, content, type, tags)`) + Sleep Maintenance for batch consolidation | Continuous micro-compile (not batch wiki rewrite) — schema-enforced output |
| **Index** | Triple index: **BM25** (keyword) + **Qdrant** (semantic) + **Hebbian graph** (relational) — all auto-maintained | No manual `index.md` upkeep |
| **Query** | Hybrid Search + AI Reranker + `explore` graph traversal | Beyond markdown grep — supports semantic + relational queries |
| **Enhance** | **Hebbian learning** — every `recall()` strengthens edges between co-retrieved memories. Sleep Maintenance consolidates periodically. | Background graph evolution (zero LLM cost) vs LLM-driven page rewrites |
**Compounding loop**: Currently explicit (user/agent calls `remember()` after synthesizing answers). Auto-write-back of synthesized answers is intentionally opt-in to keep noise low.
### Adaptive Memory: Two Search Paths
Kagura separates **precision search** and **discovery** into two independent paths, each optimized for its purpose:
```
recall() ──→ Hybrid Search (semantic + BM25) ──→ [Reranker] ──→ Precise results
│
└──→ Hebbian Learning (background) ──→ Graph edges grow
│
explore() ──→ Graph Traversal (Neural Memory) ←─────────────────┘ Related discoveries
```
- **`recall()`** — Precision search. Hybrid (semantic 60% + BM25 40%) with optional AI reranking. Returns the most relevant memories.
- **`explore()`** — Discovery. Traverses the Neural Memory graph to find related memories that keyword search would miss.
- **Hebbian learning** — Every `recall()` silently strengthens edges between co-retrieved memories. No explicit training needed — the graph grows organically as you use the system.
This separation is intentional: mixing graph signals into recall degrades precision ([validated via benchmarks](docs/neural-memory-evaluation.md)). Instead, each path does what it's best at.
**Data isolation:** All data is filtered by `workspace_id → context_id → user_id`. Memories never leak across boundaries. Single Qdrant collection with payload filtering.
**Tech stack:** FastAPI (async) · PostgreSQL · Qdrant · Redis · Next.js 16 · OAuth2 · MCP over Streamable HTTP
**Vector backend:** Qdrant by default. A single-process self-hosted / CLI / edge deployment can instead run the embedded **LanceDB backend — "Kagura Lite" (preview)** with no separate Qdrant server (`KAGURA_VECTOR_BACKEND=lance`, `cd backend && uv sync --locked --extra lite`). Not for multi-worker / SaaS (LanceDB is single-writer). See [Deployment → Embedded Vector Backend](docs/deployment.md#embedded-vector-backend-kagura-lite-preview).
## Quick Start
### System Requirements
| | Minimum | Recommended |
|--|---------|-------------|
| CPU | 2 cores | 4+ cores |
| RAM | 4 GB | 8+ GB |
| Disk | 10 GB free | 20+ GB free |
### Prerequisites
- Docker & Docker Compose
- Python 3.11+
- Node.js 20+
- OpenAI API key (for embeddings) — or a self-hosted inference server (e.g. Ollama) for local embeddings
- OAuth2 credentials (optional — password + MFA login available without OAuth)
### Setup
**One-line setup:**
```bash
git clone https://github.com/kagura-ai/memory-cloud.git
cd memory-cloud
./setup.sh
```
**With Claude Code:**
```bash
git clone https://github.com/kagura-ai/memory-cloud.git
cd memory-cloud
claude # then run /setup
```
**Step-by-step setup:**
```bash
# 1. Clone
git clone https://github.com/kagura-ai/memory-cloud.git
cd memory-cloud
# 2. Configure environment (generates secrets, prompts for API keys)
(cd backend && python3 -m src.cli.setup_env)
# 3. Start all services
docker compose up -d
# 4. Run migrations
(cd backend && alembic upgrade head)
# 5. Create admin account (interactive — sets password, MFA, API key, embedding provider)
(cd backend && python3 -m src.cli.create_admin)
# Backend API: http://localhost:8080
# Frontend UI: http://localhost:3000
# API docs: http://localhost:8080/redoc
```
**`.env.local` settings** (auto-configured by `setup_env`):
| Setting | Required | Description |
|---------|----------|-------------|
| `API_KEY_SECRET` | **Yes** | Secret for API key encryption (auto-generated) |
| `JWT_SECRET` | **Yes** | Secret for JWT tokens (auto-generated) |
| `OPENAI_API_KEY` | **Yes**\* | OpenAI API key for embeddings |
| `SELF_HOSTED_BASE_URL` | No | Self-hosted backend URL (default: `http://localhost:11434`) |
| `EMBEDDING_PROVIDER` | No | `openai` (default) or `What people ask about memory-cloud
What is kagura-ai/memory-cloud?
+
kagura-ai/memory-cloud is mcp servers for the Claude AI ecosystem. Adaptive memory for AI agents & teams — beyond RAG. Self-hosted MCP server that gets smarter every time you search: hybrid search + a neural memory graph that learns. Works with Claude, ChatGPT & any MCP client. It has 12 GitHub stars and its last recorded update is dated 2026-09-27.
How do I install memory-cloud?
+
You can install memory-cloud by cloning the repository (https://github.com/kagura-ai/memory-cloud) or following the README instructions on GitHub. ClaudeWave also provides quick install blocks on this page.
Is kagura-ai/memory-cloud safe to use?
+
Our security agent has analyzed kagura-ai/memory-cloud and assigned a Trust Score of 95/100 (tier: Verified). See the full breakdown of passed checks and flags on this page.
Who maintains kagura-ai/memory-cloud?
+
kagura-ai/memory-cloud is maintained by kagura-ai. The last recorded GitHub activity is dated 2026-09-27, with 11 open issues.
Are there alternatives to memory-cloud?
+
Yes. On ClaudeWave you can browse similar mcp servers at /categories/mcp, sorted by popularity or recent activity.
Deploy memory-cloud to your cloud
Ship this repo to production in minutes. Each platform spins up its own environment with editable env vars.
Maintain this repo? Add a badge to your README
Drop the badge into your GitHub README to show it's tracked on ClaudeWave. Each badge links back to this page and reflects the live Trust Score.
[](https://claudewave.com/repo/kagura-ai-memory-cloud)<a href="https://claudewave.com/repo/kagura-ai-memory-cloud"><img src="https://claudewave.com/api/badge/kagura-ai-memory-cloud" alt="Featured on ClaudeWave: kagura-ai/memory-cloud" width="320" height="64" /></a>More MCP Servers
Fair-code workflow automation platform with native AI capabilities. Combine visual building with custom code, self-host or cloud, 400+ integrations.
User-friendly AI Interface (Supports Ollama, OpenAI API, ...)
An open-source AI agent that brings the power of Gemini directly into your terminal.
Real-time global intelligence dashboard. AI-powered news aggregation, geopolitical monitoring, and infrastructure tracking in a unified situational awareness interface
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
The fastest path to AI-powered full stack observability, even for lean teams.