Real-time, provenance-invalidated cognitive cache for AI agents and RAG — build understanding once, reuse it everywhere, keep it fresh.
- ✓Open-source license (Apache-2.0)
- ✓Actively maintained (<30d)
- ✓Clear description
- ✓Documented (README)
- !README contains suspicious pattern: eval\s*\(
git clone https://github.com/Vectorlink-Labs/coalent && cp coalent/*.md ~/.claude/agents/Subagents overview
<p align="center">
<img src="https://raw.githubusercontent.com/Vectorlink-Labs/coalent/main/brand/wordmark.png" alt="Coalent" width="320">
</p>
<p align="center">
<b>Real-time, provenance-invalidated context for AI agents & RAG.</b><br>
<i>Build understanding once. Reuse it everywhere. Keep it fresh — automatically.</i>
</p>
<p align="center">
<img alt="pypi" src="https://img.shields.io/pypi/v/coalent?color=5145E5">
<img alt="python" src="https://img.shields.io/badge/python-3.10%2B-4F46E5">
<img alt="license" src="https://img.shields.io/badge/license-Apache%202.0-22D3EE">
<img alt="typed" src="https://img.shields.io/badge/mypy-strict-2DD4BF">
<img alt="tests" src="https://img.shields.io/badge/tests-passing-10B981">
<a href="https://discord.gg/v3hvg3nwr"><img alt="discord" src="https://img.shields.io/badge/Discord-join-5865F2?logo=discord&logoColor=white"></a>
</p>
<p align="center">
<b>📖 <a href="https://coalent.ai/docs">Documentation</a></b> · <a href="https://coalent.ai">coalent.ai</a> · <a href="https://discord.gg/v3hvg3nwr">💬 Discord</a>
</p>
<p align="center">
<a href="#quickstart">Quickstart</a> ·
<a href="#whats-new-in-v06">What's new in v0.6</a> ·
<a href="#the-read-path--a-ladder-of-gates">Gate ladder</a> ·
<a href="#bring-your-own-stack">Bring your own stack</a> ·
<a href="#use-it-from-claude-code--cursor-mcp">MCP</a> ·
<a href="#langchain">LangChain</a> ·
<a href="#benchmark">Benchmark</a> ·
<a href="#cli">CLI</a>
</p>
---
> **Your agent re-reads the same sources on every call — and the moment a source changes, every cached answer is silently wrong.**
>
> Coalent builds the *understanding* once, caches it by what the query **means**, and invalidates it **surgically** the instant an underlying source changes. As correct as re-reading everything, at a fraction of the cost — and never stale.
## Why Coalent
Every context layer is forced to trade off three things. Coalent is built to hold all three at once:
- 🧠 **Extractive understanding, not chunks.** It caches a *query-independent* set of atomic, source-grounded **claims** your LLM extracted — keeping every number and fact — so one cached unit answers many *different* later questions. The raw evidence is retained with each unit, so a hit that under-covers a query falls back to retrieval instead of answering thin.
- ♻️ **Reuse across queries, agents — and documents.** A semantic cache keyed by query *meaning*: ask again, or from another agent, and it's a warm hit. **Cross-unit recall** pools claims across units to answer **multi-hop** questions whose evidence spans documents — at **zero extra LLM calls**.
- 🌿 **Fresh by provenance.** Every unit remembers the exact sources it used. When one changes, only the units that actually used it go stale — precisely, automatically, and lazily.
Coalent sits **above retrieval** — bring any retriever (vector DB, hybrid search, GraphRAG, tools, APIs). It's the freshness-and-reuse layer, not another retriever — deliberately the *opposite* of GraphRAG's build-the-whole-graph-upfront tax: **lightweight, independent units, built lazily only when a query actually needs one**, and refreshed by dirtying a single unit (no graph surgery).
> **New in v0.6** — the **pool read path** (`read_path="pool"`): every read serves the token-budgeted, globally ranked fresh-claim pool. Measured on a 605-question news benchmark (strict grading): **0.731 accuracy @ 981 context tokens** — matching naive top-9 (0.711 @ 1,311) at **~25% fewer tokens**, and naive's best measured point (top-12: 0.731 @ 1,729) at **~43% fewer**. Plus a default-OFF **behavioral stack** — residual spans → refusal fallback → append-only repair → query keys — measured at **−33% refusals** and **+3.1 pts** on the same store. All opt-in; the default read path is unchanged v0.5 behavior. See [What's new](#whats-new-in-v06).
>
> **New in v0.6.1** — the **MCP server**: `coalent-mcp` puts the cache one line away from Claude Code, Cursor, or any MCP client ([Use it from Claude Code / Cursor](#use-it-from-claude-code--cursor-mcp)), and **[`langchain-coalent`](#langchain)** makes your existing LangChain stack the cache's substrate. Both additive-only.
## Install
```bash
pip install coalent # the core has zero required dependencies
```
## Quickstart
Runs as-is — `StubSynthesizer` needs no API key, so you can feel the loop in ten seconds:
```python
from coalent import SemanticCache, InMemoryRetriever, StubSynthesizer
# 1. Any retriever — a vector DB, a tool, an API. (In-memory here for the demo.)
retriever = InMemoryRetriever()
retriever.add("confluence:hr", "Leave policy: 21 days of annual leave per year.")
# 2. Build the cache. Swap StubSynthesizer for a real LLM below.
cache = SemanticCache(retriever, StubSynthesizer())
# 3. Ask. The first call builds understanding and caches it; the next is a warm hit.
result = cache.get("what is our leave policy?")
print(result.context["understanding"])
print(result.cache_hit) # False (cold) -> True on the next call
# 4. A source changed? Only the units that used it go stale — surgically.
cache.source_changed("confluence:hr", text="Leave policy: now 25 days.")
# the next matching read rebuilds just that one unit, lazily
```
Wire in a real model — any text-in / text-out LLM works. In v0.4 the synthesizer builds **extractive** understanding by default (query-independent atomic claims that keep every fact), and the cache does **cross-unit recall** — both on automatically:
```python
from coalent import SemanticCache, LLMSynthesizer, OpenAIProvider, OpenAIEmbedder
cache = SemanticCache(
retriever,
LLMSynthesizer(OpenAIProvider(), model="gpt-4o-mini"), # extract=True by default (v0.4)
embedder=OpenAIEmbedder(), # match queries by MEANING (recommended for real use)
)
# Multi-hop across documents? recall is already on; raise its trigger to bridge units:
# SemanticCache(retriever, synth, embedder=..., recall_threshold=0.7)
```
**The v0.6 pool read path** — opt in, and every read serves the budget-packed, globally ranked fresh-claim pool instead of one routed unit. Attribution is the one thing to wire: a 3-line `pool_header` callable mapping each unit to `[title | source | date]` from your own corpus metadata. This is the measured golden path — on a 605-question news benchmark (strict grading), **0.68** accuracy with the bare built-in header vs **0.73** with this callable, same store, same queries:
```python
DOC_META = { # your corpus metadata, keyed by artifact id
"docs:azure-refresh": {"title": "Azure region refresh", "source": "CloudWire", "date": "2026-05-02"},
}
def pool_header(unit) -> str: # the [title | source | date] golden path — 3 lines
meta = DOC_META.get(unit.evidence[0].artifact_id if unit.evidence else "")
return f"[{meta['title']} | {meta['source']} | {meta['date']}]" if meta else f"[source: {unit.id}]"
cache = SemanticCache(retriever, synthesizer, embedder=OpenAIEmbedder(),
read_path="pool", pool_header=pool_header)
result = cache.get("which regions got the refresh?")
result.context["pool"] # the packed, attributed claim payload — hand it to your answer model
```
Runnable no-API-key demo, including the refusal loop: [examples/pool_read_path.py](examples/pool_read_path.py).
## Use it from Claude Code / Cursor (MCP)
<!-- mcp-name: io.github.nisarg-pujara-vectorlink/coalent -->
`coalent-mcp` serves fresh, attributed facts from a Coalent cache to any MCP client —
and the facts are invalidated the instant their source changes. One line to wire it into
Claude Code:
```bash
pip install "coalent[mcp,openai]"
claude mcp add coalent -- coalent-mcp --cache-factory my_cache:build
```
(Cursor / Claude Desktop / any MCP client: register the same `coalent-mcp ...` command in
its MCP config.)
**Bring your own cache (`--cache-factory module:function`) — the primary mode.** Your
factory returns a fully constructed `SemanticCache`: your vector DB, your embedder, your
LLM, every knob. The server adds protocol glue only — and the glue is measured to add
**zero quality loss**: factory mode reproduced the library's own benchmark result
byte-identically (0.710 on a 100-question validation run drawn from our n=605 news
benchmark — identical CIs, 100/100 serves, 98/100 answer payloads byte-equal to the
library run).
```python
# my_cache.py — importable from the directory you launch in
from coalent import (SemanticCache, LLMSynthesizer, OpenAIProvider,
OpenAIEmbedder, SQLiteCognitionStore)
def build() -> SemanticCache:
return SemanticCache(
my_vector_retriever, # YOUR vector DB / retriever
LLMSynthesizer(OpenAIProvider()), # YOUR synthesis model
embedder=OpenAIEmbedder(), # YOUR embedder
read_path="pool",
residual_spans=True, query_keys=True, # the behavioral stack, opt-in as ever
pool_header=my_metadata_header, # [title | source | date] — the measured golden path
store=SQLiteCognitionStore("kb.db"), # persistence is yours too
)
```
Freshness here is signal-driven: your ingestion pipeline calls the `source_changed` tool
when a document changes and the affected facts invalidate immediately. (Adding
`--watch DIR` alongside the factory also fires it on file edits — invalidation only; it
never ingests into your index, and it matches only when your artifact ids equal the
watch-relative paths.)
**Zero-config folder mode (`--watch DIR`) — the demo wedge.** Point it at a folder of
docs and you get the recommended v0.6 deployment (pool path, residual spans, query keys,
SQLite persistence, automatic `[path | modified date]` attribution) with no code at all:
```bash
claude mcp add coalent --env OPENAI_API_KEY=$OPENAI_API_KEY -- coalent-mcp --watch ./docs
```
Every read rescans the watched files (mtime + content hash) before serving — you cannot
get a stale answer after sWhat people ask about coalent
What is Vectorlink-Labs/coalent?
+
Vectorlink-Labs/coalent is subagents for the Claude AI ecosystem. Real-time, provenance-invalidated cognitive cache for AI agents and RAG — build understanding once, reuse it everywhere, keep it fresh. It has 9 GitHub stars and its last recorded update is dated 2026-08-05.
How do I install coalent?
+
You can install coalent by cloning the repository (https://github.com/Vectorlink-Labs/coalent) or following the README instructions on GitHub. ClaudeWave also provides quick install blocks on this page.
Is Vectorlink-Labs/coalent safe to use?
+
Our security agent has analyzed Vectorlink-Labs/coalent and assigned a Trust Score of 77/100 (tier: Trusted). See the full breakdown of passed checks and flags on this page.
Who maintains Vectorlink-Labs/coalent?
+
Vectorlink-Labs/coalent is maintained by Vectorlink-Labs. The last recorded GitHub activity is dated 2026-08-05, with 0 open issues.
Are there alternatives to coalent?
+
Yes. On ClaudeWave you can browse similar subagents at /categories/agents, sorted by popularity or recent activity.
Deploy coalent to your cloud
Ship this repo to production in minutes. Each platform spins up its own environment with editable env vars.
Maintain this repo? Add a badge to your README
Drop the badge into your GitHub README to show it's tracked on ClaudeWave. Each badge links back to this page and reflects the live Trust Score.
[](https://claudewave.com/repo/vectorlink-labs-coalent)<a href="https://claudewave.com/repo/vectorlink-labs-coalent"><img src="https://claudewave.com/api/badge/vectorlink-labs-coalent" alt="Featured on ClaudeWave: Vectorlink-Labs/coalent" width="320" height="64" /></a>More Subagents
The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.
The agent that grows with you
Java 面试 & 后端通用面试指南,覆盖计算机基础、数据库、分布式、高并发、系统设计与 AI 应用开发
Build Agentic workflows, RAG pipelines, with rich AI model and tool support on one collaborative workspace. Deploy on cloud, VPC, or self-hosted, so teams move from prototype to production without rebuilding the stack.
The agent engineering platform.
Turn any codebase, with its docs, SQL schemas, configs, and PDFs, into a queryable knowledge graph. A /graphify skill for Claude Code, Cursor, Codex, and Gemini CLI: local deterministic AST parsing, every edge explained, no vector store.