A framework for giving any AI agent persistent, browser-trainable SSM memory — Mamba models, swappable LLM bridges, and online distillation. The hippocampus behind BuilderForce.ai.
- ✓Actively maintained (<30d)
- ✓Clear description
- ✓Documented (README)
- !No standard license detected
git clone https://github.com/SeanHogg/builderforce-memory && cp builderforce-memory/*.md ~/.claude/agents/Subagents overview
# BuilderForce Agent Memory
A framework for giving **any AI agent** persistent, browser-trainable SSM memory. It is provider-neutral — swappable LLM bridges (Anthropic / OpenAI / Fetch), an inference router, online distillation, and a persistent memory store — so any agent or app can depend on it, not just BuilderForce.
**[BuilderForce.ai](https://builderforce.ai) is the flagship deployment**, not the boundary: it trains a custom SSM in the browser, exports a model, and pushes it to an agent runtime where it runs as **Evermind — the model itself**, not just memory bolted onto someone else's LLM. Evermind is the whole brain: its own shared-expert generator (the *cortex* that does reasoning and language), a write-through knowledge memory (the *hippocampus*), and a trainable affective layer (the *limbic* system). A request can be served by Evermind directly — `evermind/<ref>` traffic is generated on-device, not forwarded to Claude/GPT — and external frontier models stay an *optional* routing choice rather than a hard dependency. The same framework is reusable by other agents off the shelf.
This monorepo consolidates two packages that previously lived in separate repos (`mambacode.js` and `ssmjs`) whose names described the *technique* rather than the *role*. They are renamed and unified here so that "what they do" is legible: they are Agent Memory.
## Packages
| Package | Layer | Was | Responsibility |
|---|---|---|---|
| [`@seanhogg/builderforce-memory-engine`](packages/memory-engine) | **Engine** | `@seanhogg/mambacode.js` | WGSL/WebGPU Mamba SSM kernels, model blocks (Mamba1/2/3 + attention), autograd, trainer, BPE tokenizer, quantization. Zero runtime deps. |
| [`@seanhogg/builderforce-memory`](packages/memory) | **Runtime** | `@seanhogg/ssmjs` | SSM execution, Transformer orchestration (Anthropic/OpenAI/Fetch bridges), online distillation, inference router, sessions, the persistent `MemoryStore`, and **Write-Through Cognition** (`EvermindCognition`). Depends on `memory-engine`. |
The two-package split is deliberate: the engine is zero-dep and WebGPU-pure and can be consumed standalone; the runtime pulls in LLM-vendor bridges. Flattening them would force engine-only consumers to drag in vendor code and vice versa. They release in lockstep from one pipeline, which kills the publish-drift bug class that the separate-repo setup suffered (bumping one version without publishing + regenerating the consumer lockfile).
## Cutting token cost
Agent Memory ships three layers that reduce LLM spend, in increasing power. All are **portable** — the same code runs in the browser (WebGPU SSM) and in Node (the agent's `@webgpu/node` SSM) because the embedder and storage are injected, never hard-wired.
| Layer | What it does | Saves |
|---|---|---|
| **Prompt caching** ([`AnthropicBridge`](packages/memory/src/bridges/AnthropicBridge.ts) `cacheSystem`) | Marks the stable system prompt as an Anthropic `cache_control` block | ~90% on the cached **input** prefix (cost, not count) |
| **Exact-match cache** ([`CachingBridge`](packages/memory/src/bridges/CachingBridge.ts) + [`ResponseCache`](packages/memory/src/bridges/ResponseCache.ts)) | Reuses byte-identical completions | Eliminates duplicate calls (retries, identical fan-out) |
| **Semantic cache** ([`SemanticCache`](packages/memory/src/cache/SemanticCache.ts) + [`SemanticCachingBridge`](packages/memory/src/bridges/SemanticCachingBridge.ts)) | Reuses a prior answer when the new prompt is within a cosine **threshold** of one already answered — catches paraphrases | Avoids frontier calls entirely on semantically-repeated prompts |
The semantic cache is the real lever. It is **read-through with two tiers**, mirroring an L1-Map / L2-KV cache:
- **L1** — an in-process vector list, scanned locally with on-device SSM embeddings (free, offline-capable).
- **L2** — an optional shared backend ([`FetchSemanticCacheBackend`](packages/memory/src/cache/FetchSemanticCacheBackend.ts) → the BuilderForce.ai gateway), so a paraphrase answered by the **web app** is reusable by an **agent** and vice-versa.
```ts
import { SemanticCache, FetchSemanticCacheBackend, AnthropicBridge } from '@seanhogg/builderforce-memory';
const cache = new SemanticCache({
embed: (t) => runtime.embed(t), // on-device SSM (free)
l2: new FetchSemanticCacheBackend({ baseUrl, apiKey }), // shared via the gateway
threshold: 0.92,
});
const { response, cached, tier } = await cache.getOrGenerate(
prompt,
() => bridge.generate(prompt), // only runs on a miss
);
```
Memory-backed fact injection is semantic too: [`SSMAgent`](packages/memory/src/agent/SSMAgent.ts) defaults to `factSelection: 'semantic'`, injecting only the top-`maxFacts` embedding-relevant facts each turn (paraphrase-robust, smaller prompts) instead of an exact key-substring match.
## Write-Through Cognition (Evermind)
Caching keeps *answers* fresh; **cognition keeps *knowledge* fresh.** A frozen model — or an append-only memory — drifts: stale and current facts pile up under different keys until something reconciles them by hand. [`EvermindCognition`](packages/memory/src/cognition/EvermindCognition.ts) closes that gap. It is the model-knowledge analogue of a write-through cache with a conflict resolver, so a belief is **replaced on write**, never appended into a reconciliation backlog. This is what lets Evermind's knowledge stay current without a manual reconcile step.
Every candidate fact flows through one pipeline:
> Canonicalize (stable subject key) → recall incumbent → evaluate evidence → reconcile (**augment | confirm | supersede | reject**) → write-through
- **Stable subject key** — facts about the same subject collide and replace, instead of accumulating under per-run ids. This is the anti-drift fix.
- **Evidence-gated** — a *conflicting* fact only supersedes the incumbent when ground-truth evidence favours it. The [`EvidenceGatherer`](packages/memory/src/cognition/types.ts) is injected, so the evidence source is surface-specific (IDE file tools, a DB probe, an HTTP check); `workspacePresenceGatherer` ships for the common "is it still on disk?" case.
- **Write-through recall** — `recall()` is served from a version-token cache that invalidates the instant knowledge changes — the same invalidate-on-write rule as the semantic cache, applied to beliefs.
```ts
import { EvermindCognition, workspacePresenceGatherer, MemoryStore } from '@seanhogg/builderforce-memory';
const cog = new EvermindCognition({ store: new MemoryStore() });
// Seed a (soon-to-be) stale belief.
await cog.commit({ subjectKey: 'pkg:ssm-stack', content: 'SSM stack = MambaKit + SSMjs' });
// Re-observe: evidence from the real workspace decides the conflict.
const r = await cog.commit(
{ subjectKey: 'pkg:ssm-stack', content: 'SSM stack = builderforce-memory monorepo' },
workspacePresenceGatherer({
list: () => listWorkspace(), // e.g. the IDE `list_files` control tool
mustExist: ['builderforce-memory/'],
mustBeAbsent: ['MambaKit/', 'SSMjs/'],
}),
);
// r.verdict === 'supersede' → exactly one belief held, the stale one retired (replace, not append)
```
`EvermindCognition` is store- and surface-agnostic — the `CognitionFactStore` interface is satisfied structurally by `MemoryStore` (no adapter) — so the same loop runs in the IDE, on-prem, cloud, and the browser.
## Enterprise architecture
Caching cuts the bill and cognition keeps knowledge current, but neither answers the questions an enterprise buyer actually asks before signing: *can it read our data, will it leak across tenants, what did that answer cost, and how do we know it got better?* Those four questions are what this layer exists to answer. Each is a port with adapters behind it, so a customer's existing stack is an adapter choice rather than a rewrite.
| Layer | Entry point | Subpath export | Answers |
|---|---|---|---|
| **Ingestion** | `IngestionPipeline` | `@seanhogg/builderforce-memory/ingest` | Can it read our data — structured *and* unstructured — and stay current? |
| **Vector store** | `VectorStore` port + adapters | `…/vectorstore` | Can it run against the database we already have, with our tenancy rules? |
| **Retrieval** | `EnterpriseRetriever` | `…/rag` | Are the answers grounded, scoped, and citable? |
| **Orchestration** | `AgentGraph` + patterns | `…/orchestration` | Can agents coordinate, pause for a human, and resume after a crash? |
| **Telemetry** | `Tracer` + `MetricsRegistry` | `…/telemetry` | What did it cost, how fast was it, and where did it go wrong? |
| **Evaluation** | `EvalHarness` | `…/eval` | How do we know it is good enough to launch — and still is? |
### Ingestion — structured and unstructured through one path
A support-ticket export (rows, typed columns) and a policy document (prose) reach the index through the same pipeline; the difference is which parser ran, not which system was built. Structure that survives parsing becomes **filterable metadata**, which is what makes *"summarise open P1 tickets about billing"* answerable — `status` and `priority` are database predicates rather than something the embedding has to imply.
```ts
import { IngestionPipeline, InMemoryIngestManifest } from '@seanhogg/builderforce-memory/ingest';
import { MemoryVectorStore } from '@seanhogg/builderforce-memory/vectorstore';
const pipeline = new IngestionPipeline({
store : new MemoryVectorStore(),
embed : (texts) => embedder.embedBatch(texts),
manifest: new InMemoryIngestManifest(), // enables `diff` sync
tracer,
});
await pipeline.ingest([
{ id: 'policy.md', tenantId: 'acme', title: 'Retention Policy',
content: { kind: 'text', text: markdown, mediaType: 'text/markdown' } },
{ id: 'tickets', tenantId: 'acme', acl: ['support'],
content: { kind: 'rows', rows: ticketRows, rowIdField: 'id' } },
]);
```
What makes it a pipeline rather than a loader scWhat people ask about builderforce-memory
What is SeanHogg/builderforce-memory?
+
SeanHogg/builderforce-memory is subagents for the Claude AI ecosystem. A framework for giving any AI agent persistent, browser-trainable SSM memory — Mamba models, swappable LLM bridges, and online distillation. The hippocampus behind BuilderForce.ai. It has 2 GitHub stars and its last recorded update is dated 2026-09-27.
How do I install builderforce-memory?
+
You can install builderforce-memory by cloning the repository (https://github.com/SeanHogg/builderforce-memory) or following the README instructions on GitHub. ClaudeWave also provides quick install blocks on this page.
Is SeanHogg/builderforce-memory safe to use?
+
Our security agent has analyzed SeanHogg/builderforce-memory and assigned a Trust Score of 62/100 (tier: OK). See the full breakdown of passed checks and flags on this page.
Who maintains SeanHogg/builderforce-memory?
+
SeanHogg/builderforce-memory is maintained by SeanHogg. The last recorded GitHub activity is dated 2026-09-27, with 0 open issues.
Are there alternatives to builderforce-memory?
+
Yes. On ClaudeWave you can browse similar subagents at /categories/agents, sorted by popularity or recent activity.
Deploy builderforce-memory to your cloud
Ship this repo to production in minutes. Each platform spins up its own environment with editable env vars.
Maintain this repo? Add a badge to your README
Drop the badge into your GitHub README to show it's tracked on ClaudeWave. Each badge links back to this page and reflects the live Trust Score.
[](https://claudewave.com/repo/seanhogg-builderforce-memory)<a href="https://claudewave.com/repo/seanhogg-builderforce-memory"><img src="https://claudewave.com/api/badge/seanhogg-builderforce-memory" alt="Featured on ClaudeWave: SeanHogg/builderforce-memory" width="320" height="64" /></a>More Subagents
The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.
The agent that grows with you
Java 面试 & 后端通用面试指南,覆盖计算机基础、数据库、分布式、高并发、系统设计与 AI 应用开发
Build Agentic workflows, RAG pipelines, with rich AI model and tool support on one collaborative workspace. Deploy on cloud, VPC, or self-hosted, so teams move from prototype to production without rebuilding the stack.
The agent engineering platform.
Makes your AI agent think like the laziest senior dev in the room. The best code is the code you never wrote.