Memory Layer — Master Architecture
Persistent, cross-channel, auto-training memory for the BoostEcom AI Teams. Status: this is the May 2026 design plan, and the three phases below all SHIPPED (the header said "implementation in three phases" until…
Persistent, cross-channel, auto-training memory for the BoostEcom AI Teams. Status: this is the May 2026 design plan, and the three phases below all SHIPPED (the header said "implementation in three phases" until 2026-09-05, long after they were in the tree). Keep this page for decisions D1 to D10 and their rationale; for what runs today, module by module, read
src/features/ai/memory/README.md, which is the living doc.
North star
A merchant should never have to re-explain context. After three conversations, @Atlas knows the brand voice, the catalog footprint, the operator's preferences and tone, the partner stack, the recurring constraints, and uses that knowledge on every channel (web chat, WhatsApp, future Slack / SMS) for years. The relationship deepens like a human agency relationship that compounds over time.
What already exists (DO NOT duplicate)
The Explore agent surfaced infrastructure we keep:
| Surface | Path | Role |
|---|---|---|
Conversation + Message Prisma models | prisma/schema.prisma — model Conversation / model Message | Cross-channel persistence already keyed on (externalId, channel) with orgId/storeId/userId joins. Cascade delete. Indexes ready. |
StoreContext Json blob | prisma/schema.prisma — model StoreContext | Per-store singleton with knowledge, todos, pipeline, agentState, conversation, archivedConversations. Multi-agent handoff trail lives here. |
| Upstash Vector + OpenAI embeddings | src/features/ai/memory/rag.ts, client in vector-index.ts | Knowledge RAG with text-embedding-3-small (1536d). Store-scoped namespace, 800-char chunks. Exposed to the model as knowledgeSearch — this row said searchKnowledge until 2026-09-06, a name nothing answers to, which is the same transposition that put a phantom tool on all six agent profiles (ai-platform-tools-29). rag.ts and embed.ts each declared their own client, singleton and embedding constant against the same index until ai-platform-tools-19; there is now one, in vector-index.ts. |
MemorySystem class | gone | This row named src/lib/memory/index.ts until 2026-09-05. The directory was removed by ai-platform/0056 and no MemorySystem class exists anywhere; the memory code lives in src/features/ai/memory/. |
| 3-layer prompt composition | src/features/ai/prompts/layer-system.ts | Layer 1 (base) + Layer 2 (@Atlas kernel, src/features/ai/prompts/kernel.ts) + Layer 3 (Store.systemPrompt + Store.instructions + Store.knowledgeSources). composeAtlasPrompt({ kernel, ctx }) is the canonical assembler. This row said "iRen kernel" and pointed at src/services/algorithms/prompts/ until ai-platform/0603 and 0599: iRen is the orchestrator's pre-2026 name, and the composer now lives in the pillar that owns the prompt (ADR 0022). |
Store.systemPrompt + Store.instructions | prisma/schema.prisma | Per-store overrides already injected into Layer 3. |
| Layer detection | src/features/ai/context/layer.ts | detectLayer() returns "guest" | "connected" | "full_agent" and gates tool access. This row said "same file" as the composer above; it never was. |
| Cron infra | src/app/api/cron/* + vercel.json | Dozens of jobs already running (intelligence-tick, reset-credits, weekly-digest, ...); vercel.json is the count, and this row said 10 until 2026-09-05. Daily cadence available. |
Implication: we extend the three-layer prompt and the StoreContext blob with structured facts rather than introducing a parallel memory store. We reuse the Upstash Vector pipeline for vector retrieval. We add one cron for memory consolidation (no infra needed).
Gaps to fill (this project's scope)
- No structured facts:
StoreContext.knowledgeandStoreContext.agentStateare freeform Json with no schema, no scoring, no idempotency, no audit trail. Reading "what does @Atlas know about this org" is a manual JSON dig. - No cross-conversation memory: messages persist but the orchestrator only loads the current conversation's last N turns. A 3-month-old insight from a different conversation is invisible.
- No cross-channel surfacing — WhatsApp and web each load their own conversation thread. Same org, same user, same store → two parallel siloed memories.
- No vector retrieval over messages: Upstash Vector only indexes knowledge sources, not conversation messages.
- No auto-training loop: user corrections in chat don't strengthen anything. Same mistake can recur.
- No memory UI: no
/[orgSlug]/~/memorypage. RGPD requires it. - System prompt has no facts injection point:
composeAtlasPromptdoesn't load memory; only static layers.
Architecture decisions
D1. Three scope-specific fact tables, not a polymorphic table
OrgFact, StoreFact, UserFact: separate Prisma models. Reasons:
- Native
onDelete: CascadetoOrganization/Store/User. - Indexes tunable per scope (org has thousands of facts, user has dozens).
- Type-safe at the call site (
OrgFact[]vsUserFact[]). - Cost of duplication: ~15 lines per model. Acceptable.
We do NOT extend StoreContext.knowledge with a typed shape because: (a) it's already in production use with freeform Json, mutating its shape risks runtime crashes, (b) a typed table gives us indexes and queries the JSON path can't.
D2. Append-only with confirmation count, never overwrite
From mem0's lessons + Anthropic's memory tool pattern. When a new fact contradicts an old one:
- Write the new row with
confirmationCount = 1. - Decay the old fact's score (see D5).
- After
≥ 3confirmations of the new fact, soft-archive the old (archivedAtnot null). - Surface contradictions to the memory UI for human review.
This preserves an audit trail and avoids destructive corrections on a single noisy turn.
D3. Per-agent namespace = (orgId | storeId | userId, agentId, scope, key)
Adds agentId to the namespace so @Maya (Marketing), @Marco (Merchandising), @Otis (Operations), @Faye (Intelligence), @Sam (Support) each have their own memory pocket. agentId = "atlas" for the orchestrator. Lets @Maya remember "the merchant prefers UGC creatives" without polluting @Marco's catalogue recall. (The roles are the runtime's, specialists.ts; this line carried the retired profile set until ai-platform/3016.)
For Phase 1 we ship with agentId = "atlas" only and add the column NOW (cheap to add, expensive to backfill later).
D4. Two-stage writes — synchronous extraction, asynchronous consolidation
Inline (request path):
- After
streamTextfinishes, call Haiku via Gateway to extract 0..5 candidate facts from the last turn pair. The call isgenerateText+Output.object(), so the Gateway sendsresponseFormat: jsonand a non-JSON answer is not representable. It used to be a plaingenerateTextwhose text was de-fenced with two regexes andJSON.parsed, and a badly closed fence dropped every fact of the turn under one warning (ai-platform-tools-21). - Persist raw to the fact table with
confirmationCount = 1. No conflict resolution yet.
Async (daily cron consolidate-memory):
- Group facts by
(scope+id, agentId, key). - Apply decay formula (D5) → compute composite score.
- Merge duplicates (cosine similarity > 0.92 on value embeddings: Phase 2).
- Increment
confirmationCountfor merged rows. - Archive low-score facts (composite < 0.2).
- Write
MemoryEventaudit rows for every action.
Mem0's production failures (write-time token cost, indexing inconsistency) all trace to one-stage pipelines. We bake the split in from day one.
D5. Decay + TTL tiers
score = baseConfidence × decayFactor
decayFactor = 0.5 ^ (ageDays / halfLifeDays)
TTL tier defaults (overrideable per fact via the extractor):
| Tag | half-life (days) | hard TTL (days) |
|---|---|---|
legal, allergy, accessibility | ∞ (no decay) | ∞ |
brand, target-market, catalog | 180 | none |
preference, tone, language | 90 | none |
season, campaign, current-focus | 30 | 180 |
casual | 14 | 60 |
The retrieval scoring uses the composite score; the consolidation cron archives facts below 0.2. Hard TTL = full delete after N days, RGPD-friendly.
D6. Prompt-cache stability — memory injected AFTER the static prefix
Anthropic prompt caching (5-min TTL, $0.30/M vs $3/M = 10× lever) requires byte-stable prefixes. The cached prefix is base + atlas-kernel. Memory injection is a NEW Layer 4 block appended AFTER the cached prefix:
[Layer 1 base + Layer 2 kernel] ← cached, stable
[Layer 3 store overlay] ← cached, stable
[Layer 4 memory facts] ← non-cached, mutates per request
[Channel context] ← non-cached
[Conversation history] ← non-cached
If we ever pollute Layer 1-3 with mutable memory, the cache invalidates every turn and cost balloons 10×.
D7. Vector retrieval reuses Upstash Vector, not pgvector
Upstash Vector is already wired, with OpenAI text-embedding-3-small, store-scoped namespaces, and 1536-dim vectors. We extend the same pipeline to index fact values and message contents. We do NOT introduce pgvector to Neon: that would split the vector surface across two stores and double the operational complexity.
Phase 2 adds two new namespaces:
facts:${scopeId}— fact values embedded for cross-fact dedup + semantic memory lookup.messages:${storeId}— message embeddings for "find earlier conversations about X".
D8. Layer / Letta-inspired tiered memory
Three tiers loaded at request time, each with its own budget:
| Tier | Source | Budget | Refresh |
|---|---|---|---|
| Core | High-confidence stable facts (composite > 0.7, tag in brand/target-market/tone/language) | ~300 tokens | Every request |
| Working | Current conversation summary (last 12 turns OR Haiku rollup if > 12) | ~500 tokens | Every request |
| Archival | Top-k semantic retrieval over facts + message embeddings | ~600 tokens | Every request, Phase 2 |
Total ~1400 tokens of memory per request. Combined with the static 3-layer prompt (~2500 tokens), every request stays well under Anthropic's input budget while keeping the prefix cacheable.
D9. Memory UI in Phase 1, not Phase 3
RGPD Article 17 (right to be forgotten) requires per-fact visibility + delete on user demand. Shipping the UI in Phase 3 leaves us legally exposed for months. Ship the page in Phase 1 alongside the schema.
D10. Opt-out per request via #noremember
A trailing #noremember in a user message disables extraction for that turn pair. Cheap to implement (string check before scheduling extraction), useful for sensitive topics. Documented in the WhatsApp / web chat help.
D11. La veille : ce qu'on apprend chez les autres, jamais pris pour nous
Un audit de concurrent ou de marque d'inspiration sert a ameliorer la
boutique connectee, donc ses conclusions meritent d'etre retenues. Elles
ne doivent jamais l'etre COMME la boutique : « avis sous le prix » vu sur
competitor.com est une piste pour nous, pas un fait sur notre theme.
Pas de table dediee. Un fait de veille est un fait ordinaire (scope
store en pratique, puisque c'est elle qu'on ameliore) qui porte les deux
marqueurs a la fois : la key watch.<site>.<sujet> et le tag veille.
src/features/ai/memory/watch.ts les fait concorder sur chaque brouillon
extrait (l'un sans l'autre est complete, ou le fait est jete s'il ne tient
plus dans 80 caracteres) ; format.ts les rend dans une section
« Veille » a part, donc aucune ligne sous « Store » ne decrit le site d'un
autre. Demi-vie : 60 jours, une page concurrente bouge plus vite qu'une
marque.
Three-phase plan (revised, all three shipped)
The week estimates below are what was planned in May 2026, not work remaining. Where each phase landed:
Plan Shipped as Phase 1 extraction / write / retrieval / formatting src/features/ai/memory/extract.ts,write.ts,retrieve.ts,format.tsPhase 1 UI + cron src/app/(dashboard)/[orgSlug]/~/memory/page.tsx, cronconsolidate-memoryinvercel.jsonPhase 2 vector retrieval src/features/ai/memory/embed.ts(Upstash Vector namespaces)Phase 3 corrections + rules src/features/ai/memory/corrections.ts,src/app/(dashboard)/[orgSlug]/~/settings/memory/rules/page.tsx
Phase 1 — Structured facts + injection + UI (2-3 weeks)
- Prisma :
OrgFact,StoreFact,UserFact,MemoryEvent, schema in D1 with namespace per D3. - Idempotency:
@@unique([scope+id, agentId, key, sourceMessageId]). src/features/ai/memory/extract.ts— Haiku extractor per D4 stage 1.src/features/ai/memory/write.ts— idempotent writer withMemoryEventaudit.src/features/ai/memory/retrieve.ts— tag + confidence + recency filter, ~1000 token budget, < 50ms p95.src/features/ai/memory/format.ts— renders the Layer 4 memory block.- Integration:
composeAtlasPromptgains a Layer 4 slot. Both web chat handler and WhatsApp bot read memory before prompt assembly, schedule extraction after stream finish. - Cron
consolidate-memory(no-op for Phase 1 dedup, just decay + audit + archive). - Page
/[orgSlug]/~/memory— list facts grouped by scope, edit / delete / export. RGPD-compliant. - Opt-out
#norememberkeyword.
Phase 2 — Vector retrieval + cross-channel memory (1-2 weeks)
- Extend Upstash Vector with
facts:${scopeId}andmessages:${storeId}namespaces. - Embed every persisted fact value + every assistant/user message (just-in-time on first retrieval miss, not eager).
getMemoryContextgains an "archival" tier per D8: cosine similarity > 0.7 against the current user turn, top-k 5 facts + top-k 5 messages.- Consolidation cron upgraded: cosine > 0.92 → merge facts, increment
confirmationCount. - Cross-channel: WhatsApp inbound (org/store linked via
WhatsappChannel) loads the same memory as the user's web chat.
Phase 3 — Auto-training + per-agent memory + skill rules (2 weeks)
- Per-agent
agentIdpopulated for the 5 specialists. Routing viaatlas.routewrites to the right pocket. Correctiondetection (heuristic: "no", "actually", "non plutot" within first 60s after an assistant turn), auto-creates a high-confidence fact withsource.reason = "user-correction"..boostecomrulesper-org file pattern (Cursor-style):/[orgSlug]/~/settings/memory/rules, Markdown editable rules that always inject into Layer 3.- A/B harness: per-org flag
memory.enabled = true|false(defaulttrue). Measure response quality (thumbs / regenerate rate) with and without.
Scale-up plan (deferred)
The pipeline is wired with raw void promise.catch(...) (WhatsApp
webhook) and next/server's after() (web Route Handler). This
holds up to roughly 1 000 turn pairs / day on Vercel's default
function runtime grace period.
| Trigger | Migration target |
|---|---|
| > 1 k msg/day OR p95 extract latency > 4 s | Upstash QStash queue — push the turn pair to a queue inside extractAndPersistFacts, drain with a /api/jobs/memory-extract Route Handler |
| > 10 k msg/day across multiple regions | Vercel Queues (when GA) for built-in retries + dead-letter routing |
Cross-region read traffic > 50 RPS on getMemoryContext | Add Upstash Redis cache layer in front of the core tier query (TTL ~60 s) |
Migration is non-disruptive: extractAndPersistFacts already
swallows errors, so the queue can run async without changing the
caller contract. The MemoryEvent audit table makes replay safe
(every action is idempotent on (scopeId, agentId, key, sourceMessageId)).
We don't pre-build the queue today because:
- Current volume is < 100 turns/day across all orgs.
- QStash adds a per-message billable line item ($0.0005/msg → $15/mo at 1k/day; acceptable but unjustified at < 100/day).
after()+void promisealready give us the lifecycle guarantees we need at our scale.
Cost model
Haiku extraction per turn pair: ~200 input + ~100 output tokens at $1/M + $5/M = ~$0.0007 / turn.
| Volume | Worst case | With batching (every 2 turns) | With noisy turn skip (~40%) |
|---|---|---|---|
| 1k orgs × 1k turns/mo | $700/mo | $350/mo | $210/mo |
| 10k orgs × 1k turns/mo | $7k/mo | $3.5k/mo | $2.1k/mo |
Embedding cost (Phase 2): facts only, ~50 tokens each, $0.02/M = $0.000001/fact. Negligible.
Storage cost (Neon): Fact rows are ~500 bytes each. 1M facts = 500 MB. Negligible inside our existing Neon plan.
RGPD posture + tenant security (Phase 1 + hardening)
Cross-tenant isolation
- Every fact-table query is gated by a tenant scope (
orgId/storeId/userId) at the API + retrieve layer. Enforced in CI byscripts/audit-memory-scope.sh(see Hardening below), a PR that adds aprisma.orgFact.findManywithout awhere.orgIdfails the run. - Upstash Vector namespaces are scope-keyed:
facts:org:<orgId>/facts:store:<storeId>/facts:user:<userId>/messages:<storeId>. Cross-namespace queries are impossible at the client API level. - The scope id used by
persistFactscomes from the server-side context (the session'sorgId/storeId/userId), NEVER from the model's output. A malicious user can't make the extractor write to another org by including a tag in their message. getOrgAccess(userId, orgId)is called on every memory API route; member role required to read, admin+ to mutate, owner to reset.
Prompt-injection resistance
- The Zod
FactDraftSchemaconstrainskeyto dotted-lowercase (regex/^[a-z0-9_.]+$/),tagsto ≤ 8 strings ≤ 40 chars each,confidenceto[0,1],reasoningto ≤ 240 chars. Anything outside the schema is dropped silently — fact by fact, which is why theOutput.object()envelope (ExtractionEnvelopeSchema) is deliberately looser thanFactDraftSchema: an output specification is validated as a whole, so putting thekeyregex there would make one malformed fact discard the entire turn. The envelope pins the shape,FactDraftSchemapins the contents. extract.tsrejects facts whose serialised value exceeds 4 KB (defense in depth against payload smuggling).- The extraction prompt is read-only and treats user content as data: the model is told to extract claims, not execute instructions. Even if the model is jailbroken, the worst-case output is a low-quality fact that lands in the right org's pocket and stays subject to RGPD delete + audit trail.
RGPD posture
onDelete: Cascadeon all fact tables to parent Org / Store / User. Deleting the parent wipes their memory atomically.- Per-user delete: account deletion from
/account/settings(POST /api/auth/delete-account) drops theUserrow, andUserFact.useris declaredonDelete: Cascade, so every fact goes with it. There is no memory-only per-user wipe. - Org-scope reset in
/[orgSlug]/~/memory→ wipes allOrgFact+StoreFactfor the org's stores (owner-only). - Per-fact delete (Article 17) via the memory dashboard. Cleans up
Upstash Vector via
deleteFactFromVector. MemoryEventrows do not cascade with their fact, and this page said they did until 2026-09. The model declares no relation at all:factIdandscopeIdare bare strings, denormalized for filtering, so a fact delete leaves its events,oldValue/newValuepayloads included. What actually removes them is the account-delete route, which callsdeleteManyonscopeIdfor the user plus every org and store being removed. Read as written, this line said no explicit cleanup was needed anywhere.src/test/erasure-leaves-no-pii.test.tsnow fails if the model gains no relation and the route stops naming it.- Export endpoint (
GET /api/organizations/[orgId]/memory/export) returns a JSON dump for Article 20 (data portability).
Hardening
scripts/audit-memory-scope.sh runs in CI via
.github/workflows/audit-memory-scope.yml. It grep-walks the
source tree, flags any prisma.<factTable>.<readOrMutate>( call
that does not include (orgId|storeId|userId|scopeId|id:) in the
following 12 lines. Two exempt files (the consolidation cron and
the export endpoint) operate by primary key after a scoped lookup
and are explicitly allowlisted.
Open questions (decide before coding)
agentIdcolumn type — string ("atlas", "maya", ...) vs enum. String wins for forward-compat; we add a Zod parser at the call site.- Confidence floor for storage: drop facts below
0.4, or persist them flagged? Drop is cheaper; persist enables learning over time. Recommend drop in Phase 1. #noremembergranularity — disables extraction for this turn only, or for the whole conversation? Recommend per-turn (less surprising).- WhatsApp memory recall visibility : say "Je me souviens que tu m'as dit X" or stay invisible? Recommend invisible (trust-safer).
- Memory UI default state: show ALL facts even low-confidence, or only
> 0.5? Recommend> 0.5with a "show all" toggle.
Anti-patterns (from research, never do)
- Eager embedding every message. Just-in-time only — saves ~10× cost.
- One-stage extract-and-upsert. Corrupts on contradiction. Always async-consolidate.
- Memory injected into the cached prompt prefix. Kills the 10× cache discount.
- Per-message memory without a token budget. Cap at ~1400 tokens / turn.
- Memory in Upstash Redis as the source of truth. Volatile + no joins. Keep Redis for working-memory hot cache only.
- No user UI. RGPD Article 17 violation.
References
- mem0 production lessons (token optimization playbook 2026).
- Letta hierarchical memory (core / working / archival).
- Anthropic prompt caching docs (5-min TTL, cache key design).
- Anthropic memory tool docs (append-only, confirmation count pattern).
- ChatGPT Memory + memory toast pattern (transparency UX).
- Cursor
.cursorrulespattern (per-project Markdown rules: inspires.boostecomrules). - FadeMem selective forgetting research (decay formulas).