ArchitectureMemory Layer — Master Architecture

Memory Layer — Master Architecture

Persistent, cross-channel, auto-training memory for the BoostEcom AI Teams. Status: this is the May 2026 design plan, and the three phases below all SHIPPED (the header said "implementation in three phases" until…

Persistent, cross-channel, auto-training memory for the BoostEcom AI Teams. Status: this is the May 2026 design plan, and the three phases below all SHIPPED (the header said "implementation in three phases" until 2026-09-05, long after they were in the tree). Keep this page for decisions D1 to D10 and their rationale; for what runs today, module by module, read src/features/ai/memory/README.md, which is the living doc.

North star

A merchant should never have to re-explain context. After three conversations, @Atlas knows the brand voice, the catalog footprint, the operator's preferences and tone, the partner stack, the recurring constraints, and uses that knowledge on every channel (web chat, WhatsApp, future Slack / SMS) for years. The relationship deepens like a human agency relationship that compounds over time.

What already exists (DO NOT duplicate)

The Explore agent surfaced infrastructure we keep:

SurfacePathRole
Conversation + Message Prisma modelsprisma/schema.prisma — model Conversation / model MessageCross-channel persistence already keyed on (externalId, channel) with orgId/storeId/userId joins. Cascade delete. Indexes ready.
StoreContext Json blobprisma/schema.prisma — model StoreContextPer-store singleton with knowledge, todos, pipeline, agentState, conversation, archivedConversations. Multi-agent handoff trail lives here.
Upstash Vector + OpenAI embeddingssrc/features/ai/memory/rag.ts, client in vector-index.tsKnowledge RAG with text-embedding-3-small (1536d). Store-scoped namespace, 800-char chunks. Exposed to the model as knowledgeSearch — this row said searchKnowledge until 2026-09-06, a name nothing answers to, which is the same transposition that put a phantom tool on all six agent profiles (ai-platform-tools-29). rag.ts and embed.ts each declared their own client, singleton and embedding constant against the same index until ai-platform-tools-19; there is now one, in vector-index.ts.
MemorySystem classgoneThis row named src/lib/memory/index.ts until 2026-09-05. The directory was removed by ai-platform/0056 and no MemorySystem class exists anywhere; the memory code lives in src/features/ai/memory/.
3-layer prompt compositionsrc/features/ai/prompts/layer-system.tsLayer 1 (base) + Layer 2 (@Atlas kernel, src/features/ai/prompts/kernel.ts) + Layer 3 (Store.systemPrompt + Store.instructions + Store.knowledgeSources). composeAtlasPrompt({ kernel, ctx }) is the canonical assembler. This row said "iRen kernel" and pointed at src/services/algorithms/prompts/ until ai-platform/0603 and 0599: iRen is the orchestrator's pre-2026 name, and the composer now lives in the pillar that owns the prompt (ADR 0022).
Store.systemPrompt + Store.instructionsprisma/schema.prismaPer-store overrides already injected into Layer 3.
Layer detectionsrc/features/ai/context/layer.tsdetectLayer() returns "guest" | "connected" | "full_agent" and gates tool access. This row said "same file" as the composer above; it never was.
Cron infrasrc/app/api/cron/* + vercel.jsonDozens of jobs already running (intelligence-tick, reset-credits, weekly-digest, ...); vercel.json is the count, and this row said 10 until 2026-09-05. Daily cadence available.

Implication: we extend the three-layer prompt and the StoreContext blob with structured facts rather than introducing a parallel memory store. We reuse the Upstash Vector pipeline for vector retrieval. We add one cron for memory consolidation (no infra needed).

Gaps to fill (this project's scope)

  1. No structured facts: StoreContext.knowledge and StoreContext.agentState are freeform Json with no schema, no scoring, no idempotency, no audit trail. Reading "what does @Atlas know about this org" is a manual JSON dig.
  2. No cross-conversation memory: messages persist but the orchestrator only loads the current conversation's last N turns. A 3-month-old insight from a different conversation is invisible.
  3. No cross-channel surfacing — WhatsApp and web each load their own conversation thread. Same org, same user, same store → two parallel siloed memories.
  4. No vector retrieval over messages: Upstash Vector only indexes knowledge sources, not conversation messages.
  5. No auto-training loop: user corrections in chat don't strengthen anything. Same mistake can recur.
  6. No memory UI: no /[orgSlug]/~/memory page. RGPD requires it.
  7. System prompt has no facts injection point: composeAtlasPrompt doesn't load memory; only static layers.

Architecture decisions

D1. Three scope-specific fact tables, not a polymorphic table

OrgFact, StoreFact, UserFact: separate Prisma models. Reasons:

  • Native onDelete: Cascade to Organization / Store / User.
  • Indexes tunable per scope (org has thousands of facts, user has dozens).
  • Type-safe at the call site (OrgFact[] vs UserFact[]).
  • Cost of duplication: ~15 lines per model. Acceptable.

We do NOT extend StoreContext.knowledge with a typed shape because: (a) it's already in production use with freeform Json, mutating its shape risks runtime crashes, (b) a typed table gives us indexes and queries the JSON path can't.

D2. Append-only with confirmation count, never overwrite

From mem0's lessons + Anthropic's memory tool pattern. When a new fact contradicts an old one:

  • Write the new row with confirmationCount = 1.
  • Decay the old fact's score (see D5).
  • After ≥ 3 confirmations of the new fact, soft-archive the old (archivedAt not null).
  • Surface contradictions to the memory UI for human review.

This preserves an audit trail and avoids destructive corrections on a single noisy turn.

D3. Per-agent namespace = (orgId | storeId | userId, agentId, scope, key)

Adds agentId to the namespace so @Maya (Marketing), @Marco (Merchandising), @Otis (Operations), @Faye (Intelligence), @Sam (Support) each have their own memory pocket. agentId = "atlas" for the orchestrator. Lets @Maya remember "the merchant prefers UGC creatives" without polluting @Marco's catalogue recall. (The roles are the runtime's, specialists.ts; this line carried the retired profile set until ai-platform/3016.)

For Phase 1 we ship with agentId = "atlas" only and add the column NOW (cheap to add, expensive to backfill later).

D4. Two-stage writes — synchronous extraction, asynchronous consolidation

Inline (request path):

  1. After streamText finishes, call Haiku via Gateway to extract 0..5 candidate facts from the last turn pair. The call is generateText + Output.object(), so the Gateway sends responseFormat: json and a non-JSON answer is not representable. It used to be a plain generateText whose text was de-fenced with two regexes and JSON.parsed, and a badly closed fence dropped every fact of the turn under one warning (ai-platform-tools-21).
  2. Persist raw to the fact table with confirmationCount = 1. No conflict resolution yet.

Async (daily cron consolidate-memory):

  1. Group facts by (scope+id, agentId, key).
  2. Apply decay formula (D5) → compute composite score.
  3. Merge duplicates (cosine similarity > 0.92 on value embeddings: Phase 2).
  4. Increment confirmationCount for merged rows.
  5. Archive low-score facts (composite < 0.2).
  6. Write MemoryEvent audit rows for every action.

Mem0's production failures (write-time token cost, indexing inconsistency) all trace to one-stage pipelines. We bake the split in from day one.

D5. Decay + TTL tiers

score = baseConfidence × decayFactor
decayFactor = 0.5 ^ (ageDays / halfLifeDays)

TTL tier defaults (overrideable per fact via the extractor):

Taghalf-life (days)hard TTL (days)
legal, allergy, accessibility∞ (no decay)∞
brand, target-market, catalog180none
preference, tone, language90none
season, campaign, current-focus30180
casual1460

The retrieval scoring uses the composite score; the consolidation cron archives facts below 0.2. Hard TTL = full delete after N days, RGPD-friendly.

D6. Prompt-cache stability — memory injected AFTER the static prefix

Anthropic prompt caching (5-min TTL, $0.30/M vs $3/M = 10× lever) requires byte-stable prefixes. The cached prefix is base + atlas-kernel. Memory injection is a NEW Layer 4 block appended AFTER the cached prefix:

[Layer 1 base + Layer 2 kernel]   ← cached, stable
[Layer 3 store overlay]            ← cached, stable
[Layer 4 memory facts]             ← non-cached, mutates per request
[Channel context]                  ← non-cached
[Conversation history]             ← non-cached

If we ever pollute Layer 1-3 with mutable memory, the cache invalidates every turn and cost balloons 10×.

D7. Vector retrieval reuses Upstash Vector, not pgvector

Upstash Vector is already wired, with OpenAI text-embedding-3-small, store-scoped namespaces, and 1536-dim vectors. We extend the same pipeline to index fact values and message contents. We do NOT introduce pgvector to Neon: that would split the vector surface across two stores and double the operational complexity.

Phase 2 adds two new namespaces:

  • facts:${scopeId} — fact values embedded for cross-fact dedup + semantic memory lookup.
  • messages:${storeId} — message embeddings for "find earlier conversations about X".

D8. Layer / Letta-inspired tiered memory

Three tiers loaded at request time, each with its own budget:

TierSourceBudgetRefresh
CoreHigh-confidence stable facts (composite > 0.7, tag in brand/target-market/tone/language)~300 tokensEvery request
WorkingCurrent conversation summary (last 12 turns OR Haiku rollup if > 12)~500 tokensEvery request
ArchivalTop-k semantic retrieval over facts + message embeddings~600 tokensEvery request, Phase 2

Total ~1400 tokens of memory per request. Combined with the static 3-layer prompt (~2500 tokens), every request stays well under Anthropic's input budget while keeping the prefix cacheable.

D9. Memory UI in Phase 1, not Phase 3

RGPD Article 17 (right to be forgotten) requires per-fact visibility + delete on user demand. Shipping the UI in Phase 3 leaves us legally exposed for months. Ship the page in Phase 1 alongside the schema.

D10. Opt-out per request via #noremember

A trailing #noremember in a user message disables extraction for that turn pair. Cheap to implement (string check before scheduling extraction), useful for sensitive topics. Documented in the WhatsApp / web chat help.

D11. La veille : ce qu'on apprend chez les autres, jamais pris pour nous

Un audit de concurrent ou de marque d'inspiration sert a ameliorer la boutique connectee, donc ses conclusions meritent d'etre retenues. Elles ne doivent jamais l'etre COMME la boutique : « avis sous le prix » vu sur competitor.com est une piste pour nous, pas un fait sur notre theme.

Pas de table dediee. Un fait de veille est un fait ordinaire (scope store en pratique, puisque c'est elle qu'on ameliore) qui porte les deux marqueurs a la fois : la key watch.<site>.<sujet> et le tag veille. src/features/ai/memory/watch.ts les fait concorder sur chaque brouillon extrait (l'un sans l'autre est complete, ou le fait est jete s'il ne tient plus dans 80 caracteres) ; format.ts les rend dans une section « Veille » a part, donc aucune ligne sous « Store » ne decrit le site d'un autre. Demi-vie : 60 jours, une page concurrente bouge plus vite qu'une marque.

Three-phase plan (revised, all three shipped)

The week estimates below are what was planned in May 2026, not work remaining. Where each phase landed:

PlanShipped as
Phase 1 extraction / write / retrieval / formattingsrc/features/ai/memory/extract.ts, write.ts, retrieve.ts, format.ts
Phase 1 UI + cronsrc/app/(dashboard)/[orgSlug]/~/memory/page.tsx, cron consolidate-memory in vercel.json
Phase 2 vector retrievalsrc/features/ai/memory/embed.ts (Upstash Vector namespaces)
Phase 3 corrections + rulessrc/features/ai/memory/corrections.ts, src/app/(dashboard)/[orgSlug]/~/settings/memory/rules/page.tsx

Phase 1 — Structured facts + injection + UI (2-3 weeks)

  • Prisma : OrgFact, StoreFact, UserFact, MemoryEvent, schema in D1 with namespace per D3.
  • Idempotency: @@unique([scope+id, agentId, key, sourceMessageId]).
  • src/features/ai/memory/extract.ts — Haiku extractor per D4 stage 1.
  • src/features/ai/memory/write.ts — idempotent writer with MemoryEvent audit.
  • src/features/ai/memory/retrieve.ts — tag + confidence + recency filter, ~1000 token budget, < 50ms p95.
  • src/features/ai/memory/format.ts — renders the Layer 4 memory block.
  • Integration: composeAtlasPrompt gains a Layer 4 slot. Both web chat handler and WhatsApp bot read memory before prompt assembly, schedule extraction after stream finish.
  • Cron consolidate-memory (no-op for Phase 1 dedup, just decay + audit + archive).
  • Page /[orgSlug]/~/memory — list facts grouped by scope, edit / delete / export. RGPD-compliant.
  • Opt-out #noremember keyword.

Phase 2 — Vector retrieval + cross-channel memory (1-2 weeks)

  • Extend Upstash Vector with facts:${scopeId} and messages:${storeId} namespaces.
  • Embed every persisted fact value + every assistant/user message (just-in-time on first retrieval miss, not eager).
  • getMemoryContext gains an "archival" tier per D8: cosine similarity > 0.7 against the current user turn, top-k 5 facts + top-k 5 messages.
  • Consolidation cron upgraded: cosine > 0.92 → merge facts, increment confirmationCount.
  • Cross-channel: WhatsApp inbound (org/store linked via WhatsappChannel) loads the same memory as the user's web chat.

Phase 3 — Auto-training + per-agent memory + skill rules (2 weeks)

  • Per-agent agentId populated for the 5 specialists. Routing via atlas.route writes to the right pocket.
  • Correction detection (heuristic: "no", "actually", "non plutot" within first 60s after an assistant turn), auto-creates a high-confidence fact with source.reason = "user-correction".
  • .boostecomrules per-org file pattern (Cursor-style): /[orgSlug]/~/settings/memory/rules, Markdown editable rules that always inject into Layer 3.
  • A/B harness: per-org flag memory.enabled = true|false (default true). Measure response quality (thumbs / regenerate rate) with and without.

Scale-up plan (deferred)

The pipeline is wired with raw void promise.catch(...) (WhatsApp webhook) and next/server's after() (web Route Handler). This holds up to roughly 1 000 turn pairs / day on Vercel's default function runtime grace period.

TriggerMigration target
> 1 k msg/day OR p95 extract latency > 4 sUpstash QStash queue — push the turn pair to a queue inside extractAndPersistFacts, drain with a /api/jobs/memory-extract Route Handler
> 10 k msg/day across multiple regionsVercel Queues (when GA) for built-in retries + dead-letter routing
Cross-region read traffic > 50 RPS on getMemoryContextAdd Upstash Redis cache layer in front of the core tier query (TTL ~60 s)

Migration is non-disruptive: extractAndPersistFacts already swallows errors, so the queue can run async without changing the caller contract. The MemoryEvent audit table makes replay safe (every action is idempotent on (scopeId, agentId, key, sourceMessageId)).

We don't pre-build the queue today because:

  1. Current volume is < 100 turns/day across all orgs.
  2. QStash adds a per-message billable line item ($0.0005/msg → $15/mo at 1k/day; acceptable but unjustified at < 100/day).
  3. after() + void promise already give us the lifecycle guarantees we need at our scale.

Cost model

Haiku extraction per turn pair: ~200 input + ~100 output tokens at $1/M + $5/M = ~$0.0007 / turn.

VolumeWorst caseWith batching (every 2 turns)With noisy turn skip (~40%)
1k orgs × 1k turns/mo$700/mo$350/mo$210/mo
10k orgs × 1k turns/mo$7k/mo$3.5k/mo$2.1k/mo

Embedding cost (Phase 2): facts only, ~50 tokens each, $0.02/M = $0.000001/fact. Negligible.

Storage cost (Neon): Fact rows are ~500 bytes each. 1M facts = 500 MB. Negligible inside our existing Neon plan.

RGPD posture + tenant security (Phase 1 + hardening)

Cross-tenant isolation

  • Every fact-table query is gated by a tenant scope (orgId / storeId / userId) at the API + retrieve layer. Enforced in CI by scripts/audit-memory-scope.sh (see Hardening below), a PR that adds a prisma.orgFact.findMany without a where.orgId fails the run.
  • Upstash Vector namespaces are scope-keyed: facts:org:<orgId> / facts:store:<storeId> / facts:user:<userId> / messages:<storeId>. Cross-namespace queries are impossible at the client API level.
  • The scope id used by persistFacts comes from the server-side context (the session's orgId / storeId / userId), NEVER from the model's output. A malicious user can't make the extractor write to another org by including a tag in their message.
  • getOrgAccess(userId, orgId) is called on every memory API route; member role required to read, admin+ to mutate, owner to reset.

Prompt-injection resistance

  • The Zod FactDraftSchema constrains key to dotted-lowercase (regex /^[a-z0-9_.]+$/), tags to ≤ 8 strings ≤ 40 chars each, confidence to [0,1], reasoning to ≤ 240 chars. Anything outside the schema is dropped silently — fact by fact, which is why the Output.object() envelope (ExtractionEnvelopeSchema) is deliberately looser than FactDraftSchema: an output specification is validated as a whole, so putting the key regex there would make one malformed fact discard the entire turn. The envelope pins the shape, FactDraftSchema pins the contents.
  • extract.ts rejects facts whose serialised value exceeds 4 KB (defense in depth against payload smuggling).
  • The extraction prompt is read-only and treats user content as data: the model is told to extract claims, not execute instructions. Even if the model is jailbroken, the worst-case output is a low-quality fact that lands in the right org's pocket and stays subject to RGPD delete + audit trail.

RGPD posture

  • onDelete: Cascade on all fact tables to parent Org / Store / User. Deleting the parent wipes their memory atomically.
  • Per-user delete: account deletion from /account/settings (POST /api/auth/delete-account) drops the User row, and UserFact.user is declared onDelete: Cascade, so every fact goes with it. There is no memory-only per-user wipe.
  • Org-scope reset in /[orgSlug]/~/memory → wipes all OrgFact + StoreFact for the org's stores (owner-only).
  • Per-fact delete (Article 17) via the memory dashboard. Cleans up Upstash Vector via deleteFactFromVector.
  • MemoryEvent rows do not cascade with their fact, and this page said they did until 2026-09. The model declares no relation at all: factId and scopeId are bare strings, denormalized for filtering, so a fact delete leaves its events, oldValue / newValue payloads included. What actually removes them is the account-delete route, which calls deleteMany on scopeId for the user plus every org and store being removed. Read as written, this line said no explicit cleanup was needed anywhere. src/test/erasure-leaves-no-pii.test.ts now fails if the model gains no relation and the route stops naming it.
  • Export endpoint (GET /api/organizations/[orgId]/memory/export) returns a JSON dump for Article 20 (data portability).

Hardening

scripts/audit-memory-scope.sh runs in CI via .github/workflows/audit-memory-scope.yml. It grep-walks the source tree, flags any prisma.<factTable>.<readOrMutate>( call that does not include (orgId|storeId|userId|scopeId|id:) in the following 12 lines. Two exempt files (the consolidation cron and the export endpoint) operate by primary key after a scoped lookup and are explicitly allowlisted.

Open questions (decide before coding)

  1. agentId column type — string ("atlas", "maya", ...) vs enum. String wins for forward-compat; we add a Zod parser at the call site.
  2. Confidence floor for storage: drop facts below 0.4, or persist them flagged? Drop is cheaper; persist enables learning over time. Recommend drop in Phase 1.
  3. #noremember granularity — disables extraction for this turn only, or for the whole conversation? Recommend per-turn (less surprising).
  4. WhatsApp memory recall visibility : say "Je me souviens que tu m'as dit X" or stay invisible? Recommend invisible (trust-safer).
  5. Memory UI default state: show ALL facts even low-confidence, or only > 0.5? Recommend > 0.5 with a "show all" toggle.

Anti-patterns (from research, never do)

  • Eager embedding every message. Just-in-time only — saves ~10× cost.
  • One-stage extract-and-upsert. Corrupts on contradiction. Always async-consolidate.
  • Memory injected into the cached prompt prefix. Kills the 10× cache discount.
  • Per-message memory without a token budget. Cap at ~1400 tokens / turn.
  • Memory in Upstash Redis as the source of truth. Volatile + no joins. Keep Redis for working-memory hot cache only.
  • No user UI. RGPD Article 17 violation.

References

  • mem0 production lessons (token optimization playbook 2026).
  • Letta hierarchical memory (core / working / archival).
  • Anthropic prompt caching docs (5-min TTL, cache key design).
  • Anthropic memory tool docs (append-only, confirmation count pattern).
  • ChatGPT Memory + memory toast pattern (transparency UX).
  • Cursor .cursorrules pattern (per-project Markdown rules: inspires .boostecomrules).
  • FadeMem selective forgetting research (decay formulas).