ArchitectureCode protection: what the browser can and cannot learn about the platform

Code protection: what the browser can and cannot learn about the platform

Scope: the web app bundles, source maps, API responses and the surface the BoostEcom Chrome extension talks to. Neighbouring contracts: extension-presence-contract.md (how the extension announces itself to the page) and…

Scope: the web app bundles, source maps, API responses and the surface the BoostEcom Chrome extension talks to. Neighbouring contracts: extension-presence-contract.md (how the extension announces itself to the page) and panel-ingest-contract.md (the one write door the extension uses).

The rule, stated honestly

Anything a browser executes, a person with that browser can read. Minified JavaScript, an extension package unzipped from a .crx, a bundle chunk saved from DevTools: all of it is readable by a patient reader, and no obfuscation changes that for long. So the protection of the platform's proprietary logic (scoring, ranking, classification, prompts, thresholds, the paid data itself) is not client-side hiding. It is three server-side facts:

  1. The computation runs on the server. A client receives a result (health_score: 71), never the weights, thresholds or prompt that produced it.
  2. Every read is authenticated and attributed to a session, an organization key or, for the free tier, an IP.
  3. Every read is metered and plan-gated on the server: per-minute rate limits, and a daily budget of full records per account. A stolen extension has no credentials of its own and no plan: it is a UI that calls the same public routes anyone can call.

Client-side measures (no source maps, server-only markers, the boundary guard) are deterrence and hygiene: they remove the free, annotated copy and stop accidents. They are listed below as what they are.

Threat model

AdversaryGoalWhat stops them
Someone who installs the extension and unzips itRead the UI and logic, clone the productNothing hides the extension's own code (and policy forbids obfuscating it, see below). What it can reach is only what the public routes serve to any caller, at the caller's plan.
Someone who reads the web bundles in DevToolsFind scoring weights, prompts, thresholdsEngines are kept out of client chunks (guard test), prompts and scoring modules are server-only, no source maps.
A script behind one paid accountPage the whole catalogue outDaily budget of full records per account (src/lib/hub/read-budget.ts), burst alert to an operator, per-minute limits.
A script with no accountSame, for freeFree callers only ever receive the capped, redacted preview; per-IP limit; offsets pinned to 0.
A page on another originMake a victim's browser call our APINo CORS headers on /api/extension/*, no credentialed CORS anywhere, CSRF origin check on state-changing routes.
Another extension that copies our idObtain a tokenThere is no extension token. See "Authentication of the extension".

What is hidden (and how)

  • Engines and prompts stay out of browser chunks. A multi-source walk from every "use client" file fails the test suite if it reaches a module under an engine directory (src/services/algorithms/, src/services/discovery/, src/features/ai/agents/agents/, src/features/ai/prompts/, src/features/ai/orchestrator/, src/features/tracking/scan/, src/lib/scraper/, src/lib/fetch-providers/). The first run found a real leak: the chat UI imported one function from team-roster.ts, which imports the specialists' system prompts, so all five prompts shipped in the chat chunk. The handles now live in team-handles.ts (types only), and the prompts and scoring modules carry import "server-only", so a future mistake is a compiler error.
  • A few modules are public on purpose and listed with a reason in the test (the sponsor price list, a unit conversion, URL helpers, check counts).
  • No source maps. productionBrowserSourceMaps is not enabled, nothing in public/ is a .map, and no build script turns them on.
  • Server plumbing is not in the browser either. @/env/server (the name of every secret variable is a map of the infrastructure), the crypto and token helpers, the plan gate and the read budget are asserted unreachable from a Client Component. serverEnv had reached the browser through the logger, the plan catalogue and the site config; those now read process.env.NEXT_PUBLIC_* directly, which the build inlines (values were never exposed, names were).
  • API responses carry results, not machinery. The HubStore contract declares no weights, components, prompts or debug fields (checked by name), the ranking responses no longer name the internal modules that produce a ranking, and error responses carry a code, never a caught error's message or stack. API answers also carry X-Robots-Tag: noindex, nofollow and X-Content-Type-Options: nosniff (next.config.mjs), because src/proxy.ts does not run on /api.

What cannot be hidden

  • The extension's own code. Chrome Web Store policy forbids obfuscated code in extensions (minification is allowed, deliberate obfuscation is not), and an extension that is rejected or pulled protects nothing. We do not obfuscate.
  • The shape of every public API response. If a field is returned to a caller, that caller has it. Which fields a plan receives is decided on the server (src/lib/hub/field-gate.ts); the extension's own plan table is a mirror used to render lock states and is not a control.
  • The plan names and which feature needs which plan: this is the pricing page. FEATURE_MIN_TIER is reachable from client code for rendering; the server stays authoritative.
  • Public numbers such as the sponsor CPM and the check counts shown to visitors.

Authentication of the extension

There is no shared secret and no extension token in this repository, on purpose:

  • The extension never holds a credential. Writes and account-scoped reads go through a signed-in platform tab (the bridge page, same origin), so the platform session cookie is the credential and the extension origin is never trusted for these (src/app/api/extension/CLAUDE.md).
  • chrome-extension://<id> is accepted as an Origin by exactly two routes, POST /api/intelligence/panel/ingest and POST /api/extension/analyze (the service worker's enqueue, no platform tab needed), and only as CSRF protection: the route still requires a session (allowExtensionOrigin is opt-in per route, matched by exact string against the single published id, src/lib/security/csrf-origin.ts). A different or prefix-extended id is refused. A non-browser client can forge any Origin, which is why the origin is never authentication.
  • Because no token is minted for an extension id, there is nothing for an unknown extension id to obtain. If a token handshake is ever added, it must go through an explicit server-side allowlist of ids and keep the session as the credential.

Enforced by tests

TestWhat it pins
src/test/engines-stay-out-of-client-bundles.test.tsNo engine directory module reachable from a "use client" file (transitive, with a reasoned exceptions list); @/env/server, crypto, token helpers, plan gate and read budget unreachable; the prompt and scoring modules carry server-only; the public handles module never imports the prompts.
src/test/client-server-boundary.test.tsA Client Component does not import a value from a module that reaches the database (one hop).
src/test/no-browser-source-maps.test.tsSource maps stay off.
src/test/api-response-headers.test.tsX-Robots-Tag and nosniff on /api, share images excluded.
src/test/extension-api-posture.test.tsEvery /api/extension route needs a session and a per-user limiter (two documented exceptions); no CORS headers there, no credentialed CORS anywhere; the extension origin opt-in stays on one route; every public Intelligence route has the shared guard; routes serving full records spend the daily read budget; no error message or stack in JSON bodies.
src/lib/hub/public-payload-hygiene.test.tsThe HubStore contract and a redacted store carry no weights, components, prompts or internal module ids.
src/lib/hub/field-gate.test.ts, src/app/api/intelligence/plan-gating*.test.ts, read-budget-routes.test.tsWhich fields each plan receives, and that bulk reads are metered.

Gap closed during this review

The competitor watchlist (/api/intelligence/hub/tracker) returned full store records for every tracked domain without touching the daily read budget: add a domain, read the list, delete it, repeat, and the catalogue could be walked at the per-minute rate limit. Adding a NEW domain now spends one unit of the account's daily full-read budget and answers 429 daily_read_budget once it is spent. Reading the list again costs nothing, since it reveals no new record.

Rich lists and the scan door

  • GET /api/intelligence/similar?enrich=1 is the paid list. A caller below Pro (anonymous, free account, or a key of a free org) receives the first two lookalikes (FREE_LIST_LIMITS.similar_stores) plus total, locked (how many were withheld) and required_plan: "pro". The names of the rest are not in the response, so there is nothing left to blur in a page. The lean list (no enrich) is unchanged: public, domain and score only. Pro and above are unchanged and still metered by the read budget.
  • GET /api/intelligence/graph/{domain} already withholds the whole view below Pro ({ locked: true, required_plan }); seo, predict and supplier do the same per plan and charge the read budget. The plan is resolved on the server from the session or the key, never from a request parameter or header.
  • POST /api/intelligence/hub/scan enforces its quota on the server whatever the client sends: 8 per hour per IP (the /64 for IPv6), also 8 per hour per signed-in account (rotating IPs does not reset it), and a platform-wide ceiling of 5 000 live crawls per day (503 scan_capacity_reached with Retry-After; cached answers never spend it). A live crawl still needs a session or a solved challenge, and the paid probes keep their own daily reservation.
  • Interactive and automatic scans have separate budgets. A client marks a scan it started without a user gesture with the optional header X-Boost-Scan-Trigger: auto; no header, or any other value, is click. Automatic scans count in their own per-caller buckets (so they never spend a person's 8 per hour) and in their own daily budget of 1 500 live crawls (AUTO_LIVE_SCANS_PER_DAY), which is also counted inside the 5 000 ceiling. Automatic collection can therefore never take more than 1 500 of the day's crawls: at least 3 500 stay reserved for web search and explicit clicks, and automatic requests are the first to receive 503 scan_capacity_reached. No new environment variable: both numbers are constants in src/lib/hub/scan-capacity.ts. The extension's enqueue door (POST /api/extension/analyze, contract in intelligence-extension-analysis.md) spends the same automatic budget when it queues a new domain; it never crawls inside the request.
  • The scan also enforces the owner's opt-out of a connected store (Store.intelligenceVisibility other than PUBLIC, null included) before any crawl, for a domain with no canonical record yet. The answer is the same not_listed a non-public record gets, and a client needs no guard of its own.

Extension scale

  • Rate-limit identity: a public Intelligence route buckets a bei_ key on its organization, a signed-in session on user:<id>, and only a caller with neither on its IP (the /64 for IPv6). People behind one office NAT or mobile CGNAT address no longer throttle each other. The ceiling per bucket is unchanged.
  • GET /api/intelligence/lookup?q=<hostname>&rich=1 resolves a full hostname by its unique domain (one indexed read) and runs the substring search only when no exact PUBLIC row exists or the term is partial. The response shape is the same either way; limit still caps stores.
  • One plan, two questions. GET /api/usage keeps limits.plan (the request organization's subscription, which prices tokens) and adds intel_plan: the effective plan the Intelligence routes apply to this session, resolved by getCallerPlan (the best unlocked membership across the user's organizations). A client shows intel_plan to describe what the server will serve. Store-scoped features (the tracker board, /api/extension/watch) are deliberately billed to the store's own organization and keep that rule.
  • access.reason: when the account's daily budget of full records is spent, a record is served as the free preview and carries access: { plan, unlocked: false, reason: "daily_read_budget" }, with the response header X-Intelligence-Read-Budget: exhausted. Where it appears: every store in lookup?rich=1 stores[], similar?enrich=1 items[].store, hub/scan store, and the ranking and tracker lists; the per-store views (graph, seo, predict, supplier) answer { domain, withheld: true, access: { plan, unlocked: false, reason } }. A client must render the preview and may explain the reason; it must not retry in a loop. (The reason at the top of a similar answer with an empty list, such as no_neighbors, is a different field: why there is nothing to show.)

Follow-ups (not done here)

  • Proposal, not implemented: a server-rendered per-plan brief (GET /api/intelligence/brief?domain=) that returns one store's already redacted summary and the radar driver labels, so a client renders what it is given instead of redacting with a local plan table.

  • 14 non-public API routes still echo a caught error's message in a JSON body (store assets, branches, mirror sections, admin devstore pool, fleet runs, jobs, ops dev, approvals decide, wizard routes, workflow test-node, preview browserbase). They are authenticated and outside the extension surface, but should return a code and log the message.

  • POST /api/intelligence/beta/consent has no rate limiter (session required, one idempotent upsert); low risk.

  • src/features/growth/growth-kinds.ts ships SCORE_WEIGHTS to the ops console (a client component) so the form can preview a unit's score. It is an internal growth heuristic, not customer-facing; move it behind an endpoint if the weights become sensitive.

  • The marketing page /intelligence/stores/[domain] prints each category's algorithm label (internal module names). It is a deliberate trust signal; decide whether the wording should be product copy instead.

  • Extension-side items for the extension team: keep scoring thresholds and plan tables out of the package (render only what the server sends); make sure no shared secret (the legacy INGEST_SECRET HMAC described in panel-ingest-contract.md) is bundled; treat the locally mirrored plan table as cosmetic; handle 429 daily_read_budget on tracker add.

Adding a new engine or prompt

  1. Put the weights, thresholds and prompts in a module under an engine directory, and add import "server-only" to it.
  2. Expose results through a route that applies the field gate and, if it returns full records, the read budget.
  3. If a Client Component needs one small constant from the same area, put that constant in its own dependency-free module outside the engine directory (the pattern of team-handles.ts), never import the engine for it.
  4. Run npx vitest run src/test/engines-stay-out-of-client-bundles.test.ts.