Tenant isolation

One Kubernetes namespace per tenant, and one process — platform-api — that decides which namespace a caller is in. Everything else follows from those two facts.

The tenant comes from the principal, never from the request

Every RPC authenticates with a Firebase ID token sent as Authorization: Bearer <idToken>. The server verifies it with the Firebase Admin SDK — signature against Google’s JWKS, iss = https://securetoken.google.com/<project>, and aud = <project> — then resolves uid → tenant from a cached Firestore lookup. On first sign-in it provisions the tenant record and its namespace (tenant- + tenant id).

On the tenant-facing services, no request carries a namespace or a tenant id. agent_id is a bare name, a payment method is only ever the caller’s own. One tenant cannot name another’s agent because there is no field in which to name it.

Only AdminService takes a tenant_id, and it re-checks the caller’s email against the server’s ADMIN_EMAILS allowlist on every call. The comparison is case-insensitive, an unverified email address never counts, and an empty allowlist denies everyone — the correct failure mode for a misconfigured deployment, and the server warns about it at startup. IdentityService.WhoAmI.is_admin is a hint for the UI, not the enforcement point: the admin console is a static bundle on a public CDN, so anyone can patch that flag in a devtools console and get nothing for it.

The interceptors are full interceptors, not the unary helper

Both auth and admin are connect.Interceptor implementations rather than connect.UnaryInterceptorFunc. That is deliberate: the unary helper wraps unary calls only, and every stream in this API is server-streaming, so a unary-only interceptor would leave StreamRunEvents, StreamNotifications and StreamDashboardStats unauthenticated.

The chain is OpenTelemetry, then error logging, then auth, then admin — OTel first so a rejected request is still traced, admin last so it reads the principal auth has just placed and fails closed if it is absent. CORS wraps the whole mux, because the SPAs and the API are different origins and a failed preflight never reaches a handler.

The tenant baseline

The baseline converges the tenant’s runtime prerequisites: a gVisor WorkerPool, the spooky-kagent Harness, a tenant snapshot location, credential-provider access, catalog ModelConfigs, platform/broker MCP discovery and prompt libraries. The custom runtime already contains its pinned skills; there is no per-tenant OCI skill-pull refresh loop.

Baseline convergence runs for newly resolved tenants and existing tenants. A completed baseline is separate from agent readiness or successful session execution. Dependency failures must not be reported as a ready tenant merely because the namespace exists.

References and service identity

Tenant-facing writes validate model and delegation references in the caller’s namespace. The API must not accept arbitrary namespace-qualified references merely because Kubernetes can resolve them. Agent-specific bindings and credentials are distinct from discovery credentials, which can list a catalog without being authorized to execute an agent’s tools.

Platform MCP agent tokens bind the namespace, agent name and immutable agent UID. Scheduling tools check that UID against the current agent and, for existing wakes, their recorded owner incarnation. Name-only legacy agent tokens are refused. Startup rotates a verified current agent’s exact legacy Secret entry before scheduling and billing workers start; it does not accept an old token as proof of ownership. Already prepared actors may need a fresh session or reprepare to load the new credential. Changing a Secret alone is not a guarantee that a running actor has refreshed its configuration.

Substrate actor identity, credential-provider grants and gateway authorization cooperate to enforce runtime service access. Network reachability alone is not caller identity. For model attribution and brokerage ownership, see The model gateway and The paper brokerage.

Quotas and suspension

Tenant quotas bound supported resource counts. Their effective defaults must remain visible to the operator; they are not monetary spend caps. Tenant suspension gates new work while preserving permitted reads. Agent/session suspension uses runtime lifecycle APIs rather than scaling each agent’s Deployment. Payment-standing gates are separate from both resource quotas and operator suspension. See Metering and billing.

Firebase authorized domains

Sign-in is signInWithPopup with GoogleAuthProvider, and authDomain stays at the project’s default spooky-labs.firebaseapp.com. Pointing it at developer.spookylabs.ai breaks sign-in: a custom authDomain requires Firebase Hosting, and GitHub Pages cannot reverse-proxy /__/auth/.

What does need configuring is Firebase’s authorized domains list — the hosts that sign users in, developer.spookylabs.ai and admin.spookylabs.ai, plus localhost — which is Terraform, as google_identity_platform_config.authorized_domains, and is a whole-list replacement.