Architecture overview

The browser reaches platform-api. That API owns tenant authorization and platform workflows; kagent and Substrate run sessions; agentgateway handles model traffic; Metronome rates usage and Stripe collects payment.

Browser → GKE Gateway → platform-api
  ├─ Kubernetes: Agent + AgentTemplate + Harness, Schedule, BrokerAccount
  ├─ kagent APIs: sessions, tasks and live A2A observations
  ├─ Firestore: tenants, billing state, usage evidence and delivery operations
  ├─ Metronome: customer contracts, usage rating and credits
  └─ Stripe: payment methods, invoice collection and payment standing

Session → tenant Substrate WorkerPool → pinned kagent runtime
  ├─ agentgateway → Vertex AI
  ├─ platform MCP → scheduling
  └─ broker MCP → paper trading

Runtime and tenancy boundaries

Firebase identity is resolved to a tenant by platform-api. Browser clients do not choose their Kubernetes namespace or connect directly to kagent’s control plane. Admin operations add a server-side allowlist check. See Tenant isolation.

The pinned kagent alpha5 integration uses api.kagent.dev/v1alpha3 Agent, AgentTemplate and Harness resources. The Harness references a tenant WorkerPool and snapshot location. Sessions and tasks are reached through kagent’s APIs; there are no platform Session or Task CRDs. See The agent runtime for the version boundary with current upstream docs.

Substrate supplies session execution and suspend/resume. The custom runtime has baked Git-package skills and file tools, with native Bash removed. The platform MCP endpoint supplies scheduling. See Runtime isolation.

State and evidence

Concern Owner
Agent definitions and runtime configuration Kubernetes resources reconciled by the stock kagent and Substrate controllers at their pinned chart versions
Schedules and brokerage account ownership Platform custom resources and reconcilers
Session/task projections and execution state kagent/Substrate APIs and their persistence
Tenants, notifications, payment state and usage delivery evidence Firestore through platform-api
Customer usage rating and credits Metronome
Payment collection Stripe
Leaderboard ordering Redis index with platform-owned source records
Traces and metrics OpenTelemetry collectors, Cloud Trace and Cloud Monitoring

Live output, durable usage evidence, customer-rated charges and vendor-cost estimates are different records. The streaming and billing contracts state their coverage and reconciliation limits. Compute running seconds are derived from Substrate’s actor lifecycle records, routed through Cloud Logging and Pub/Sub into platform-api, which folds them into running intervals and ledger rows. Dashboard resource capacity is not a substitute for that measurement.

Infrastructure dependencies

Terraform owns shared infrastructure, not each tenant’s agent. The tracked CI dependency graph makes gcp/platform a producer for both github and gcp/workloads; changes select downstream consumers and execute producers first. Workloads providers need a real cluster even during a plan. The separate stripe root manages native collection configuration, not usage meters.

The deploy pin file binds both platform and runtime image digests and the baked skills commit. See The deploy path for promotion and verification boundaries.

Observability

The gateway exports access telemetry used by the inference ledger and publishes scrape metrics. Substrate’s actor lifecycle log stream, in its stdout and OTLP copies, is the compute meter’s evidence, so it is routed and never sampled. The collectors also observe controller/runtime health; these metrics do not replace usage evidence. Delegation graphs are reconstructed from configured trace reads and may be unavailable or incomplete. Request content is not required for usage attribution, so logging policy should keep prompt and completion content out of the gateway usage stream.