The model gateway
The configured model path is an in-cluster agentgateway proxy to Vertex AI. It centralizes provider credentials, model routing, usage attribution and policy enforcement for the platform's runtime.
Model access
The tenant’s ModelConfigs point to the gateway’s OpenAI-compatible API. The request chooses the model from the supported catalog; the gateway routes it to Vertex. The proxy uses its GKE Workload Identity for Google credentials. Provider authorization belongs to the dataplane identity.
The gateway is not a public browser API. Runtime egress and credential-provider rules, gateway authentication and authorization, and tenant attribution must agree. A private Service alone is not proof of actor identity, and a caller-supplied attribution header is not an authority to charge a tenant. Gateway policy must derive billing identity from trusted authentication context.
What telemetry proves
Access telemetry supplies attributed inference token usage through Cloud Logging and Pub/Sub to the platform’s evidence ledger. Scraped metrics show request counts, latency and policy outcomes. They serve different purposes: a healthy request counter does not prove that every billable event reached Metronome.
Gateway cost telemetry is a pricing cross-check, not a Google invoice. The customer charge comes from Metronome’s configured pricing. Current vendor-list estimates, rated customer usage and provider-billed cost must remain labeled separately. See Metering and billing.
Limits and standing
Every model call carries the calling agent’s own gateway key. A key that names no agent is refused by the model route with a 403. The admission service behind the route refuses a call for a suspended account or for no known account; that refusal is a 429 the runtime ends the turn on without retrying. The admission check runs before the call with no token count and is amended after the response, so the call that crosses a limit completes and the next one is refused. The policy fails closed: when the admission service is unreachable, model calls fail with an HTTP 500 rather than being admitted unmetered.
The platform’s one money limit is a prepaid balance enforced on that admission path: model calls are refused once the balance reaches zero. That limit is the design landing now, not a served contract yet. The console shows the limits that apply to an account today; the dollar ceiling comes from the prepaid balance and payment standing, not from the gateway’s presence.
The gateway also routes brokerage traffic and broker MCP calls. Those routes use their own authentication, account ownership and market-policy checks; a working model call is not permission to place a paper order. See The paper brokerage.