Metering and billing

Gateway telemetry records inference usage and Substrate's actor lifecycle records measure compute. A durable ledger carries delivery evidence to Metronome, which rates usage. Stripe collects payment. These are separate responsibilities, and a missing measurement must remain unknown rather than become a zero charge.

Usage to collection

agentgateway access telemetry          Substrate actor lifecycle records
  → Cloud Logging → Pub/Sub              → Cloud Logging → Pub/Sub
  → platform-api gateway subscriber      → platform-api compute subscriber and folder
  → Firestore usage evidence and the usage_events ledger
  → elected usage emitter → Metronome usage ingestion and rating
  → Metronome invoice / native Stripe integration → Stripe collection

BILLING_USAGE_PROVIDER supports metronome or disabled. The removed Stripe Billing Meters reporter, hourly Stripe bucket seals and runtime-event recorder are not alternative producers. In Metronome mode, startup requires a resolved, persisted gateway billing cutover. Missing cutover state must stop startup instead of running an evidence-only subscriber that silently bills nothing.

The subscriber uses gateway request identity and attribution to deduplicate evidence and ledger writes. The emitter records delivery outcomes and retries within the provider’s accepted window. An accepted ingest response, rated usage and a paid invoice are different facts. Expired or ambiguous delivery needs reconciliation; it cannot be repaired by blindly sending the same usage forever. Metronome’s idempotency reference documents separate retention rules for usage transaction IDs and API request keys.

What each money value means

Value Basis and limit
Customer usage and rated amount Metronome’s customer contract and pricing; reads carry source, observation time and coverage
Gateway cost field Telemetry cross-check; it is not proof of the vendor’s invoice or the final customer charge
Vendor list estimate Usage multiplied by a supported current vendor list price; not an event-time invoice, negotiated discount or billed vendor cost
Cloud Billing export A separate reconciliation source when configured, subject to export delay and adjustments
Missing measurement Absent/unknown; never silently converted to measured zero

Native managed pricing is the Metronome contract/rate-card path. The admin managed-pricing APIs read, preview and apply it. Legacy markup rate-card RPCs must not be treated as the authority for Metronome charges or as proof that historical vendor costs can be reconstructed.

Compute running seconds

Compute is the time a session’s Substrate actor spends in the running state. It is measured from the lifecycle records Substrate’s API server writes on every committed state change. Running time at or after the pinned bill-from, 2026-09-30 05:00 UTC, is billable at the compute rate shown on the pricing page. Suspended time is not charged. The unit is persisted actor state, not CPU consumption: slow crash detection extends the running state until the control plane changes it.

ate-api-server "Actor state changed" / "Actor crashed", written twice (stdout and OTLP)
  → Cloud Logging sink compute-usage → Pub/Sub topic compute-usage
      (a dead-letter topic after ten refused deliveries; a second sink with the same
       filter keeps every record 400 days in the compute-usage-evidence log bucket)
  → platform-api subscriber: one Create-only compute_usage_records document per copy
  → platform-api folder, on one elected replica: running intervals in compute_usage_intervals
  → usage_events rows of kind agent_compute, one per UTC-hour slice
  → Metronome metric "Agent compute seconds" (unit pod_seconds, group key agent_id)

The records. After committing an actor state transition, Substrate’s ate-api-server logs one record twice with the same event time: a JSON line on stdout, which GKE’s workload logging collects, and an OTLP log record exported through the in-cluster collector (ate.actor.state_changed, or ate.actor.crashed for a crash). Nothing in Substrate is forked or patched for this; the record shapes are pinned by fixtures in platform-api. Both copies are routed, either alone is enough to bill, and the folder reports a record whose copies disagree or whose second copy never arrives. Upstream’s guidance for this stream is “Do not sample this stream. A dropped record leaves the last known state wrong, with nothing to show that a record is missing” (Substrate telemetry); the sink filter selects the two record shapes and never samples.

Intervals. The folder merges the two copies, orders an actor’s records by event time and opens one interval per running record. The interval closes at the actor’s next record once that record has settled (10 minutes), with the end decided on what the evidence shows:

Next record End Basis
suspending, pausing, crashed, deleting or reverting that record’s time direct
suspended, paused or deleted within 120 s of running (the transition record was lost) that record’s time release_within_gap
anything else, or a released state later than 120 s the start: 0 s billed beyond it lower_bound
no record for 24 h the start: 0 s billed beyond it finalized_open

Every basis other than direct is reported, and an interval still open after 15 minutes is reported. Corrections only add: a later record that would reduce an already billed interval is reported and ignored.

Attribution. An interval is billed only for a tenant namespace, only for a session actor (session-<id>) and only through the session’s immutable binding, which names the agent incarnation UID the rows are charged to. A binding that is missing, invalid or in another namespace refuses the interval; refused intervals are reported and never billed. Running time before the bill-from, or in a non-tenant namespace such as the golden snapshot build, stays evidence only.

Rows and delivery. A billed interval is cut at every UTC hour boundary into usage_events rows of kind agent_compute, each holding integer milliseconds under the id actor_<uid>_<sliceStart>. That id is also the Metronome transaction_id, so a replay of the subscription writes the same rows and changes nothing. The emitter sends each row as one event carrying agent_id and compute_seconds as an exact decimal (1.831 for 1831 ms); the metric sums those values and the product converts the sum to hours once. A slice older than Metronome’s backdate window is not written and is reported. A slice written after its month’s invoice finalized is still written and reported, because no invoice carries it until that invoice is regenerated.

The honest limitation. Delivery of the records is best effort: Substrate keeps no outbox, so a record lost by both copies leaves no interval. A lost running record under-bills the whole session, and only the hourly task check (every finished task must fall inside a running interval of its session’s actor) notices it. A lost end record bills nothing beyond the proven start unless a released state arrives within the 120 s gap heuristic, in which case at most that much suspend time is over-counted. The event time is the API server’s clock just after the commit, on whichever replica ran the workflow, so replica clock skew can reorder two close records. What stands against each of these is alerting, not an assumption: alerts fire for sink export errors, subscription backlog, dead letters, an interval open past 15 minutes, copies that disagree, a task no interval covers, a meter silent while tasks ran, an uncertain end, a refused interval, a correction that would reduce billed time, a slice too late for its invoice, a pipeline fault and a folder that stops heartbeating. The raw records are kept 400 days in the compute-usage-evidence log bucket and in Firestore, so a disputed charge can be checked against what Substrate wrote.

The bill-from is pinned once in Firestore and is one-way: running time before it is never backfilled, and a configured instant that differs from the pin is refused, not applied. This pipeline does not implement measured storage billing; session snapshots and configured storage capacity are not storage-usage measurements.

Payment standing gates new work

Stripe’s Payment Element captures a payment method using a server-created SetupIntent. The server verifies ownership and webhook-derived billing state; a browser success screen alone does not grant paid work. Missing billing configuration or unverified standing fails closed.

Admission considers all outstanding invoice standing. The default policy allows scheduled retries; pause_on_any_failure is stricter. No method, action required, hard decline or unverified state blocks new paid work. This gate does not cancel an in-flight run or liquidate paper positions.

The console must distinguish a known zero balance or charge from a failed, stale or partial read. GetCurrentUsage reports the provider’s available quantities and amounts with provenance. Its estimate is not a guarantee that the eventual invoice will be identical.

Credits and payment retries

Credit purchases use a stable purchase UUID, including when resuming an operation from another browser tab. The server persists the operation and its original terms before calling Metronome, then binds the result to its invoice. Repeating or resuming the operation reuses that identity even if another tab has cleared its local pending state. A changed selection requires a new purchase, but an unresolved purchase must be reconciled before starting another. A timeout or ambiguous provider response stays pending until evidence establishes the result. It must not be described as proof that no charge occurred.

Provider creation retries are bounded; older unresolved purchases require reconciliation. Invoice webhooks project standing without regressing a paid invoice to an earlier failure. A failed manual payment-gated commit is voided and is not retried automatically. Once failure is confirmed, the customer can deliberately start a new purchase. An ambiguous result stays pending and keeps its original identity. See Metronome’s payment-gated commit lifecycle.

Metronome’s Stripe integration describes the provider relationship; the platform source determines its configured behavior. Live acceptance of the purchase path on a test tenant is tracked as release work, not claimed by this page.