Metering and billing
Gateway telemetry records inference usage and Substrate's actor lifecycle records measure compute. A durable ledger carries delivery evidence to Metronome, which rates usage. Stripe collects payment. These are separate responsibilities, and a missing measurement must remain unknown rather than become a zero charge.
Usage to collection
agentgateway access telemetry Substrate actor lifecycle records
→ Cloud Logging → Pub/Sub → Cloud Logging → Pub/Sub
→ platform-api gateway subscriber → platform-api compute subscriber and folder
→ Firestore usage evidence and the usage_events ledger
→ elected usage emitter → Metronome usage ingestion and rating
→ Metronome invoice / native Stripe integration → Stripe collection BILLING_USAGE_PROVIDER supports metronome or disabled. The removed Stripe Billing Meters
reporter, hourly Stripe bucket seals and runtime-event recorder are not alternative producers.
In Metronome mode, startup requires a resolved, persisted gateway billing cutover. Missing cutover
state must stop startup instead of running an evidence-only subscriber that silently bills nothing.
The subscriber uses gateway request identity and attribution to deduplicate evidence and ledger writes. The emitter records delivery outcomes and retries within the provider’s accepted window. An accepted ingest response, rated usage and a paid invoice are different facts. Expired or ambiguous delivery needs reconciliation; it cannot be repaired by blindly sending the same usage forever. Metronome’s idempotency reference documents separate retention rules for usage transaction IDs and API request keys.
What each money value means
| Value | Basis and limit |
|---|---|
| Customer usage and rated amount | Metronome’s customer contract and pricing; reads carry source, observation time and coverage |
| Gateway cost field | Telemetry cross-check; it is not proof of the vendor’s invoice or the final customer charge |
| Vendor list estimate | Usage multiplied by a supported current vendor list price; not an event-time invoice, negotiated discount or billed vendor cost |
| Cloud Billing export | A separate reconciliation source when configured, subject to export delay and adjustments |
| Missing measurement | Absent/unknown; never silently converted to measured zero |
Native managed pricing is the Metronome contract/rate-card path. The admin managed-pricing APIs read, preview and apply it. Legacy markup rate-card RPCs must not be treated as the authority for Metronome charges or as proof that historical vendor costs can be reconstructed.
Compute running seconds
Compute is the time a session’s Substrate actor spends in the running state. It is measured from
the lifecycle records Substrate’s API server writes on every committed state change. Running time
at or after the pinned bill-from, 2026-09-30 05:00 UTC, is billable at the compute rate shown
on the pricing page. Suspended time is not charged. The unit is persisted actor state, not CPU
consumption: slow crash detection extends the running state until the control plane changes it.
ate-api-server "Actor state changed" / "Actor crashed", written twice (stdout and OTLP)
→ Cloud Logging sink compute-usage → Pub/Sub topic compute-usage
(a dead-letter topic after ten refused deliveries; a second sink with the same
filter keeps every record 400 days in the compute-usage-evidence log bucket)
→ platform-api subscriber: one Create-only compute_usage_records document per copy
→ platform-api folder, on one elected replica: running intervals in compute_usage_intervals
→ usage_events rows of kind agent_compute, one per UTC-hour slice
→ Metronome metric "Agent compute seconds" (unit pod_seconds, group key agent_id) The records. After committing an actor state transition, Substrate’s ate-api-server logs one
record twice with the same event time: a JSON line on stdout, which GKE’s workload logging
collects, and an OTLP log record exported through the in-cluster collector
(ate.actor.state_changed, or ate.actor.crashed for a crash). Nothing in Substrate is forked or
patched for this; the record shapes are pinned by fixtures in platform-api. Both copies are routed,
either alone is enough to bill, and the folder reports a record whose copies disagree or whose
second copy never arrives. Upstream’s guidance for this stream is “Do not sample this stream. A
dropped record leaves the last known state wrong, with nothing to show that a record is missing”
(Substrate telemetry);
the sink filter selects the two record shapes and never samples.
Intervals. The folder merges the two copies, orders an actor’s records by event time and opens
one interval per running record. The interval closes at the actor’s next record once that record
has settled (10 minutes), with the end decided on what the evidence shows:
| Next record | End | Basis |
|---|---|---|
suspending, pausing, crashed, deleting or reverting | that record’s time | direct |
suspended, paused or deleted within 120 s of running (the transition record was lost) | that record’s time | release_within_gap |
| anything else, or a released state later than 120 s | the start: 0 s billed beyond it | lower_bound |
| no record for 24 h | the start: 0 s billed beyond it | finalized_open |
Every basis other than direct is reported, and an interval still open after 15 minutes is
reported. Corrections only add: a later record that would reduce an already billed interval is
reported and ignored.
Attribution. An interval is billed only for a tenant namespace, only for a session actor
(session-<id>) and only through the session’s immutable binding, which names the agent
incarnation UID the rows are charged to. A binding that is missing, invalid or in another namespace
refuses the interval; refused intervals are reported and never billed. Running time before the
bill-from, or in a non-tenant namespace such as the golden snapshot build, stays evidence only.
Rows and delivery. A billed interval is cut at every UTC hour boundary into usage_events rows
of kind agent_compute, each holding integer milliseconds under the id actor_<uid>_<sliceStart>.
That id is also the Metronome transaction_id, so a replay of the subscription writes the same rows
and changes nothing. The emitter sends each row as one event carrying agent_id and compute_seconds as an exact decimal (1.831 for 1831 ms); the metric sums those values and the
product converts the sum to hours once. A slice older than Metronome’s backdate window is not
written and is reported. A slice written after its month’s invoice finalized is still written and
reported, because no invoice carries it until that invoice is regenerated.
The honest limitation. Delivery of the records is best effort: Substrate keeps no outbox, so a
record lost by both copies leaves no interval. A lost running record under-bills the whole
session, and only the hourly task check (every finished task must fall inside a running interval of
its session’s actor) notices it. A lost end record bills nothing beyond the proven start unless a
released state arrives within the 120 s gap heuristic, in which case at most that much suspend
time is over-counted. The event time is the API server’s clock just after the commit, on whichever
replica ran the workflow, so replica clock skew can reorder two close records. What stands against
each of these is alerting, not an assumption: alerts fire for sink export errors, subscription
backlog, dead letters, an interval open past 15 minutes, copies that disagree, a task no interval
covers, a meter silent while tasks ran, an uncertain end, a refused interval, a correction that
would reduce billed time, a slice too late for its invoice, a pipeline fault and a folder that stops
heartbeating. The raw records are kept 400 days in the compute-usage-evidence log bucket and in
Firestore, so a disputed charge can be checked against what Substrate wrote.
The bill-from is pinned once in Firestore and is one-way: running time before it is never backfilled, and a configured instant that differs from the pin is refused, not applied. This pipeline does not implement measured storage billing; session snapshots and configured storage capacity are not storage-usage measurements.
Payment standing gates new work
Stripe’s Payment Element captures a payment method using a server-created SetupIntent. The server verifies ownership and webhook-derived billing state; a browser success screen alone does not grant paid work. Missing billing configuration or unverified standing fails closed.
Admission considers all outstanding invoice standing. The default policy allows scheduled retries; pause_on_any_failure is stricter. No method, action required, hard decline or unverified state blocks
new paid work. This gate does not cancel an in-flight run or liquidate paper positions.
The console must distinguish a known zero balance or charge from a failed, stale or partial read. GetCurrentUsage reports the provider’s available quantities and amounts with provenance. Its
estimate is not a guarantee that the eventual invoice will be identical.
Credits and payment retries
Credit purchases use a stable purchase UUID, including when resuming an operation from another browser tab. The server persists the operation and its original terms before calling Metronome, then binds the result to its invoice. Repeating or resuming the operation reuses that identity even if another tab has cleared its local pending state. A changed selection requires a new purchase, but an unresolved purchase must be reconciled before starting another. A timeout or ambiguous provider response stays pending until evidence establishes the result. It must not be described as proof that no charge occurred.
Provider creation retries are bounded; older unresolved purchases require reconciliation. Invoice webhooks project standing without regressing a paid invoice to an earlier failure. A failed manual payment-gated commit is voided and is not retried automatically. Once failure is confirmed, the customer can deliberately start a new purchase. An ambiguous result stays pending and keeps its original identity. See Metronome’s payment-gated commit lifecycle.
Metronome’s Stripe integration describes the provider relationship; the platform source determines its configured behavior. Live acceptance of the purchase path on a test tenant is tracked as release work, not claimed by this page.