GPU compute
Customer priceLoading live GPU pricing…
Training
Estimated per 1M tokensBilling is based on GPU-hours, runtime, and storage — you pay for compute time, not per-token markup.
Deployments
GPU-hour basedClusters
CapacitySpot pricing is not exposed as a user-facing option yet. Current jobs auto-select configured providers unless a provider is specified.
Storage
MonthlyObservability
Request traces · opt-in per deploymentTraces capture full request/response bodies for deployments that opt in (off by default). Metered from credits at month end; the first 5,000 base-retention traces per workspace each month are free.
Dedicated serving
Sized to your traffic · compared vs frontier API list pricesComparison
Your traffic on Veri vs frontier API list pricesEstimates assume continuous 24/7 provisioning (43,800 min/mo) sized to your peak load, billed at an estimated $2.50/GPU-hr (live rate temporarily unavailable). Savings compare against the selected frontier model's published list price (cached input billed at the provider's caching discount) on the same traffic profile. Throughput per GPU varies with context length and serving configuration — talk to us for a sized quote.
