Per-token rates
See models for context windows and capabilities.
Prompt caching
When a request reuses a prefix that’s already cached, those cache-read input tokens are billed at 25% of the model’s standard input rate. Cache creation (the first request that populates the cache) is charged at the standard input rate — there’s no cache-write premium. Output tokens are always billed at the standard rate. Caching is automatic; cached and uncached input tokens are reported separately in the responseusage.
Credits
Usage is metered in credits, and one balance works across every model. At full-price rates, credits work out to about 1 credit per 500 tokens onabliterated-model and about 1 credit per 300 tokens on abliterated-model-large, with a minimum of 1 credit per call. Cache-read input tokens are metered at their discounted rate, so they consume proportionally fewer credits. Image and video inputs — supported on the base model only — are metered as tokens on the same scale.