abliteration.ai is usage-based: you pay for the tokens you send and receive. Input and output are billed at the same flat per-model rate; recognized cache reads are billed at a discount. Rates below are per 1M tokens in USD.
Per-token rates
See models for context windows and capabilities.
Prompt caching
When a request reuses a prefix that’s already cached, those cache-read input tokens are billed at 10% of the model’s standard input rate. Cache creation (processing tokens and making them eligible for possible later reuse) is charged at the standard input rate — there’s no cache-write premium, and a cache write does not guarantee a later cache read. Output tokens are always billed at the standard rate.
Caching is automatic; cached and uncached input tokens are reported separately in the response usage. Cache hits are best-effort, and cached-token counts are measured per request rather than accumulated across a conversation. Use the reported cache-read tokens to determine whether the discount applied. See prompt caching.
Billing and balance
Usage is billed in USD against a single prepaid balance that works across every model. You pay the per-token rates above; cache-read input tokens are billed at the discounted rate, so they draw down your balance proportionally less. Image and video inputs — supported on the base model only — are metered as tokens on the same per-token scale.
Subscriptions and credit packs
Monthly subscriptions and prepaid credit packs are available, along with custom volume and enterprise pricing. See the pricing page for current plans and to buy credits. Last modified on September 2, 2026