Skip to main content
abliteration.ai is usage-based: you pay for the tokens you send and receive. Input and output are billed at the same flat per-model rate; recognized cache reads are billed at a discount. Rates below are per 1M tokens in USD.

Per-token rates

See models for context windows and capabilities.

Prompt caching

When a request reuses a prefix that’s already cached, those cache-read input tokens are billed at 25% of the model’s standard input rate. Cache creation (the first request that populates the cache) is charged at the standard input rate — there’s no cache-write premium. Output tokens are always billed at the standard rate. Caching is automatic; cached and uncached input tokens are reported separately in the response usage.

Credits

Usage is metered in credits, and one balance works across every model. At full-price rates, credits work out to about 1 credit per 500 tokens on abliterated-model and about 1 credit per 300 tokens on abliterated-model-large, with a minimum of 1 credit per call. Cache-read input tokens are metered at their discounted rate, so they consume proportionally fewer credits. Image and video inputs — supported on the base model only — are metered as tokens on the same scale.

Subscriptions and credit packs

Monthly subscriptions and prepaid credit packs are available, along with custom volume and enterprise pricing. See the pricing page for current plans and to buy credits.
Last modified on July 29, 2026