Charged on actual token usage; how failed requests are handled.
Last updated
Chat models are billed on input + output tokens at the prices shown on the Models page. Cache hits, reasoning tokens, audio and images are priced separately and itemised on your bill.
The test is whether upstream actually spent compute, not whether you got a 200.