LLM API token & billing: estimation, reading, budgeting, and examples
Using an LLM API is like using utilities: if you don't check the meter, the end-of-month bill will surprise you; if you understand it, you know exactly what you're paying for. Here, the "meter" is token counting. This article explains what tokens are, how to roughly estimate Chinese and English usage, how to read the usage field in responses, how to install a budget gate in your code, and finally uses three examples with clear assumptions to apply the input $0.25 / output $1.00 (per million tokens) pricing to concrete scenarios.
Updated on
Key points
- Billing formula: multiply input tokens by $0.25, add output tokens multiplied by $1.00, then divide by 1,000,000.
- Estimation is for pre-estimation; actual usage is based on the usage in the response. Streaming responses include usage in the final chunk.
- For chat apps, the biggest cost is sending history repeatedly. Trimming history is the most cost-effective approach.
- Three main budget control tactics: limit max_tokens, limit history length, and accumulate spend in the program with a daily cap.
What tokens are and how to calculate costs
A token is the smallest unit of text a model processes; it is neither a character nor a word, but a "fragment" between the two. Think of it as a scale mark on a water meter: the more you use, the faster the mark moves. Billing is split into two parts:
| Item | Unit price (per million tokens) | Includes |
|---|---|---|
| Input | $0.25 | All content in messages: system, history, and this request |
| Output | $1.00 | Model-generated response |
Note that the output price is four times the input price, so having the model "talk less" is often cheaper than you "sending less context." However, in conversational scenarios, input grows due to history accumulation, so you need to manage both sides. The payment method is prepaid credit, not a subscription, and the balance never expires; new accounts get $0.50 in free trial credit, valid within 7 days.
One point of confusion: output tokens are the actual number generated by the model, not the max_tokens you set. max_tokens is just the upper limit; if the model finishes in 200 tokens, you are charged only for those 200 tokens. However, when estimating budget, calculate based on the upper limit for the worst-case scenario so you are not caught off guard by occasional long outputs.
Pre-estimation: rough estimates for Chinese and English
Before sending a request, you can only estimate. The following are approximations and may vary by content:
| Text | Rough estimation rule | Example (assumed) |
|---|---|---|
| Chinese | Approximately 1 to 1.5 tokens per character | 3,000 characters ≈ 3,000 to 4,500 tokens |
| English | Approx. 1 token per 4 characters, or 1.3 tokens per word | 1,000 words ≈ 1,300 tokens |
| Mixed Chinese/English, code | Estimate using the higher value | Content with JSON and symbols consumes more |
The correct use of estimation is to make an order-of-magnitude assessment: Is this article roughly 10,000 or 100,000 tokens? Will it break the 100,000 context window? How much does one call cost roughly? For precision, read the usage. A good habit is to sample dozens of runs for each business scenario, calculate the average usage, and use that instead of an estimated value.
Here is an example combining estimation with actual measurement. Suppose you need to process a batch of Chinese customer feedback of about 2,000 characters. Roughly estimate 1.3 tokens per character, making a single request about 2,600 tokens. Add your prompt of 200 tokens, making the input about 2,800. First, run 30 samples, read the usage, and take the average. If the actual average is 2,500, adjust your coefficient to about 1.15. Then calculate the budget for the whole batch using the adjusted coefficient; the error will be much smaller. This measured number is just a hypothetical example of the method; your data will have its own values.
Reading usage: the real "meter reading"
Every successful response carries usage, containing three items: prompt_tokens, completion_tokens, and total_tokens. Streaming responses do not require extra parameters; the final chunk automatically appends a data block with usage. The code below demonstrates both ways of reading and converts the reading to USD:
import os
from openai import OpenAI
client = OpenAI(base_url="https://api.apidamoxing.com/v1", api_key=os.environ["API_KEY"])
PRICE_IN, PRICE_OUT = 0.25, 1.00 # 美元 / 百万 token
def cost(u):
return (u.prompt_tokens * PRICE_IN + u.completion_tokens * PRICE_OUT) / 1_000_000
# 非流式:usage 在响应对象上
r = client.chat.completions.create(
model="uncensored", max_tokens=200,
messages=[{"role": "user", "content": "用三句话解释什么是通货膨胀。"}],
)
print(r.usage.prompt_tokens, r.usage.completion_tokens, f"${cost(r.usage):.6f}")
# 流式:最后一个数据块带 usage,其余块的 usage 为空
usage_last = None
stream = client.chat.completions.create(
model="uncensored", max_tokens=200, stream=True,
messages=[{"role": "user", "content": "再用三句话解释什么是通货紧缩。"}],
)
for chunk in stream:
if chunk.usage:
usage_last = chunk.usage
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
print()
if usage_last:
print(usage_last.prompt_tokens, usage_last.completion_tokens, f"${cost(usage_last):.6f}")Be aware when streaming: only the last chunk has usage; in previous chunks it is empty, so you can judge with if chunk.usage. Log every reading to a log or database. Reconciling at the end of the month, checking for anomalies, and calculating per-user costs all rely on this record.
Another good habit: tag each business function and log costs together. For example, log "summarization", "customer service", and "translation" separately. At month-end, you can see which function costs the most and whether you should optimize prompts, trim history, or lower max_tokens. Without tags, you only have a total, making optimization difficult.
Three examples (assumptions stated)
Example 1: Customer service Q&A
Assumption: 800 tokens per request input (including system and knowledge snippets), 200 tokens output; 10,000 requests per day.
Cost: 800 × 0.25 ÷ 1,000,000 = $0.0002, output 200 × 1.00 ÷ 1,000,000 = $0.0002, total $0.0004. Daily $4.00, 30 days $120. At this rate, $0.50 trial credit supports ~1250 calls.
Example 2: Long article summarization
Assumption: 20,000 tokens input per article, 600 tokens output for the summary; 500 articles in total.
Per article: 20,000 × 0.25 ÷ 1,000,000 = $0.005. Output: 600 × 1.00 ÷ 1,000,000 = $0.0006. Total: $0.0056. 500 articles = $2.80. Long input, short output tasks are very cheap.
Example 3: Multi-turn chat and history pruning
Assumption: 100 tokens for system; 60 tokens user input per turn, 150 tokens reply per turn; 20 turns per conversation.
| Approach | Cumulative input tokens | Cumulative output tokens | Total cost |
|---|---|---|---|
| With full history each time | 43,100 | 3,000 | Approx. $0.0138 |
| Only last 3 turns | 14,540 | 3,000 | Approx. $0.0066 |
Full history input: 20 × (100 + 60) + 210 × (0 + 1 + … + 19) = 3,200 + 39,900 = 43,100. Pruned approach: rounds 1-3 are 160, 370, 580; rounds 4-20 are 790 each. Total: 14,540. Cost difference is ~50%, widening with more rounds because full history input grows quadratically.
From these examples, we derive a rule: cost is driven by request count × input/output volume. History and attached docs are the biggest variables. Reduce knowledge snippet length for support, prune history for chat, and chunk-parallelize for summarization. Don't pay for unnecessary tokens. Note: these numbers are based on the stated assumptions; use actual usage for your data.
Reverse calc: $30 monthly budget for 20-turn chats. With 'last 3 turns' plan, ~$0.0066/call → ~4545 calls. With full history, $0.0138/call → enough for ~2174 calls. This is the real difference.
Budget control: installing a gate in your program
Estimation is just a forecast; the gate is the insurance. Below is a daily budget class that tracks costs in USD: predict using the “worst-case scenario” (output uses max_tokens) before the request, and record actual usage after the request.
class Budget:
"""按美元计的日预算。超出时拒绝新请求。"""
def __init__(self, daily_usd):
self.limit = daily_usd
self.spent = 0.0
def check(self, est_prompt_tokens, max_tokens):
# 最坏情况:输出用满 max_tokens
worst = (est_prompt_tokens * 0.25 + max_tokens * 1.00) / 1_000_000
if self.spent + worst > self.limit:
raise RuntimeError(f"预算不足:已用 ${self.spent:.4f},本次最坏 ${worst:.4f},上限 ${self.limit}")
def record(self, usage):
self.spent += (usage.prompt_tokens * 0.25 + usage.completion_tokens * 1.00) / 1_000_000
budget = Budget(daily_usd=5.0)
budget.check(est_prompt_tokens=1200, max_tokens=500) # 请求前
# ……发请求……
# budget.record(response.usage) # 请求后Besides programmatic gates, there are three configuration-level methods:
- Set max_tokens per task, don’t just set it to the maximum; the default is 2048, the maximum is 32,000;
- Set a length limit for conversation history, see the approach in Building a Chatbot;
- Only top up the amount you plan to use in your account, and add more when it runs out. The prepaid credit model naturally provides a hard upper limit.
Unit prices and balance rules are subject to the Pricing page, and API details are in Parameter Reference.
Finally, summarized into a checklist: before, estimate the scale, set max_tokens, and check history length; during, watch the usage of the last chunk when streaming; after, record usage, categorize by function, and compare with the budget. After these three steps, the account is clear.
Frequently Asked Questions
Do both input and output tokens cost money?
Yes. Input is $0.25 per million tokens, output is $1.00 per million tokens, calculated separately based on prompt_tokens and completion_tokens in the usage.
Can I calculate the exact cost in advance?
You can only estimate the scale. The exact number depends on the usage in the response. It is recommended to sample and calculate the average for each scenario to make predictions.
Does the balance expire?
The topped-up prepaid credit never expires. The $0.50 free trial credit for new accounts is valid for 7 days.
How to avoid conversations getting more expensive?
Limit the length of history included, keep only the last few turns, or compress earlier content into summaries, and set appropriate max_tokens per task.
Fill out the form to get your API key
Create an account, copy the API key, and modify the Base URL. Configuration is that simple.