LLM API FAQ: Find answers by topic
When developers first use an LLM API, they have many questions: where to get a key, image support, downtime, max length. We divide them into five groups. Answers are concise; click to jump to the section you care about.
Updated on
Key points
- Registration via email and password. The API key is displayed instantly. One account, one key, regenerable.
- Capabilities are text chat only, single model, supporting streaming and function calling. No images, audio, vectors, or fine-tuning.
- Combined prompt and output per request must not exceed 100,000 tokens. Each API key allows 300 requests per minute.
- Legal adult content is allowed, but sexual content involving minors is blocked. The service is for adults only.
Accounts & keys
How do I get my first API key?
Go to the Get API key page, register with email and password. The key is displayed after registration; no payment info is required.
Can I create multiple API keys per account?
No, one per account. To track usage per project, tag requests in your own backend instead of relying on multiple keys.
What if my API key is leaked?
Regenerate it immediately. The old key becomes invalid instantly. Deploy the new value everywhere before regenerating.
Can new users try it first?
Yes. New accounts get $0.50 in free trial credit, valid within 7 days. No payment method is required.
To see the full process from registration to the first request, follow the complete project setup in the Chatbot guide.
If you manage multiple projects, don't rely on multiple API keys for isolation. Tag each call in your backend and log the usage. Summarize by tag at month-end for clearer stats than console granularity.
Capabilities & limits
What is the endpoint and model name?
The endpoint is https://api.apidamoxing.com/v1,模型名固定为 uncensored. The service offers only this model.
Which endpoints are provided?
Two: POST /v1/chat/completions for chat, and GET /v1/models to list models. Other paths return 404.
Can I send images, audio, or do vector search?
No. We process text only. No images, audio, video, vector embeddings, or fine-tuning.
Does it support streaming and function calling?
Yes. Add stream: true to the request for streaming via SSE. Function calling uses the OpenAI tools format.
Can I use the existing OpenAI SDK?
Yes. Point base_url to the endpoint above and set api_key to your API key. Works with Python, Node, or curl.
A common misconception is that OpenAI format compatibility means full feature parity. Compatibility refers to request and response formats, not capabilities. Since we only support text chat, code relying on image understanding, speech-to-text, or vector search must be adapted.
Length, frequency & parameters
How much content fits in one request?
Input plus output is limited to 100,000 tokens. The request body is also capped at 8 MB, but tokens usually hit the limit first.
Why does the response cut off mid-sentence?
You likely hit max_tokens, default 2048. Increase it; the maximum is 32,000. Ensure the total does not exceed 100,000 tokens.
How many requests per minute?
300 per key. For batch jobs, queue them on your side first to avoid hitting the rate limit.
Do temperature, top_p, and stop take effect?
Yes, these standard sampling fields are passed to the model as-is, behaving as you would expect from the chat completion endpoint.
Where can I see the details for each field?
Detailed API parameter explanation is available in the table. The documentation page provides a concise reference.
Regarding “what counts as long”: 100,000 tokens is roughly tens of thousands of Chinese characters, depending on content. This is an order-of-magnitude estimate, not an exact conversion. For very long documents, splitting them and then summarizing is a safe approach.
Pricing and balance
How is it charged?
Per token: $0.25 per million input tokens, $1.00 per million output tokens. No monthly fee, no subscription required.
Do I pay upfront or later?
Upfront. Top up prepaid credit in your account, and it is deducted as you use it. The balance never expires.
What happens when the money runs out?
Requests fail with status code 402 and error code no_credit. The same happens when free trial credit expires. Service resumes after you top up.
How do I estimate how much a feature costs?
Refer to the examples in tokens and billing. The key is to read the usage in the response and multiply by the unit price.
Another budgeting habit: top up based on your expected usage for the period. You don’t need to top up too much at once. Since it is prepaid, your balance is your hard limit; it stops when used up, avoiding surprise bills.
Deployment and troubleshooting
Can dev and production share the same key?
Since an account has only one key, they effectively share it. Keep the key in server-side environment variables, not in code repositories or frontend pages. Note that regenerating it invalidates it everywhere at once.
Should I retry immediately on a 503?
No. This means the service is temporarily busy. Wait a few seconds and retry, with a max retry limit. Bursting requests just wastes your request quota.
The endpoint is slow; is something wrong?
Long text generation takes time, and longer outputs take longer. Enable streaming so users see the first words immediately, and set a generous client-side timeout.
The conversation gets long; how do I avoid hitting the limit?
Truncate history on your backend, keeping only the most recent turns. Summarize older content. Estimate the total before sending the request; don’t wait for a 400 response to handle it.
How do I confirm how many tokens a request used?
Read the usage field in the response, which includes input, output, and total. For streaming responses, this data is in the last chunk; no extra parameters are needed.
These questions keep coming up because they happen right after “the code works.” Treat this section as a pre-deployment checklist to verify your app handles each case.
Content boundaries and users
What content is rejected?
Legal adult content, fiction, and controversial topics are not rejected. Sexual content involving minors is always blocked with a 403, including in fiction and roleplay.
Who can use it?
Adults 18 and older. If your product is public-facing, add age verification at the entry point.
Will my prompts be used for training?
No, prompts are not used for training.
Where do I look first when I see an error?
Read the error code and message in the response body, then check the status code: 401 is key issue, 402 is balance, 403 is content, 429 is rate limit, 503 is service busy. Wait a few seconds and try again.
Frequently asked questions
What do I need to call the endpoint?
A key from a registered account, plus an HTTPS request to https://api.apidamoxing.com/v1 with a Bearer token in the Header. The rest follows the chat completion format.
What features are supported and unsupported?
Supported: text chat, streaming, and function calling. Unsupported: images, audio, video, vector embeddings, and fine-tuning.
Why is the answer truncated?
Usually because max_tokens was reached (default 2048, adjustable up to 32,000). The combined total of prompt and output cannot exceed 100,000 tokens.
What happens when the balance runs out?
Requests return 402 and no_credit. Top up prepaid credit to continue; the balance never expires.
Who cannot use this service?
People under 18. Sexual content involving minors is blocked with a 403, whether fictional or not.
Fill out the form to get your key
Create an account, copy the API key, modify the Base URL. That’s simple.