LLM API Parameters Explained: From Request to Response Fields
First-time readers of the chat completion API docs often get intimidated by the long list of parameters. Think of a call like ordering at a restaurant: messages are your conversation with the waiter, temperature is how much you want the chef to improvise, max_tokens is the maximum portion size, and tools let the chef call the kitchen for ingredients. We’ll break down every field using this approach, one table at a time.
Updated on
Key Takeaways
- messages consist of system, user, assistant, and tool roles. The model has no memory; you must provide the history yourself.
- temperature controls randomness, top_p controls the candidate range. Their effects are similar, so usually only one is adjusted.
- finish_reason determines how to handle the result: stop means normal completion, length means truncation, and tool_calls means you need to execute functions.
- The usage field is the only reliable basis for billing and budgeting; read it every time.
What Does a Request Look Like?
The endpoint is POST https://api.apidamoxing.com/v1/chat/completions, with auth header Authorization: Bearer <key>. The request body is JSON, compatible with OpenAI’s chat completion format. Here is a complete request using common fields:
curl https://api.apidamoxing.com/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [
{"role": "system", "content": "你是一位耐心的天文科普作者。"},
{"role": "user", "content": "为什么月亮总是同一面朝向地球?"}
],
"temperature": 0.7,
"top_p": 0.9,
"max_tokens": 400,
"stop": ["###"]
}'Note that model must be uncensored because there is only one model available. You can verify this with GET /v1/models. Below, we explain each field individually.
You will find these parameters fall into three categories: what to say (messages, tools), how to say it (temperature, top_p), and how much/how long (max_tokens, stop, stream). Remember this classification; when you encounter unfamiliar fields, identifying the category helps you guess their purpose.
messages: The Conversation Log
messages is an array with role and content. Like a meeting log, the model reads from start to write the next page. It has no memory, so you must provide full history for multi-turn conversations.
| role | Author | Purpose |
|---|---|---|
| system | You (Developer) | Sets identity, rules, and output format; usually placed first |
| user | End User | Questions or instructions |
| assistant | Model (or your historical entries) | Previous answers, used for multi-turn context |
| tool | Your Program | Function call execution results; requires tool_call_id |
A common trick: to make the model continue writing in a certain tone, construct an assistant message and add it to the history. Remember that the combined prompt and output must not exceed 100,000 tokens; longer history takes up more space.
For example, you build a customer service bot. The system prompt says "Only answer order-related questions, max three sentences." Round 1: user asks about shipping time. Round 2: user asks "What about returns?" The Round 2 messages list must include: system, Round 1 user, Round 1 assistant response, Round 2 user. If any is missing, the model won't know what "that" refers to.
Sampling Parameters: Tuning the “Room to Improvise”
When generating each token, the model assigns probability scores to all candidates and then samples. The following two parameters adjust the sampling rules:
| Parameter | Analogy | How to Understand | Common Values |
|---|---|---|---|
| temperature | Creativity of the chef | Lower values are more conservative and stable; higher values are more divergent. | 0 to 1.2; use lower for Q&A, higher for creative tasks |
| top_p | Sample only from the top candidates | Sample only from candidates whose cumulative probability reaches p | 0.8 to 1; the default is usually sufficient |
Both control "how random" the output is, just from different angles. It is hard to tell which parameter caused an effect when both are adjusted, so we recommend changing only one at a time. These standard sampling fields are passed through as-is; the API does not rewrite them.
Another intuitive example: assume the model's top three likely next tokens have probabilities of 60%, 30%, and 10%. Lowering temperature makes the 60% option dominate, making outputs always pick the top choice; raising it brings them closer, making rare tokens more likely. With top_p at 0.9, only the top two tokens (cumulative 90%) are kept; the third is excluded. These numbers are hypothetical examples, not real probabilities.
Control length and stopping: max_tokens, stop, stream
| Parameter | Function | Key points |
|---|---|---|
| max_tokens | Limits the maximum number of tokens generated in this request | Default is 2048, maximum per request is 32,000; combined with the prompt, it must not exceed 100,000 |
| stop | Stops when a specified string is encountered | Accepts a string array; suitable for segmentation or truncating fixed formats |
| stream | Whether to return streaming responses | Set to true to push data in chunks via SSE; a final chunk with usage is automatically appended. |
max_tokens is like the size of a serving plate. If the plate is too small, the dish is taken away unfinished, and finish_reason will be length. stop is like a signal; the chef stops when they hear it. stream does not change the content, only the delivery method: from "serve everything at once" to "serve bite by bite," making the response feel faster to the user.
stop has a very practical use: if you ask the model to output in a fixed format like "Question: ... Answer: ...###", and set ### as the stop string, generation stops automatically at the delimiter. This saves tokens and prevents the model from adding unnecessary text afterwards. Note that when stop is triggered, finish_reason is also stop; if you need to distinguish the cause, you must check the content yourself.
tools and tool_choice: Let the model "call you"
The model itself cannot access real-time data. Function calling works by: you first tell the model which functions are available. When the model decides a function is needed, it does not answer directly but returns "Call function X with parameters..." Your program executes it and returns the result; the model then organizes the answer based on that. The format is consistent with OpenAI.
| Parameter | Value | Description |
|---|---|---|
| tools | Array of function descriptions | Each contains name, description, and parameters in JSON Schema format |
| tool_choice | "auto" / "none" / specific function | auto lets the model decide; none disables calling; specifying a function forces it to be called |
import json
import os
from openai import OpenAI
client = OpenAI(base_url="https://api.apidamoxing.com/v1", api_key=os.environ["API_KEY"])
tools = [{
"type": "function",
"function": {
"name": "get_tide_time",
"description": "查询某港口今天的高潮时刻",
"parameters": {
"type": "object",
"properties": {"port": {"type": "string", "description": "港口名称"}},
"required": ["port"],
},
},
}]
messages = [{"role": "user", "content": "青岛今天几点涨潮?"}]
first = client.chat.completions.create(
model="uncensored", messages=messages, tools=tools, tool_choice="auto"
)
msg = first.choices[0].message
if msg.tool_calls:
call = msg.tool_calls[0]
args = json.loads(call.function.arguments)
result = {"port": args["port"], "high_tide": "14:20"} # 这里换成你自己的查询
messages.append(msg)
messages.append({"role": "tool", "tool_call_id": call.id, "content": json.dumps(result, ensure_ascii=False)})
final = client.chat.completions.create(model="uncensored", messages=messages, tools=tools)
print(final.choices[0].message.content)
else:
print(msg.content)The process involves two rounds: first, get tool_calls; second, append the role: tool result and request again. The more specific the description, the better the model knows when to call. The parameter is a JSON string; always parse it with json.loads and validate it; do not trust it blindly.
Response fields: Understanding the returned "receipt"
A successful response looks roughly like this, with fixed fields:
{
"id": "chatcmpl-xxxx",
"object": "chat.completion",
"model": "uncensored",
"choices": [
{
"index": 0,
"message": {"role": "assistant", "content": "因为潮汐锁定……"},
"finish_reason": "stop"
}
],
"usage": {"prompt_tokens": 38, "completion_tokens": 212, "total_tokens": 250}
}| Field | Meaning |
|---|---|
| choices | Result array, usually containing only one item; the body is in choices[0].message.content |
| finish_reason | Reason for ending: stop indicates normal completion; length indicates truncation due to max_tokens; tool_calls indicates the model requested a function call |
| usage.prompt_tokens | Tokens consumed for input |
| usage.completion_tokens | Tokens consumed for output |
| usage.total_tokens | Sum of the two |
In code, check finish_reason first: if it is length, prompt the user that content is truncated or auto-continue; if it is tool_calls, go to the function execution branch. The usage field's role is detailed in the tokens and billing article for budgeting methods.
Common parameter pitfalls
- Treating max_tokens as an "input limit". It only controls output; input is controlled by you.
- Assuming the model remembers the previous request. Each request is independent; history must be provided by you.
- Setting temperature to 0 assumes every result will be exactly the same. It will be more stable, but do not assume character-by-character consistency.
- Reading choices[0].message directly in streaming mode. The fields in streaming chunks are delta; you must concatenate them yourself.
- Forgetting to append the assistant message containing tool_calls after a tool call, causing errors in the second round.
Error code meanings and retry strategies can be found in the documentation. For a complete example tying these parameters together, refer to the chatbot setup article.
Frequently Asked Questions
Can I set both temperature and top_p?
Yes, but it is difficult to isolate their effects. In most scenarios, adjusting temperature is sufficient; adjust top_p only when you need fine-grained control over the candidate range.
What should I do if finish_reason is length?
This indicates the output reached max_tokens and was truncated. You can increase max_tokens (up to 32,000) or have the model generate in segments.
Is the format of tools the same as OpenAI's?
Yes, it uses OpenAI-format tools and tool_choice, and the tool_calls in the response have the same structure.
How do I get usage in a streaming response?
When stream is enabled, a final chunk containing usage is automatically appended. You can read it directly without extra parameters.
Just fill out the form to get your key
Create an account, copy the key, and modify the Base URL. Configuration is that simple.