Build a chatbot with the LLM API: FastAPI backend with streaming frontend
To build a chatbot that types out character by character, you only need three components: a backend to forward requests, a web page to read the stream, and a storage to remember the conversation. This guide uses FastAPI for the backend, browser-native fetch for streaming, and stores history in your own SQLite. The whole process takes less than 100 lines of code, suitable for an adult-oriented product prototype.
Updated on
Key points
- Keep the API key only in the backend; the browser never touches it. The frontend only communicates with your own /chat endpoint.
- The model won't remember conversations for you: each request retrieves the most recent history from your storage and sends it to the endpoint.
- The key to streaming output is the backend forwarding with StreamingResponse, and the frontend looping to concatenate with reader.read().
- For an adult-oriented product, implement age verification at the entry point and limit history length and single input length.
First, draw a small diagram: three roles, each handling their own part
Think of the chatbot as a takeout shop: the browser is the customer, your backend is the front desk, and the LLM endpoint is the kitchen. Customers never enter the kitchen directly; all orders are forwarded by the front desk. This is why the API key cannot be placed in the frontend.
| Role | Responsible for | Must not |
|---|---|---|
| Browser | Display, collect input, read stream | Hold the API key |
| FastAPI backend | Save history, assemble messages, forward requests | Return the key to the frontend |
| Model endpoint | Generate replies based on messages | —it only processes what you send it |
A conversation flow is: the web page sends a message, the backend saves it, retrieves recent history, appends the system prompt, requests the endpoint, forwards the stream to the web page as it arrives, and saves the full reply after completion.
Why choose backend forwarding instead of direct frontend connection? Besides key security, there are two practical reasons: first, you need to add history, enforce rate limits, and filter during forwarding, which can only be done on the backend; second, when switching models or services in the future, you only need to change one place in the backend, and the frontend remains untouched.
Backend: FastAPI with SQLite
Install dependencies and start first:
pip install fastapi uvicorn openai
export API_KEY=你的密钥
uvicorn server:app --reload --port 8000Then write server.py. There are four points in the code worth pausing to look at:
# server.py
import os
import sqlite3
import uuid
from fastapi import FastAPI
from fastapi.responses import FileResponse, StreamingResponse
from pydantic import BaseModel
from openai import AsyncOpenAI
app = FastAPI()
client = AsyncOpenAI(base_url="https://api.apidamoxing.com/v1", api_key=os.environ["API_KEY"])
SYSTEM = {"role": "system", "content": "你是『小墨』,一个说话简洁、爱用比喻的聊天伙伴。回答控制在 200 字内。"}
KEEP = 20 # 每次只带最近 20 条消息
db = sqlite3.connect("chat.db", check_same_thread=False)
db.execute("create table if not exists msg(id integer primary key autoincrement, sid text, role text, content text)")
def load(sid, n=KEEP):
rows = db.execute("select role, content from msg where sid=? order by id desc limit ?", (sid, n)).fetchall()
return [{"role": r, "content": c} for r, c in reversed(rows)]
def save(sid, role, content):
db.execute("insert into msg(sid, role, content) values (?,?,?)", (sid, role, content))
db.commit()
class ChatIn(BaseModel):
session_id: str
message: str
@app.get("/")
def index():
return FileResponse("index.html")
@app.post("/session")
def new_session():
return {"session_id": uuid.uuid4().hex}
@app.post("/chat")
async def chat(body: ChatIn):
save(body.session_id, "user", body.message[:4000])
messages = [SYSTEM] + load(body.session_id)
async def gen():
parts = []
try:
stream = await client.chat.completions.create(
model="uncensored", messages=messages,
stream=True, max_tokens=800, temperature=0.8,
)
async for chunk in stream:
if not chunk.choices: # 末尾的 usage 块没有 choices
continue
delta = chunk.choices[0].delta.content
if delta:
parts.append(delta)
yield delta
except Exception as e:
yield f"\n[请求失败:{type(e).__name__}]"
finally:
if parts:
save(body.session_id, "assistant", "".join(parts))
return StreamingResponse(gen(), media_type="text/plain; charset=utf-8")load()only takes the most recent 20 items to prevent the conversation from growing too long and exceeding the 100,000 token context window.gen()is an async generator that yields content as it arrives, allowing the frontend to display it character by character.- The end of the stream contains a data block with usage stats but no choices, so check for emptiness and skip it.
- In
finally, save the full reply so that even if an error occurs mid-way, the generated part is not lost.
If you want to skip the database for now, replace load and save with dictionary reads and writes; the rest of the code remains unchanged. Conversely, when traffic increases, you can replace SQLite with any database you are familiar with, as long as the interface maintains the shape of these two functions. Note: the example uses synchronous sqlite3, which is fast enough for prototypes; consider async drivers for high parallel requests scenarios.
Frontend: reading the stream with fetch
No libraries needed. resp.body.getReader() gives a reader; loop read(), decode chunks with TextDecoder, and append to the page. Save as index.html alongside server.py:
<!doctype html>
<meta charset="utf-8">
<title>小墨</title>
<div id="gate">
<label><input type="checkbox" id="adult"> 我已年满 18 周岁</label>
<button id="enter">进入</button>
</div>
<div id="app" hidden>
<div id="log" style="white-space:pre-wrap;min-height:300px"></div>
<input id="box" placeholder="说点什么"> <button id="send">发送</button>
</div>
<script>
let sid = localStorage.getItem("sid");
const log = document.getElementById("log");
document.getElementById("enter").onclick = async () => {
if (!document.getElementById("adult").checked) return;
if (!sid) {
const r = await fetch("/session", { method: "POST" });
sid = (await r.json()).session_id;
localStorage.setItem("sid", sid);
}
document.getElementById("gate").hidden = true;
document.getElementById("app").hidden = false;
};
document.getElementById("send").onclick = async () => {
const box = document.getElementById("box");
const text = box.value.trim();
if (!text) return;
box.value = "";
log.textContent += "\n你:" + text + "\n小墨:";
const resp = await fetch("/chat", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ session_id: sid, message: text }),
});
const reader = resp.body.getReader();
const decoder = new TextDecoder("utf-8");
while (true) {
const { done, value } = await reader.read();
if (done) break;
log.textContent += decoder.decode(value, { stream: true });
}
log.textContent += "\n";
};
</script>Two details: add { stream: true } during decoding because the bytes of a single Chinese character might be split across two chunks; without it, you will see garbled text. Second, the send logic in the input box should disable the button during the request to prevent double-clicking. The example omits this for brevity; add it before launch.
After running, open your browser to port 8000, check the age verification box, and send a message. If the text appears as a whole block instead of character by character, an intermediate proxy is likely buffering the response. Try connecting directly to the local port to rule this out. When deploying behind Nginx, you need to disable response buffering for this path.
Conversation history: why store it on your side
The chat endpoint is stateless, like a receptionist who only knows what is on the note you hand them. So, remembering context is your responsibility. There are three levels of implementation:
- Simplest:Store in a dictionary in memory only. Lost on restart; suitable for debugging.
- Common:Store in SQLite as shown in the example, one row per message, querying the most recent N messages by session_id.
- Advanced:When history gets too long, have the model summarize older content into a single summary paragraph, place it after the system prompt, and keep recent messages as raw text.
Another benefit of storing on your own backend is that you fully control the retention period. We recommend providing a "Clear Chat" button to truly delete the session's records, and clearly stating what you save in your privacy policy.
The trigger condition for summarization can be simple: when the estimated history exceeds 10,000 tokens, hand the oldest half of the messages to the model to summarize into a paragraph of no more than 300 words, write it back to storage, and delete the original messages. This preserves context while keeping the volume of each request stable and manageable. Note that summaries are model-generated and may omit details; important settings (such as the user's nickname or forbidden topics) are better fixed in the system prompt rather than relying on the summary to retain them.
Adult-oriented products: Entry and boundaries
This service is for users aged 18 and over, and your application should be too. The example frontend includes a simple confirmation entry; more rigorous products can add a more complete age verification flow. Here are some best practices:
- Clearly indicate the age limit on the entry page; do not show the chat interface until confirmed.
- Do not market the product to students or minors, and do not place it in scenarios intended for them.
- Sexual content involving minors will be intercepted by the API and return a 403 status, whether fictional or not. Your backend should detect this status and show the user a friendly message instead of the raw error.
From a product design perspective, you can also let users set a nickname and tone preference in settings. Write these into the system prompt to make the conversation more personalized without forcing the model to guess.
Pre-launch hardening and extensions
- Limit input:The example truncates individual messages to 4,000 characters; adjust as needed.
- Limit frequency:Cap the key at 300 requests per minute; implement rate limiting in the backend by user or IP.
- Error display:The backend catches exceptions and returns a short message to the frontend; do not expose the stack trace.
- Usage monitoring:To count tokens, read usage in non-streaming requests, or read from the last chunk of a stream. See Tokens and billing for details.
- Switching frameworks:When switching the frontend to Vue or React, the logic for reading streams remains identical. Switching the backend to Express just means replacing StreamingResponse with the corresponding streaming implementation.
For more complete parameter documentation, see API parameter details; for more questions, see FAQ.
Do a final pre-launch self-check: Is the key only in server-side environment variables? Are there upper limits on input length and history length? Are errors displayed with friendly messages? Does "Clear Chat" actually delete records? Is age confirmation at the entry point? If all five are met, this prototype is ready for real users to try.
Frequently Asked Questions
Why not let the browser call the API directly?
Because the key would be exposed in the web page, allowing anyone to copy and misuse it. Having the browser access only your own backend, with the backend holding the key, is the more secure approach.
How many conversation history messages should be stored?
Depending on the scenario, the example takes the most recent 20 items. When controlling costs and context length, it is better to include fewer items and use summarization to retain key points.
What if the streaming output fails midway?
The backend catches the exception, outputs a prompt to the frontend, and saves the generated portion. The frontend can provide a retry button to resend the last user message.
Can this bot be used by minors?
No. The service is limited to adults aged 18 and over; your application must also perform age confirmation at the entry point.
Just fill out the form to get your key
Create an account, copy the key, and modify the Base URL. That's how simple the configuration is.