Most teams don’t need another explanation of how large language models work. They need a feature that runs by Friday. So this is a build guide: make a real call, choose a model that won’t wreck your budget, force clean JSON out of it, stream the result to a user, and survive the rate limits and timeouts that appear the moment actual traffic shows up.
Everything below runs on a laptop. Grab an editor and follow along.
Your first call takes about ten minutes
Create a key in the OpenAI dashboard, then keep it out of your repository from day one. Export it as an environment variable and read it in code, never inline.
pip install openai
export OPENAI_API_KEY="sk-..."
A minimal request looks like this:
import os
from openai import OpenAI
client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])
resp = client.chat.completions.create(
model="gpt-4o-mini",
messages=[
{"role": "system", "content": "You are a concise technical editor."},
{"role": "user", "content": "Summarise this bug report in one sentence: ..."},
],
temperature=0.2,
)
print(resp.choices[0].message.content)
print(resp.usage)
That usage object is the part people skip and later regret. It reports prompt tokens, completion tokens and the total, which is exactly what you need to attribute spend to a feature. If account setup and key rotation are new territory, the walkthrough in getting from an API key to a working bot covers that side in more detail.
Pick the model after you understand the task
Model choice is mostly a cost decision, not a quality one. At the time of writing, a small model like gpt-4o-mini handles the vast majority of classification, extraction and drafting work at a fraction of flagship pricing. The bigger models earn their keep on tasks with multiple hard constraints, long documents, or anything that genuinely resembles reasoning.
- Classification, tagging, routing: gpt-4o-mini. You will run thousands of these.
- Drafting and rewriting: gpt-4o-mini, plus a sharp system prompt and two or three examples.
- Multi-step analysis or tight constraints: a frontier model, or a reasoning model if the logic has to hold up.
The practical rule: build on the cheap model, keep a list of twenty hard inputs, and only upgrade when the cheap model visibly fails them. Our breakdown of what the OpenAI API costs in 2025 has current per-million-token figures if you’re putting a budget together.
Write prompts like an API contract
Put the rules in the system message
Tone, audience, length limits and what to do when information is missing all belong there. Keep that block stable across requests so you can cache it and measure prompt changes against a fixed baseline.
Two good examples beat ten adjectives
If you want short answers, don’t write “be concise” five times. Include one ideal answer and one that misses the mark. Models imitate patterns far more reliably than they follow descriptions of patterns.
Demand JSON you can parse
Free-text output means regex and regret. Structured outputs let the schema do the enforcing:
resp = client.chat.completions.create(
model="gpt-4o-mini",
messages=messages,
response_format={
"type": "json_schema",
"json_schema": {
"name": "ticket",
"strict": True,
"schema": {
"type": "object",
"properties": {
"category": {"type": "string",
"enum": ["billing", "bug", "feature", "other"]},
"priority": {"type": "integer", "minimum": 1, "maximum": 5},
"summary": {"type": "string"}
},
"required": ["category", "priority", "summary"],
"additionalProperties": False
}
}
}
)
With strict: True, the response either matches your schema or the request fails loudly. Loud failure during development is a gift.
Stream anything a human is waiting for
A 400-token answer can take six or seven seconds to arrive in full. Streamed, the first words land in under a second and the perceived wait collapses. Set stream=True and iterate the chunks:
stream = client.chat.completions.create(
model="gpt-4o-mini",
messages=messages,
stream=True,
)
for chunk in stream:
delta = chunk.choices[0].delta.content
if delta:
print(delta, end="", flush=True)
One caveat: streaming and structured JSON rarely mix well in a UI, because the user watches half a JSON object crawl across the screen. Stream prose to people, request structured output for machine-to-machine calls.
The failures that show up in week two
- 429 rate limits. Read the
retry-afterheader and back off exponentially with jitter. Never retry in a tight loop. - Context length errors. Trim old conversation turns instead of resending the full history every time.
- Timeouts. Set an explicit client timeout and a fallback response. Users prefer a canned apology to a spinner that never ends.
- Prompt injection. Anything a user can type is untrusted input. Pass it inside a delimited block as data, never as instructions.
- Silent quality drift. Pin model versions where you can, and re-run your test set after every change, including the ones you think are cosmetic.
Keeping the invoice boring
Four habits do most of the work. Cache identical requests with a 24-hour TTL, since support tools see the same twenty questions all day. Use the Batch API for anything that doesn’t need an immediate answer, which halves cost in exchange for a 24-hour turnaround. Cap max_tokens so a runaway generation can’t drain a budget. And log token usage per endpoint, tagged by feature, so you can see which one is actually expensive.
Teams pushing this into a support desk or warehouse workflow tend to hit the same wall: costs look fine in testing and shocking in production. A staged 90-day rollout plan with real examples is a useful template for that, even if your tooling looks nothing like theirs.
Build one thing end to end
Take a single annoyance and automate only that. Ticket triage makes a good first project because it needs everything: a schema, a system prompt, a model choice and a fallback path.
def triage(message: str) -> dict:
resp = client.chat.completions.create(
model="gpt-4o-mini",
temperature=0,
messages=[
{"role": "system",
"content": "You triage support tickets. Return JSON only."},
{"role": "user",
"content": f"Ticket:\n<<<\n{message}\n>>>"},
],
response_format={"type": "json_schema", "json_schema": TICKET_SCHEMA},
)
return json.loads(resp.choices[0].message.content)
Forty lines plus a schema. From there you can add a suggested reply, route by category, or push anything marked high priority into Slack. Once the shape works, function calling lets that same endpoint check an order status or open a ticket in your own systems, which is the point where these projects start paying for themselves.
Then write down twenty inputs whose correct answers you already know, and run them after every prompt edit. That test set is the difference between a demo and something you can leave running overnight. If you want to build that habit properly, this 30-day plan with real projects is a sensible way to stack the skills. Do that for a fortnight and you’ll ship API features with the same confidence as any other part of your stack.

