Get a personalized assessment of your operational efficiency and accelerate growth for your business.
If you came here looking for GPT-4, start with this: OpenAI's deprecations page lists gpt-4, gpt-4-0613 and gpt-4-turbo for shutdown on October 23, 2026, with gpt-5.6-sol as the recommended replacement. Earlier GPT-4 snapshots went before them.
That is worth sitting with for a moment, because it is the real lesson of the last three years of building on this API: the model you pick is the most replaceable part of what you build. The prompts, the evaluation set, the failure handling and the data boundaries you put around it are what survive.
This guide covers the decisions that actually matter when a business builds on the OpenAI API: which model, what it costs, what happens to your data, how you know the thing works, and what a single well-scoped workflow looks like end to end — including what it does when it goes wrong.
Everything here was checked against OpenAI's own documentation in September 2026. Model names, prices and policies change often, so verify the specifics before you commit; the links go to the pages that stay current.
ChatGPT Is Not the API
These get conflated constantly, and the confusion is expensive.
ChatGPT is a finished product. Your team logs in, uses it, and the work stays in the seat licence. It is the right tool for drafting, research and one-off analysis, and it is the fastest way to find out whether a task is tractable at all.
The OpenAI API is a component you build into your own software. You write the prompts, you handle timeouts and bad outputs, you pay per token, and you own the security and compliance posture of the result. Nothing about it is turnkey.
The practical consequence: a successful ChatGPT pilot proves a task is possible. It tells you almost nothing about what that task costs at volume, how often it fails on your edge cases, or how it behaves when it is one step in an automated chain with nobody reading the output. Those answers come from building a bounded version and measuring it.
Choosing a Model
The current lineup on OpenAI's models page, as of September 2026:
| Model | Input / output per million tokens | Context | Where it fits |
|---|---|---|---|
GPT-6 Astra (gpt-6-astra) |
$10 / $50 | 1.05M | The hardest end-to-end work: multi-step reasoning, complex code, tasks where a wrong answer is costly |
GPT-5.6 Sol (gpt-5.6-sol) |
$4 / $20 | 1.05M | Complex professional work; the replacement OpenAI names for retiring GPT-4 deployments |
GPT-5.6 Terra (gpt-5.6-terra) |
$2 / $12 | 1.05M | The default middle ground when you want capability without flagship pricing |
GPT-5.6 Luna (gpt-5.6-luna) |
$0.20 / $1.20 | 1.05M | High-volume, cost-sensitive work: classification, routing, extraction, tagging |
Four things are worth noticing in that table.
The price spread is 50x. Astra costs fifty times what Luna does per input token. Routing every request to the most capable model is the single most common way to turn a promising pilot into an unaffordable feature. Most production systems end up mixed: a cheap model handles the volume, an expensive one handles the cases the cheap one flags.
Output costs multiples of input. On every model above, output tokens are five to six times the price of input tokens. A feature that reads a long document and returns a label is cheap. A feature that writes long prose is not.
Context windows are no longer the constraint. At roughly a million tokens, the limit on what you can put in a prompt is rarely the technical ceiling — it is the cost of the tokens and the fact that models still attend unevenly to very long inputs. Retrieving the relevant ten pages beats pasting the whole manual, on both counts.
Reasoning effort is now a dial. Current models expose reasoning levels rather than being fixed at one setting. The same model ID can be cheap and fast or slow and thorough depending on how you call it, which means "which model" and "how hard should it think" are two separate decisions.
Pick the cheapest model that passes your evaluation, not the best one on a public benchmark. Benchmarks are not your workload.
If you find yourself asking whether to move off the API entirely, that is a separate question with a real answer — we covered it in custom LLM vs API: when does renting AI stop making sense.
What It Actually Costs
Billing is per token, roughly four characters of English each. To estimate before you build:
- Take a representative request. Count the tokens in the system prompt, the retrieved context and the user input. That is your input.
- Estimate the output length you expect, honestly — models are verbose unless you constrain them.
- Multiply both by the per-million rates for your chosen model and by your monthly request volume.
- Add 15–30% for retries, failed generations and the prompt iterations you will do in the first months.
Two mechanisms change the answer materially, and both are documented:
- Prompt caching discounts repeated input. If every request carries the same long system prompt or the same reference document, cache it. The savings on a high-volume feature with a fixed preamble are substantial.
- Batch processing trades latency for cost. Work that does not need an answer in the next second — overnight enrichment, backfills, bulk classification — belongs here.
The cost failure mode to plan for is not the steady state. It is the loop: a retry that retries, an agent that calls itself, a user who pastes a book. Set hard spend limits per project, cap output tokens per request, and alert on cost per request rather than only on the monthly total.
Data Handling and Compliance
This is the section that decides whether legal signs off, so get the facts from the source. OpenAI's data controls documentation states:
- Data sent to the API is not used to train or improve OpenAI models, as of March 1, 2023, unless you explicitly opt in.
- Abuse monitoring logs may contain prompts and responses and are retained for up to 30 days by default, longer only where law or harm prevention requires it.
- Zero Data Retention and Modified Abuse Monitoring exclude your content from those logs. Both require prior approval from OpenAI and additional terms, and Zero Data Retention changes some endpoint behaviour — stored state is disabled even if your code requests it.
- PHI requires an executed Business Associate and Healthcare Addendum, and the endpoint limitations documented on that page still apply.
Before the first production call, answer these in writing:
- What is the most sensitive field that can reach the prompt, including through retrieved context and user-pasted text? Most leaks come from retrieval, not from the code path you designed.
- Do you need to redact or tokenize before sending, and who verifies that the redaction works?
- Which of your own logs capture prompts and outputs? Your observability stack is a data store too, and it is usually the one nobody audited.
- What is the contractual position: standard terms, a BAA, a data processing agreement, or an approved retention control?
- If you operate under GDPR, where is processing happening and what does your record of processing activities say about it?
None of this is exotic, but it has to be decided before launch rather than during a security review.
Evaluation: How You Know It Works
Most AI features that fail in production did not fail suddenly. Nobody was measuring.
Build the evaluation before you tune the prompt:
- Collect real cases. Fifty to a few hundred examples from your actual business, covering the boring majority and the edge cases you already know hurt.
- Label them. What is the correct output for each? If your own experts disagree on twenty percent of cases, you have learned something important about the feature before writing any code.
- Define wrong. A wrong answer in a support draft that a human reviews is an inconvenience. A wrong answer in an automated refund decision is money. Grade against consequence, not against an abstract score.
- Measure, then change one thing. Prompt changes that feel like improvements frequently are not. The test set is what tells you.
- Re-run on every change — new prompt, new retrieval, new model version, new reasoning level. This is exactly what makes a forced migration like the GPT-4 shutdown survivable.
A note on tooling: OpenAI is retiring its own Evals platform, which becomes read-only on October 31, 2026 and shuts down on November 30, 2026 per the evals documentation. Do not build your evaluation process around a product that is being withdrawn. Keep the labelled cases and the grading logic in your own repository, in whatever form your team will actually maintain, and treat any vendor's harness as replaceable.
One Bounded Workflow, End to End
Abstract advice is easy. Here is a single workflow, scoped the way we would scope it for a client.
The workflow: drafting first-response emails for inbound support tickets.
Scope it tightly. The model drafts a reply; an agent reviews and sends. It does not send anything itself, it does not issue refunds, and it does not touch tickets flagged as complaints, legal or billing disputes. That boundary is the design.
How it runs
- A ticket arrives. The system classifies it with the cheap model (
gpt-5.6-luna) into a category and a confidence score. - Categories outside the allowed set, or confidence below the threshold, go straight to a human queue untouched.
- For the rest, retrieve the three most relevant help-centre articles and the customer's last two interactions.
- Call the mid-tier model (
gpt-5.6-terra) with a system prompt that sets tone, the retrieved context, the ticket, and instructions to return structured output: a draft, the article IDs it relied on, and a self-reported confidence. - Present the draft to the agent inside the existing helpdesk, with the cited articles visible. The agent edits and sends, or discards.
- Log the ticket, the draft, the agent's final text and whether the draft was used.
What it does when things go wrong
- API timeout or error: retry twice with exponential backoff, then fall back to the blank composer. The agent never waits on the model; the queue never stalls.
- Malformed or unparseable output: one retry with a stricter instruction, then fall back. Never ship a parse failure into the UI.
- Model cites an article that does not exist: validate every article ID against your help centre before rendering. Any invalid citation discards the draft, and increments a counter you watch.
- Cost spike: hard cap on output tokens per draft, per-project spend limit, alert on cost per ticket crossing a threshold.
- Quality drift: the edit-distance between draft and sent reply is your live quality metric. When it climbs, something changed — a model update, a product change, a new class of ticket — and you re-run the evaluation set.
- Outage of the whole feature: a config flag turns drafting off without a deploy. Support carries on exactly as before, because the workflow was additive.
What you measure
Draft acceptance rate. Median edit distance. Handle time against a control group of agents without the feature. Cost per ticket. Escalations after an AI-drafted reply. That last one is the number that tells you whether you saved time or moved the work downstream.
How you roll it out
One team, two weeks, shadow mode first — generate drafts, show them to nobody, have a supervisor grade a sample. Then a small group of agents. Then the rest, if the numbers hold. If acceptance sits below roughly half, the problem is usually the retrieved context, not the model.
That is the shape of a workflow worth building: narrow, reversible, measured, with a human where the consequence is. Most of the disappointing AI projects we get called in to fix were none of those things.
Where the API Fits Well — and Where It Does Not
Patterns that tend to work, because the output is checkable and the failure is cheap:
- Classification and routing. Tickets, leads, documents, transactions. Cheap models, high volume, measurable accuracy.
- Extraction. Pulling structured fields out of unstructured documents — invoices, contracts, forms — with validation against your schema.
- Drafting with a reviewer. Replies, summaries, descriptions, first-pass translations, where a person approves before anything leaves the building.
- Search over your own content. Retrieval plus generation, with citations the user can click. The citation is the safety mechanism.
- Developer assistance. Code review suggestions and test generation, inside a workflow that already has review gates.
Patterns that tend to disappoint:
- Anything where you cannot check the answer. If nobody can tell a good output from a plausible wrong one, you have not built a feature, you have built a liability.
- Precise numerical work. Arithmetic, reconciliation, forecasting. Use the model to pull the numbers out and a deterministic system to compute them.
- Regulated decisions without a human. Credit, clinical, hiring, insurance. The regulation does not care how good the model is.
- Open-ended autonomy. Long agent chains compound their own errors; each step's mistakes become the next step's input. Bound the steps and check between them.
- Replacing a broken process. Automating a workflow nobody understands produces a faster version of the same confusion.
If your integration target is an older system, the constraint is usually the system rather than the model — we wrote about what AI integration approach works best for legacy systems separately.
A Short Checklist Before You Build
- The task is one job, with a defined input and a checkable output.
- You have a labelled test set from your own data, and a definition of "wrong".
- You have picked the cheapest model that passes it, and know what the alternative costs.
- You have costed it per unit of work, with a spend cap and an alert.
- You know what data reaches the prompt, and the contractual position covering it.
- There is a human wherever the consequence is real.
- There is a rollback that does not require a deploy.
- Someone owns the feature after launch, including re-running the evaluation when the model changes.
If you cannot tick all eight, the gap is your next piece of work — not a bigger model.
Building This With Imaginovation
We build and integrate AI features into production software, and a fair amount of our work is fixing implementations that skipped the parts above: no evaluation, no cost controls, no boundary around what the model is allowed to decide.
If you have a workflow in mind and want an honest read on whether it is worth building — including when the answer is no — talk to us. We would rather scope one workflow that works than ten that demo well.




