Open a fresh browser tab, type “write a blog post about our new pricing,” and you’ll get back something that reads like a brochure drafted by a committee. The model isn’t the problem. The instruction is.
What follows is the workflow I use to get output I can actually ship: five moves, each one illustrated with the prompt I’d literally type. No “unlock the power of AI” language. Just the steps, in the order I do them.
1. Match the model to the job before you type anything
Most frustration starts here. People reach for the biggest available model for everything, wait longer, pay more, and get a result no better than a smaller one would have produced. Choose by task type:
- Mechanical reformatting — converting a CSV into JSON, rewriting 30 product blurbs to a fixed template, extracting every date from a transcript. Use the small, fast model. The output is close to identical and the cost is a fraction.
- Nuanced writing — tone, negotiation emails, summarising a messy 40-message thread without losing who said what. This is where the flagship model earns its price.
- Multi-step reasoning — reconciling two spreadsheets, reading a stack trace, planning a six-week migration. Use a reasoning model that works through the problem before answering.
One habit pays for itself immediately: run the same prompt on two models side by side. Ninety seconds of comparison tells you more than any leaderboard, because your tasks are not the benchmark tasks.
2. Use a five-slot prompt skeleton
Long prompts aren’t better. Structured prompts are. Every prompt I keep answers five questions, in this order:
- Role — who is answering?
- Task — what single thing must they produce?
- Context — what do they need to know that they can’t guess?
- Format — what does the finished output look like, exactly?
- Constraints — what should they avoid, and how long should it be?
The difference in practice is stark. Vague version: Summarise these meeting notes.
Skeleton version: You are a project manager writing for a client who wasn’t in the room. Turn the notes below into a 200-word update with three sections: Decisions, Open Questions, Next Steps. Next Steps needs an owner and a date for each item. Don’t invent owners or dates. If the notes don’t say, write TBC. Plain English, no jargon.
The second prompt takes forty seconds to write and saves an hour of editing. The line about not inventing owners matters more than any clever phrasing, because models fill gaps silently when you leave them open.
A quick test for weak prompts
If you can’t tell whether the output is correct without re-reading the whole thing, your prompt was missing a format or a constraint. Add one, don’t add ten paragraphs.
3. Chain three short prompts instead of one enormous one
Big requests degrade. Ask for research, drafting and formatting in a single message and quality drops across all three, because the model is juggling competing objectives. Split the work. Here’s a chain I run when I need a summary that people will actually trust:
- Step one, extract only. “List every factual claim in this document as a bullet. Add the paragraph number where each one appears. Change nothing, add nothing.”
- Step two, reorganise. “Group these bullets into themes. Flag any two bullets that contradict each other.”
- Step three, write. “Using only the grouped bullets, draft a 250-word summary for a non-technical reader.”
Each step’s output becomes the next step’s input, and you get to inspect it in between. That middle checkpoint is the whole point. If step one hallucinated a claim, you catch it before it becomes the spine of your final draft.
4. Give it your real material, not a description of it
Generic output comes from generic input. Paste the actual thing: the transcript, the contract, the support tickets, the 90-row spreadsheet. Two details that change results more than most people expect.
First, placement. When you’re pasting a long document, put your instructions after the text, not before. Instructions sitting at the end of a long context get followed more reliably.
Second, persistence. If you find yourself pasting the same brand guidelines, tone notes and formatting rules every day, move them into custom instructions or a saved project so they apply automatically. That single change removes maybe 200 words of repetitive typing per prompt.
5. Recognise when ChatGPT-4 is the wrong tool entirely
This step saves more time than the other four combined, because it stops you forcing a good tool into a bad fit. Three signals that you should look elsewhere:
- The data is sensitive. Client contracts, HR records, patient notes, unreleased financials. Check what your plan allows before pasting.
- The task runs thousands of times a day. Per-call pricing and latency both bite at volume.
- You need it to work offline or inside your own network. A hosted API is a non-starter.
In those cases, running a model on your own hardware is genuinely practical now. Ollama handles the download and serving side in a couple of commands, and a self-hosted interface like Open WebUI gives you something that feels like ChatGPT but runs on your machine. One team I know swapped their hosted model out of a nightly build-check job for a small local one and the pipeline stopped failing entirely, mostly because the local model was fast, deterministic, and didn’t rate-limit at 3am.
6. Verify before you ship, every single time
Fluent output is not accurate output. Those are two different qualities and the model is excellent at only one of them. Cheap checks that catch most errors:
- Ask it to quote the exact sentence from your source that supports each claim.
- Start a new chat and ask the same question cold. If the two answers disagree, go back to the source.
- Read numbers, names and dates character by character. That’s where errors hide, and they hide well.
The failure mode to watch for isn’t a spectacular wrong answer. It’s the slow drift where you stop checking because the last twenty outputs were fine. That pattern, where convenience quietly replaces judgement, is more common than most people assume, and it costs far more than a bad paragraph.
7. Keep a prompt library you’ll actually reuse
The difference between people who get consistently good results and people who don’t usually isn’t talent. It’s whether they save what worked. Keep a plain text file. When a prompt produces output you barely had to edit, paste it in with a note about what made it work.
Worth storing alongside each prompt: the job it’s for, the model that ran it best, the version of the output you approved, and any constraint you added after a failure. Review the file every month or two and delete anything you haven’t touched. Three well-tuned prompts beat forty half-remembered ones.
The first prompt you save pays for itself the second time you need it.

