Close Menu
AI News TodayAI News Today

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    What's Hot

    Open, scalable training infrastructure for large MoEs

    Photon held a funeral for mobile apps. Now it has $4.5M to help replace them with agents

    Can an Apartment Search Agent Call the Model Fewer Times and Still Find Good Matches?

    Facebook X (Twitter) Instagram
    • About Us
    • Contact Us
    Facebook X (Twitter) Instagram Pinterest Vimeo
    AI News TodayAI News Today
    • Home
    • AI News
    • AI Reviews
    • AI Tools
    • AI Tutorials
    • Chatbots
    • Free AI Tools
    • Artificial Intelligence
    AI News TodayAI News Today
    Home»AI Tools»Your AI Bill Is a Toll Booth. Stop Paying Twice.
    AI Tools

    Your AI Bill Is a Toll Booth. Stop Paying Twice.

    By No Comments18 Mins Read
    Share Facebook Twitter Pinterest LinkedIn Tumblr Reddit Telegram Email
    Your AI Bill Is a Toll Booth. Stop Paying Twice.
    Share
    Facebook Twitter LinkedIn Pinterest Email

    Here’s a question worth asking your finance team: how much of this year’s AI budget is left?

    For a surprising number of companies, the honest answer is “not much”. Some are running dry by April. The annual budget, gone in a quarter.

    The obvious explanation is that everyone’s using AI more. That’s true, but it’s not the whole story. The bigger reason is how we’re using it. We’ve moved from asking chatbots questions to running agents. And agents don’t just answer. They call tools, read files, check results and call more tools. Every one of those steps burns tokens. One agentic request can cost many times what a plain chat message does.

    And the price tag won’t sit still

    It doesn’t help that the pricing itself keeps moving. There are subscriptions. There’s pay-as-you-go API access. There are promos that appear for a few months to stop people switching, and new top-tier models landing every few months with their own price points. Trying to keep up feels like reading a menu where the prices change while you’re ordering.

    So before we get into how to cut the bill, let’s make sure we all understand what we’re actually paying for. Because once you see how the meter works, the savings become obvious.

    Learn this step by step with the interactive AI Engineer roadmap.

    Surprise: you’re a token predictor too

    Let’s play a quick game. Finish these:

    • “Twinkle, twinkle, little…”

    • “Ready, set…”

    • “Salt and…”

    Star. Go. Pepper. You didn’t have to think. Your brain looked at what came before and predicted what comes next.

    Now try a lyric from a song you’ve never heard. You can’t. Not because you’re slow, but because you have no training data for it. That’s exactly how a large language model works, and exactly where it falls down.

    What happens under the hood

    An LLM predicts one token at a time. For each token, it runs a stack of maths. Here’s the catch: to predict the next token, it needs the maths for every single token before it. Then to predict the one after that, it needs all of that again, plus the new one. Round and round, until it produces a special “I’m done” token.

    Hold on to that idea. It’s the reason your bill looks the way it does.

    Tokens aren’t quite words

    I’ve been saying “token” like it means “word”. It doesn’t, quite.

    • Short, common words are often one token each.

    • Long or unusual words get chopped into pieces. “Understanding” might be two or three.

    • Spaces, punctuation, numbers and symbols all count as tokens too.

    Rule of thumb: about 1.33 tokens per word. But each model slices text its own way, so the same sentence can cost a different number of tokens depending on who you send it to.

    Welcome to the toll booth

    Picture every API call as a toll road. You pay to drive in (input tokens), and you pay to drive back out (output tokens). Output is the pricier lane.

    Take Anthropic’s lineup as an example. At the top sits Claude Fable 5 at $10 per million tokens in and $50 per million out. Fable costs exactly double what Opus does, and Opus in turn costs five times what the lightweight Haiku does. That’s a 10x spread between the cheapest and priciest model from the same company. Always check the provider’s current pricing page, because these figures shift.

    Now look for the smallest line on that price sheet. Cached tokens cost one tenth of normal tokens. Keep that number in your pocket. We’ll come back to it in a minute, and it’s worth more than everything else in this article combined.

    The conversation that never stops replaying

    Here’s the part most people don’t realise. Without caching, the model needs the whole conversation every single time.

    1. You send message one. It sends back message one plus its reply.

    2. You send message one, reply one, and message two. It sends all of that back plus reply two.

    3. You send the whole lot plus message three. It sends the whole lot back plus reply three.

    Every turn, you pay for the entire history again, in both directions. It’s like paying the toll for every mile you’ve already driven, every time you drive another mile. Nobody could afford that for long. Which brings us to the fix.

    Rule number one: never miss the cache

    If you skim everything else, read this section.

    Remember how each token needs the maths for all the tokens before it? Providers can store that maths for you. That store is the cache. When it’s warm, the model skips straight to the new bit. It doesn’t re-crunch the whole conversation, it just works out what comes next.

    Say the model has already processed “The quick brown”. With a warm cache, all it has to do is produce “fox”. That’s why cached tokens cost a tenth of the price. You’re paying for a shortcut instead of a marathon.

    The catch: caches go cold

    Holding all that maths in memory costs the provider real money. So they don’t keep it forever. And, just to keep things interesting, every tool has its own timer:

    Where you’re working

    How long the cache lives

    Claude apps (subscription)

    1 hour

    Anthropic API

    5 minutes

    ChatGPT

    Roughly 5 to 10 minutes

    OpenAI API

    30 minutes

    These numbers change, so check your own setup. The point isn’t to memorise the table. It’s to know your number.

    Good news: every hit resets the clock

    Each time you use the cache, the timer starts again. Keep the conversation moving and the cache stays warm indefinitely. On a one-hour window, you could go back and forth all afternoon.

    The coffee-break trap

    Be honest: how many chat tabs do you have open right now? Most of us juggle several sessions, drift away from one, and come back hours later or the next morning. The history is still there, which feels great. But the cache is long gone.

    So the model starts from token one and recomputes the entire conversation before it can write a single new word. On a long session, that’s a nasty hit.

    Here’s how nasty. Ask one follow-up a day in the same thread for a week, and you can easily pay several times more than if you’d asked all seven questions five minutes apart.

    The rule: find out your cache window. Do your back-and-forth inside it. And never walk away from a half-finished task for longer than that window.

    Trim the fat

    Since the whole conversation gets replayed, every extra sentence is a recurring charge. You’ll miss the cache sometimes, no matter how disciplined you are. When you do, a lean session hurts a lot less than a bloated one.

    Say less, and ask it to say less

    Keep your custom instructions and prompts short. Then tell the model to be brief back. One neat trick: ask it to reply in Simplified Technical English (STE). It’s a real international standard, ASD-STE100, built for clear, compact technical writing. Drop a line in your custom instructions asking for STE by default and watch the replies shrink.

    More context isn’t better context

    When million-token context windows arrived, a lot of people treated them like a skip. Throw everything in, let the model sort it out. That backfires twice:

    • Worse answers. Irrelevant material muddies the prediction. Noise in, noise out.

    • Bigger bills. A stuffed context makes every turn heavier, going up and coming back.

    Only give the model what this task actually needs.

    Don’t bring the whole toolbox

    Every tool and skill you switch on quietly lengthens the system prompt. Your session is already “long” before you’ve typed hello. Worse, each tool is a tempting side door. The more you load, the more chances the model has to wander off and call something it didn’t need.

    New task? New chat.

    Easy to say, hard to do. You’re in the flow, one question leads to another, and suddenly you’re on a different topic entirely with a context window full of leftovers. Start a fresh session when the task changes. Two short, focused chats beat one long rambling one, on cost and on quality.

    Good, better, best, ludicrous

    You wouldn’t hire a senior partner to sort the post. So why send every task to the most expensive model?

    Most providers now sell a ladder. Anthropic has Haiku, Sonnet, Opus and Fable. OpenAI’s GPT-5.6 family has Luna, Terra and Sol. Same idea, different names.

    Rung

    Examples

    Send it

    Light

    Haiku, Luna

    Sorting, tagging, pulling data out of documents, summaries

    Middle

    Sonnet, Terra

    The everyday stuff: drafting code and documents, lighter analysis

    Top

    Opus, Fable, Sol

    Deep analysis, and coordinating other agents

    To think, or not to think

    Light models like Haiku usually don’t “reason” by default. That’s a feature and a bug.

    Reasoning models write themselves a private chain of thought before they answer. That hidden scratchpad is a big part of why AI got so much better at analysis, decisions and tool use over the last couple of years. Skip it, and models are more likely to get things wrong or make things up.

    But if the job is “label these 500 emails as sales or support”, you don’t need a philosopher. You need a fast clerk. Paying for reasoning there is paying for thinking nobody asked for.

    The two-second habit

    Before you hit enter, ask: how hard is this, really? Pick the model to match. It’s a tiny pause, and it can stretch a subscription or API budget a very long way. Not everything needs the top shelf.

    Send in the interns

    Subagents have a reputation for torching budgets, and it’s deserved. Ask a swarm of a thousand agents to brainstorm a slogan and you might spend a thousand times what one agent would have.

    But used well, they do the opposite. Think of your main agent as a senior manager, and subagents as interns. You don’t hand an intern the whole project. You hand them one clear job, with just enough context to do it:

    • Summarise this file

    • Make this one yes-or-no call

    • Dig through this folder or database and bring back what’s relevant

    The intern can run on a cheaper model. And the messy work stays out of the manager’s context, so the expensive agent stays focused and lean.

    Don’t let the intern miss the bus

    There’s one trap. If your main agent runs on the Anthropic API, its cache lasts five minutes. If an intern takes seven minutes to report back, the manager’s cache has expired and it has to recompute its whole context. You’ve just spent more than you saved. Only delegate jobs that reliably finish inside the cache window.

    A worked example: the proposal agent

    Imagine an agent that builds client proposals. It runs on a top-tier model like Opus, because pulling a proposal together takes judgement. One of its steps is adding bios of the team members who fit the project.

    That step is perfect intern work. The Opus agent hands a Haiku subagent one instruction: find bios from the team folder for people who work on mobile apps and enterprise back-end systems.

    The intern then:

    1. Opens the team folder.

    2. Reads through dozens of CVs saved as Word files.

    3. Picks out the right people.

    4. Sends back only the relevant bio snippets, and throws the rest away.

    Two wins in one move. The heavy reading happens on a model that’s far cheaper per token. And those dozens of CVs never touch the expensive Opus context. The manager gets exactly what it needs and nothing else.

    Go old school

    LLMs are magic with messy, unstructured information. But plain old code still wins at anything with one right answer. If a step can be done by a script or by a model, the script will be faster, cheaper and the same every time.

    So here’s the principle: if a step can be deterministic, make it deterministic.

    Hide scripts inside your skills

    In tools that support skills, like Claude Cowork, a skill can carry its own Python scripts. The model calls the script, the script does the exact work, and the model spends almost no tokens on it.

    This is a big reason Claude Cowork now handles Word, PowerPoint and Excel files so well, arguably better than Microsoft’s own Copilot. The skills for those formats include scripts that read the files precisely, instead of asking the model to eyeball them.

    Example: let the calculator calculate

    Say your agent needs to add VAT to an invoice total. A model can probably manage it. But:

    • It’s predicting digits, not doing arithmetic, so it might not give the same answer twice.

    • It’ll spend a pile of reasoning tokens getting there.

    A three-line script does it perfectly, instantly, for free. Wrap that in a skill and let the model call it. That’s the sweet spot: models for judgement, code for anything with a right answer.

    Running the same agent daily? Give it a spine.

    If you run the same agent over and over, move it onto a framework with a deterministic backbone, like LangGraph or CrewAI, hosted in the cloud. Then turn as many steps as you can into code. Fewer tokens, fewer surprises, and a much calmer finance team.

    Quick pause for questions

    Before the last big idea, a few questions that always come up at this point.

    “So how is the cache different from the context window?”

    Think of the context window as a suitcase. It has a maximum size, and your whole conversation has to fit inside it. Every turn, you’re carrying the full suitcase to the model and back.

    The cache is what stops that from being ruinous. With it, the model remembers the work it already did on everything in the suitcase and only processes what’s new. Without it, the cost of a conversation would balloon quadratically as it grew. No one could use AI that way, which is why every provider offers caching, and also why they clear it out once it’s been idle too long. Memory in a data centre isn’t cheap.

    “I want to pick this up next week. What’s the smart move?”

    Compress it yourself. Before you log off, ask the model to summarise where you got to: decisions, facts, open questions. Short and portable. Next time, open a fresh chat and paste that in.

    You’d miss the cache either way after a week. So you might as well restart with a tidy brief instead of dragging along every twist and turn of the original conversation.

    “Do memory tools actually save tokens?”

    They can. A decent memory system, whether it’s a vector database hooked up through MCP or an off-the-shelf memory tool, fetches only what’s relevant to the task. Without one, the temptation is to paste in whole documents that are mostly irrelevant. Good memory keeps the signal high and the context small.

    “How do I see where my tokens are actually going?”

    That’s observability, and it depends heavily on your tools.

    • Claude Code gives you proper visibility: full transcripts and OpenTelemetry support.

    • Claude Cowork sits on top of Claude Code but exposes less.

    • Both have a usage command that shows basic session stats, and can produce a more detailed usage report.

    With transcripts and telemetry you can play detective: which turns ate the most tokens, where the cache was missed, which conversations were bloated. For company-wide tracking, look at tools like Langfuse or LangSmith.

    One warning. “Supports OpenTelemetry” can mean a lot or a little. Claude Code sends far more through its telemetry feed than Claude Cowork does. The standard can carry raw prompts, but not every tool that claims support actually sends them. Test before you trust.

    Own your stack

    The last lever is the long game: never get stuck with one supplier.

    The leader today isn’t the leader tomorrow

    Cast your mind back a couple of years. OpenAI looked untouchable. Then Anthropic overtook it, first in enterprise adoption and then in revenue. Almost nobody saw that coming.

    The lesson isn’t “back Anthropic”. It’s that the race is wide open. Top models are converging in quality, and that includes open-source models from China. Kimi K3 and GLM 5.3 now hold their own at coding and agent work, and they cost a fraction of the big US names to run, even on US hosting. If you can switch freely, you can always pay for the best value.

    Your own AI, on your own desk

    Local hardware is catching up fast. Apple’s newly announced Mac Studio M5 Ultra will come with up to 512 GB of memory. That’s enough to run compressed versions of open-source frontier models on your desk, or across a small cluster of them. Lower running costs and your data never leaves the building.

    Keep your crown jewels portable

    If you’re on a commercial platform and can’t leave yet, that’s fine. Just make sure you own what you’ve built on it: your skills, custom tools, prompts and test suites. That’s what makes your agents reliable, and it should be able to move. Tools like Claude Cowork only run Anthropic’s models, so if you ever need to switch, your work needs to come with you.

    Why an open-source harness is worth a look

    • You see everything. Pair OpenTelemetry with an LLM gateway and you can trace every token back to the prompt, workflow and tool call that caused it. Essential once agents run across a whole company.

    • You can switch on the fly. Swap in a cheaper or better model whenever one appears. Nothing keeps providers honest like a customer who can leave.

    • You can go small and local. Run models on your own hardware, including small models fine-tuned for one job.

    That last one is underrated. A small model trained for a single task can match or beat a giant frontier model at that task, for a fraction of the running cost.

    Here’s where this is heading. As top models get closer in quality, AI starts to look like electricity: a utility. The companies that can switch suppliers freely will pay the least. Owning the layer that runs your models is how you get there, and it protects your privacy and your IP along the way.

    Rapid-fire round

    “Should I use more than one model on the same project?”

    Yes, but think of it less as “best tool for the job” and more as “second opinion”. Different models learned from different data, so they have different blind spots. Let one write the code and another try to break it. Let one draft the proposal and another tear it apart. Developers do this routinely now, and it works just as well for business writing.

    “How much do attachments cost me?”

    More than you’d think. If there’s a skill for that file type, the file may be cleaned up before the model sees it. If not, all that raw data turns into tokens, and it rides along in every turn after. Dropping files in casually is one of the quickest ways to bloat a session.

    If you only need a scrap of information from a file, like an address or a couple of form fields, send an intern (a subagent) to fetch it. The rest of the file never touches your main conversation.

    “Which file format is cheapest?”

    • Word documents are among the worst. Open one as plain text and you’ll find a mountain of hidden markup.

    • Markdown is lean and clear.

    • HTML is a dark horse. Models have seen oceans of it, and it carries structure well.

    You can borrow that trick for your custom instructions too. Wrap each section in simple tags, say, one pair around your writing-style notes and another around your formatting rules. It helps the model keep things straight.

    “What about all those token-saving hacks?”

    There’s no shortage. Tools that make the model talk like a caveman. Scripts that squash replies. People uploading screenshots of text instead of the text itself. Some work, most will be obsolete by next quarter.

    That’s why this article sticks to principles. As long as we’re using transformer models, every token depends on the ones before it, and caching will be the thing that saves you money. That won’t change in six months.

    “Are agent loops worth the cost?”

    Sometimes. A brute-force loop that keeps reworking the same prompt until the output improves is expensive, but for something important, like polishing a conference talk submission, it can be worth it.

    A smarter version is the Karpathy loop: instead of refining the output, you refine the process against a measurable target. Given enough rounds, models sometimes find approaches you’d never have thought of.

    And a loop doesn’t have to run nonstop. A daily scheduled task that picks up new information each morning and improves on yesterday’s work is a loop too, and a much cheaper one.

    One last tip: watch for deals

    Pricing chaos cuts both ways. Promotional plans pop up regularly, some offering heavy usage at a steep discount for the first few months. If you’ve got a big project, it pays to shop around. Just remember that intro prices end.

    The cheat sheet

    If your AI bill is starting to feel like a second mortgage, here’s the whole thing on one card:

    • Never miss the cache. Know your window. Keep the conversation moving inside it. Don’t walk away mid-task.

    • Compress before you leave. Get a summary, start fresh next time.

    • Trim the fat. Short prompts, short replies, only relevant context, only the tools you need, new chat for new task.

    • Match the model to the job. Clerks for clerical work, experts for expert work.

    • Use interns wisely. Cheap subagents for narrow jobs that finish inside the cache window.

    • Let code do the maths. Scripts in skills, and a coded backbone for agents you run daily.

    • Watch the meter. Transcripts, usage reports and telemetry show where the money goes.

    • Stay free to leave. Own your skills and prompts, and keep an eye on open-source and local models.

    The hacks will keep changing. The physics won’t. Every new token leans on every token before it, so the cheapest token is the one the model never has to work out twice.

    I hope you’ve enjoyed this and learned something new. I’m always open to suggestions and discussions on LinkedIn. Hit me up with direct messages.

    If you’ve enjoyed my writing and want to keep me motivated, consider leaving stars on GitHub and endorsing me for relevant skills on LinkedIn.

    Till the next one, happy exploring!

    Bill Booth paying Stop toll
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
    Previous ArticleHere’s why NASA is celebrating the failed mission to save the Swift observatory
    Next Article How to Deploy an AI Robot in Your Business: A 5-Step Guide With Real Examples
    • Website

    Related Posts

    AI Tools

    Can an Apartment Search Agent Call the Model Fewer Times and Still Find Good Matches?

    AI Tools

    What the ReLU Revolution Revealed About Biological Plausibility

    AI Tools

    How to Build a 10-Slide Deck in Decktopus AI: A Step-by-Step Walkthrough

    Add A Comment
    Leave A Reply Cancel Reply

    Top Posts

    Open, scalable training infrastructure for large MoEs

    0 Views

    Photon held a funeral for mobile apps. Now it has $4.5M to help replace them with agents

    0 Views

    Can an Apartment Search Agent Call the Model Fewer Times and Still Find Good Matches?

    0 Views
    Stay In Touch
    • Facebook
    • YouTube
    • TikTok
    • WhatsApp
    • Twitter
    • Instagram
    Latest Reviews
    AI Tutorials

    Quantization from the ground up

    AI Tools

    David Sacks is done as AI czar — here’s what he’s doing instead

    AI Reviews

    Judge sides with Anthropic to temporarily block the Pentagon’s ban

    Subscribe to Updates

    Get the latest tech news from FooBar about tech, design and biz.

    Most Popular

    Open, scalable training infrastructure for large MoEs

    0 Views

    Photon held a funeral for mobile apps. Now it has $4.5M to help replace them with agents

    0 Views

    Can an Apartment Search Agent Call the Model Fewer Times and Still Find Good Matches?

    0 Views
    Our Picks

    Quantization from the ground up

    David Sacks is done as AI czar — here’s what he’s doing instead

    Judge sides with Anthropic to temporarily block the Pentagon’s ban

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    Facebook X (Twitter) Instagram Pinterest
    • About Us
    • Contact Us
    • Terms & Conditions
    • Privacy Policy
    • Disclaimer

    © 2026 ainewstoday.co. All rights reserved. Designed by DD.

    Type above and press Enter to search. Press Esc to cancel.